Image compression method, device, storage medium and computer device
By introducing trainable modules into the E2E neural image compression NIC framework and optimizing the parameters of the encoder and decoder, the problem of traditional encoders and decoders being difficult to optimize is solved, and more efficient image compression performance is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TENCENT AMERICA LLC
- Filing Date
- 2023-03-14
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional hybrid video codecs are difficult to optimize as a whole. Improvements to individual modules may not lead to coding gains in overall performance. Furthermore, existing image compression systems have high computational complexity during training and do not produce ideal training results for dissimilar images.
A trainable module is introduced into the end-to-end E2E neural image compression NIC framework. The parameters of the encoder and decoder are adjusted through online training, and the E2E NIC framework is optimized to adapt to various types of quality metrics. The trainable module is used for preprocessing to improve compression performance.
A flexible and universal E2E NIC framework has been implemented, which can adapt to various types of quality metrics, improving the performance and efficiency of image compression.
Smart Images

Figure CN117157646B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Application No. 63 / 325,117, filed March 29, 2022, entitled “Trainable Module for Neural Image Compression”, and U.S. Application No. 18 / 120,612, filed March 13, 2023, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application relates to the field of image processing, and more particularly to an image compression method, apparatus, storage medium, and computer device. Background Technology
[0004] Traditional hybrid video codecs are difficult to optimize as a whole. Improvements to individual modules may not lead to overall coding performance gains. Recently, standards bodies and companies have been actively exploring the potential need for standardization in future video coding technologies. These standards bodies and companies have established the JPEG-AI group, which focuses on AI-based end-to-end neural image compression using deep neural networks (DNNs). The Chinese AVS standard has also established the AVS-AI special group to study neural image and video compression technologies. The recent success of these approaches has already brought increasing industrial benefits to advanced neural image and video compression methodologies.
[0005] In existing technologies, training an image compression system based on the loss between the encoder and decoder involves high computational complexity; furthermore, if the detected image is not very similar to the training image, the training results are not ideal. Therefore, the framework of the image compression system needs to be optimized. Summary of the Invention
[0006] This application relates to an image compression method, apparatus, storage medium, and computer device, providing a flexible and universal end-to-end E2E neural image compression NIC framework that can adapt to various types of quality metrics.
[0007] According to an embodiment of this application, an image compression method is provided, including:
[0008] Receive input image;
[0009] The encoder in the end-to-end E2E neural image compression NIC framework is used to process the entire input image to obtain a first bitstream representation of the entire input image.
[0010] Using the decoder in the E2E NIC framework, the output image is reconstructed from the first bitstream representation of the entire input image;
[0011] The encoder in the E2E NIC framework is optimized by reducing the distortion loss between the input image and the output image; and,
[0012] The input image is processed using an optimized encoder in the E2E NIC framework to obtain a second bitstream representation of the input image.
[0013] According to an embodiment of this application, an image compression apparatus is provided, comprising:
[0014] The receiving module is used to receive input images;
[0015] The encoding module is used to process the entire input image using the encoder in the end-to-end E2E neural image compression NIC framework to obtain a first bitstream representation of the entire input image.
[0016] The decoding module is used to reconstruct the output image from the first bitstream representation of the entire input image using the decoder in the E2E NIC framework;
[0017] An optimization module is used to optimize the encoder in the E2E NIC framework by reducing the distortion loss between the input image and the output image;
[0018] The encoding module is further configured to process the input image using an optimized encoder in the E2E NIC framework to obtain a second bitstream representation of the input image.
[0019] According to an embodiment of this application, a non-volatile computer-readable storage medium is provided, on which instructions are stored, which, when image compression is performed by at least one processor, cause the at least one processor to execute the above-described method.
[0020] According to an embodiment of this application, a computer device is provided, including a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the above-described method.
[0021] According to embodiments of this application, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the described method.
[0022] As can be seen from the above technical solutions, the method provided in the embodiments of this application, by introducing a trainable module, makes the optimized E2E NIC framework flexible and universal, capable of adapting to various types of quality metrics, and achieving better compression performance with the help of the trainable module. Attached Figure Description
[0023] Figure 1 A schematic diagram illustrating an environment for implementing the methods, apparatus, and systems described herein according to embodiments of this application is shown;
[0024] Figure 2 It shows Figure 1 A block diagram of at least one example component of a computer device;
[0025] Figure 3 An optimized end-to-end (E2E) neural image compression (NIC) framework according to an embodiment of this application is shown;
[0026] Figure 4 An example of block-by-block image encoding is shown;
[0027] Figure 5 A flowchart illustrating an optimized end-to-end (E2E) neural image compression (NIC) method according to an embodiment of this application is shown;
[0028] Figure 6 A block diagram of an apparatus for optimizing end-to-end (E2E) neural image compression (NIC) according to an embodiment of this application is shown. Detailed Implementation
[0029] Implementations may include receiving an image, performing an optimization process to adjust certain elements of an end-to-end (E2E) neural image compression (NIC) framework, and encoding the image into a bitstream to optimize rate-distortion performance of the image coding based on the E2E NIC framework. The E2E NIC framework can be an artificial neural network (ANN)-based image or video coding framework that is pre-trained using certain fixed modules and one or more trainable modules. In an ANN-based video coding framework, by performing an online training process on one or more trainable modules, the overall performance of different modules can be further optimized to improve the final goal (e.g., rate-distortion performance), thereby obtaining an E2E-optimized NIC framework.
[0030] like Figure 1 As shown, environment 100 may include user equipment 110, platform 120, and network 130. The devices in environment 100 may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections.
[0031] User equipment 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, user equipment 110 may include computing devices (e.g., desktop computers, laptop computers, tablet computers, handheld computers, smart speakers, servers, etc.), mobile phones (e.g., smartphones, cordless phones, etc.), wearable devices (e.g., a pair of smart glasses or a smartwatch), or similar devices. In some implementations, user equipment 110 may receive information from and / or transmit information to platform 120.
[0032] Platform 120 includes one or more devices capable of generating audio output signals via a multi-band synchronous neural vocoder, as described elsewhere in this application. In some implementations, platform 120 may include a cloud server or a group of cloud servers. In some implementations, platform 120 may be designed to be modular, allowing certain software components to be swapped in or out as needed. Thus, platform 120 can be easily and / or quickly reconfigured for different purposes.
[0033] In some implementations, as shown in the figure, platform 120 may be hosted in a cloud computing environment 122. It is worth noting that although the implementations described herein depict platform 120 as hosted in a cloud computing environment 122, in some implementations, platform 120 may not be cloud-based (i.e., it may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0034] The cloud computing environment 122 includes the environment of the hosting platform 120. The cloud computing environment 122 can provide computing, software, data access, storage, and other services without requiring end users (e.g., user equipment 110) to know the physical location and configuration of one or more systems and / or devices of the hosting platform 120. As shown in the figure, the cloud computing environment 122 may include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resources 124").
[0035] Computing resource 124 includes one or more personal computers, workstations, server devices, or other types of computing and / or communication devices. In some implementations, computing resource 124 may be a hosting platform 120. Cloud resources may include computing instances executing in computing resource 124, storage devices provided in computing resource 124, data transmission devices provided by computing resource 124, etc. In some implementations, computing resource 124 may communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0036] like Figure 1As further shown, computing resources 124 include a set of cloud resources, such as one or more applications (“APP”) 124-1, one or more virtual machines (“VM”) 124-2, virtualized storage (“VS”) 124-3, one or more hypervisors (“HYP”) 124-4, etc.
[0037] Application 124-1 includes one or more software applications that can be provided to or accessed by user equipment 110 and / or sensor device 120. Application 124-1 can eliminate the need to install and execute software applications on user equipment 110. For example, application 124-1 may include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, an application 124-1 may send / receive information to / from one or more other applications 124-1 via virtual machine 124-2.
[0038] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on its usage and its correspondence to any real machine. A system virtual machine can provide a complete system platform supporting the execution of a full operating system (“OS”). A process virtual machine can execute a single program and can support a single process. In some implementations, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.
[0039] Virtualized storage 124-3 includes one or more storage systems and / or one or more devices that utilize virtualization technologies within a storage system or device that uses computing resources 124. In some implementations, the type of virtualization within the context of the storage system may include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage, enabling access to the storage system regardless of physical storage or heterogeneous architecture. Separation allows storage system administrators flexibility in how they manage end-user storage. File virtualization eliminates the dependency between data accessed at the file level and the location where the file is physically stored. This enables performance optimization for storage usage, server consolidation, and / or non-disruptive file migration.
[0040] Hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to execute concurrently on a host computer such as computing resource 124. Hypervisor 124-4 can present a virtual operating platform to the guest operating system and manage the execution of the guest operating system. Multiple instances of various operating systems can share virtualized hardware resources.
[0041] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include cellular networks (e.g., fifth-generation (5G) networks, long-term evolution (LTE) networks, third-generation (3G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMN), local area networks (LAN), wide area networks (WAN), metropolitan area networks (MAN), telephone networks (e.g., public switched telephone network (PSTN)), private networks, self-organizing networks, intranets, the Internet, fiber-optic networks, etc., and / or combinations of these or other types of networks.
[0042] Figure 1 The number and arrangement of devices and networks shown are provided as examples. In practice, there may be more. Figure 1 The devices and / or networks shown may include more devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown can be implemented as multiple distributed devices. Additionally or alternatively, a group of devices in environment 100 (e.g., one or more devices) can perform one or more functions described as being performed by another group of devices in environment 100.
[0043] Figure 2 It shows Figure 1 A block diagram of at least one example component of a computer device.
[0044] Computer device 200 may correspond to user device 110 and / or platform 120. For example... Figure 2 As shown, the computer device 200 may include a bus 210, a processor 220, a memory 230, a storage unit 240, an input unit 250, an output unit 260, and a communication interface 270.
[0045] Bus 210 includes components that allow communication between parts of computer device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing unit. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic storage, and / or optical storage) that stores information and / or instructions for use by processor 220.
[0046] Storage component 240 stores information and / or software related to the operation and use of computer device 200. For example, storage component 240 may include hard disks (e.g., magnetic disks, optical disks, magneto-optical disks, and / or solid-state disks), compact discs (CDs), digital versatile optical discs (DVDs), floppy disks, cassette tapes, magnetic tapes, and / or other types of non-volatile computer-readable storage media, and corresponding drives.
[0047] Input component 250 includes components that allow computer device 200 to receive information, such as via user input (e.g., touchscreen display, keyboard, keypad, mouse, buttons, switches, and / or microphone). Additionally or alternatively, input component 250 may include sensors for sensing information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, and / or actuator). Output component 260 includes components that provide output information from computer device 200 (e.g., a display, speaker, and / or one or more light-emitting diodes (LEDs)).
[0048] Communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable computer device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 270 may allow computer device 200 to receive information from and / or provide information to another device. For example, communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0049] Computer device 200 can execute one or more processes described in this application. Computer device 200 can execute these processes in response to processor 220 executing software instructions stored in a non-volatile computer-readable storage medium such as memory 230 and / or storage unit 240. Computer-readable storage medium is defined in this application as a non-volatile memory device. A memory device includes memory space within a single physical storage device or memory space distributed across multiple physical storage devices.
[0050] Software instructions can be read into memory 230 and / or storage unit 240 via communication interface 270 from another computer-readable storage medium or from another device. When executed, the software instructions stored in memory 230 and / or storage unit 240 can cause processor 220 to perform one or more processes described herein. Additionally or alternatively, hard-wired circuitry may be used in place of or in conjunction with software instructions to perform one or more processes described herein. Therefore, the implementation described herein is not limited to any particular combination of hardware circuitry and software.
[0051] Figure 2 The number and arrangement of components shown are provided as an example. In practice, computer device 200 may include components related to... Figure 2 The components shown may be more, fewer, different, or arranged differently than other components. Additionally or alternatively, a set of components (e.g., one or more components) of computer device 200 may perform one or more functions described as being performed by another set of components of computer device 200.
[0052] Neural Image Compression (NIC) is an end-to-end neural image compression based on artificial intelligence (AI). It uses artificial neural network techniques, such as deep neural networks (DNNs), to compress / decompress images and videos. Given an input image x, the goal of NIC is to use image x as input to a DNN encoder to compute a compressed representation. The compression representation For storage and transmission purposes, compactness is required, and then this compressed representation is used. Used as input to the DNN decoder to reconstruct the image Some NIC methods can employ a variational autoencoder (VAE) structure, where the DNN encoder directly uses the entire image x as its input, which is then processed by a set of network layers that act like a black box to compute the output representation (i.e., a compressed representation). Correspondingly, the DNN decoder will compress the entire representation. As its input, this input is used to compute the reconstructed image through another set of network layers that function like another black box. Use the following objective loss function Optimize the rate-distortion (RD) loss to achieve the reconstructed image. Distortion loss Representation of compression with equilibrium hyperparameter λ The balance between bit consumption R:
[0053]
[0054] Embodiments of this application propose that, for each input image to be compressed, a trainable image processing module, hereinafter referred to as a trainable module, is used to adjust the encoder within the E2E NIC framework to find an optimized E2E NIC framework and compress the input image into a bitstream with better rate-distortion performance. Online training is used to find the optimal parameters of the trainable module, thereby using the optimized module and the encoder in the E2E NIC framework to compress the input image. By using the trainable image processing module, the remainder of the optimized E2E NIC framework can be an image or video coding framework based on an artificial neural network (ANN) that is pre-trained with pre-configured or fixed parameters. The trainable image processing module and the associated methods performed by it can be used as a preprocessing step to improve the compression performance of any E2E NIC framework method. The pre-trained compression model in the E2E NIC framework itself or any training data does not require any training or fine-tuning. Methods and apparatus for preprocessing models will now be described in detail according to one or more embodiments.
[0055] Learning-based image compression can be viewed as a two-step mapping process. The original image x0 in high-dimensional space is mapped by the encoder to a bitstream of length R(x0), and then this bitstream is processed by a compression algorithm with distortion loss. decoder mapping back The original space.
[0056] Figure 3 This is an example diagram illustrating an optimized end-to-end (E2E) neural image compression (NIC) framework 300 according to an embodiment. The E2E NIC framework 300 includes an encoder 310 and a decoder 320. In some embodiments, the encoder 310 and decoder 320 networks have corresponding structures. For example, as... Figure 3 As depicted, encoder 310 is a multi-layer artificial neural network, including, for example, layers #1, #2, ..., #N, such that the output of one layer of artificial neural network is the input of the next layer of artificial neural network. Similarly, decoder 320 is a multi-layer artificial neural network, including, for example, layers #1, #2, ..., #N, such that the output of one layer of artificial neural network is the input of the next layer of artificial neural network.
[0057] In some embodiments, the input image x0 is first processed by the encoder 310 and compressed into a bitstream 340. Then, the decoder 320 processes the bitstream 340 into a reconstructed image. Use at least including distortion loss The loss function is calculated, and this E2E NIC framework is used to correlate several training images to adjust the parameters in the encoder 310 and decoder 320, thereby achieving an optimized E2E NIC framework for other images similar to the training images. However, when using the optimized E2E NIC framework to process images that are not very similar to the training images, the results are often not optimal. On the other hand, the training process of this framework is not only computationally intensive but also time-consuming. Sometimes, recalibrating the optimized E2E NIC framework to process only a small number of images that are similar to the training images is impractical.
[0058] In some embodiments, such as Figure 3 As shown, a trainable module 330 is introduced into an optimized E2E NIC framework. For example, a preprocessing module maps x0 to x′0, and this preprocessing module is added to the front end of the encoder 310. The output x′0 of the trainable module 330 is then processed by the same optimized E2E NIC framework to produce another reconstructed image closer to x0 based on a distance measurement or a loss function (a smaller loss function). In this way, the optimized E2E NIC framework can achieve better compression performance with the help of trainable modules. Optimal compression performance is achieved at the global minimum of equation (1) above.
[0059] The network structure of the trainable module 330 is not limited. In some embodiments, the trainable module 330 is a convolutional neural network (CNN). In other embodiments, the trainable module 330 is a set of residual blocks (ResBlocks). A ResBlock is constructed from a normal network layer connected to a rectified linear unit (ReLU) and an underlying pass-through fed with unchanged information from previous layers. The network portion of the ResBlock may consist of two or more layers of neural networks.
[0060] In some embodiments, online training is used to adjust the parameters of trainable module 330 to achieve better performance of E2E NIC framework 300. In one embodiment, trainable module 330 is a pre-trained network and will be optimized during the optimization of the entire framework. In another embodiment, the parameters of the trainable module are newly initialized and will be adjusted during optimization. For example, trainable module 330 has an initial set of parameters. A set of training images, for example, collected on the Internet, is fed into E2E NIC framework 300 including trainable module 330 to adjust the parameters of trainable module 330 while the rest of E2E NIC framework 300, such as encoder 310 and decoder 320, remains unchanged.
[0061] After the tuning process, the E2E NIC framework 300, including the optimized trainable module 330, can process images similar to the training image set used for online training. It should be noted that the trainable module 330 can be located anywhere on the encoder side. For example, it can be located at the front end of the encoder 310, such as... Figure 3 As shown. In some other cases, the trainable module 330 can be added to the encoder 310. In this way, the decoder-related operations do not need to be changed. The changes only apply to the trainable module on the encoder side.
[0062] Unlike during the model training phase, in this phase, gradients are used to update the parameters of the E2E NIC framework, while the parameters of encoder 310 and decoder 320 are fixed, and gradients can be used to update the trainable module 330 itself. The entire framework becomes differentiable by replacing non-differentiable parts with differentiable parts (e.g., replacing quantization with noise injection) (allowing gradients to propagate back). Therefore, the above optimization can be solved iteratively using gradient descent.
[0063] This preprocessing model has two key hyperparameters: stride and number of steps. The stride indicates the "learning rate" during online training. Images with different content types can correspond to different strides to achieve optimal optimization results. The number of steps indicates the number of updates performed. These hyperparameters are then compared with the target loss function. Both are used in the learning process. For example, the step size can be used in gradient descent algorithms or backpropagation calculations performed during the learning process. The number of iterations can be used as a threshold for the maximum number of iterations to control when the learning process can be terminated.
[0064] Figure 4 The illustration shows an example of block-by-block image encoding.
[0065] In the example embodiment, image 400 can first be segmented into blocks (by...) Figure 4 (as shown by the dashed lines in the image), and can compress segmented blocks instead of the image 400 itself. Compressed blocks are in... Figure 4The blocks are represented by shaded areas, and the blocks to be compressed are not shaded. The sizes of the segmented blocks can be equal or unequal. The stride of each block can be different. For this purpose, different strides can be assigned to image 400 to achieve better compression results. Block 410 is an example of a segmented block with height h and width w.
[0066] In the example embodiment, the image can be compressed without being segmented into blocks, and the entire image is the input to the E2E NIC model. Different images can have different strides to achieve optimized compression results.
[0067] In another example embodiment, the step size can be selected based on characteristics of the image (or patch) (e.g., the RGB variance of the image). In this embodiment, RGB may refer to the red-green-blue color model. Further, in another example embodiment, the step size can be selected based on the RD performance of the image (or patch). Therefore, according to embodiments thereof, the trainable module 330 can be optimized based on multiple step sizes, and the step size with better compression performance can be selected to optimize the trainable module 330.
[0068] Figure 5 This is a flowchart of an optimized end-to-end (E2E) neural image compression (NIC) method 500 according to an embodiment.
[0069] In some implementations... Figure 5 One or more process frames can be executed by platform 120. In some implementations, Figure 5 One or more process frames may be executed by another device or group of devices that are separate from or include the platform 120, such as user equipment 110.
[0070] like Figure 5 As shown, in operation S510, method 500 includes receiving an input image.
[0071] In operation S520, method 500 includes processing the entire input image using the encoder in the E2E NIC framework to obtain a first bitstream representation of the entire input image. In some embodiments, a trainable module exists on the encoder side of the E2E NIC framework. The parameters of this trainable module can be adjusted to achieve better performance of the E2E NIC framework as described above.
[0072] In operation S530, method 500 includes reconstructing an output image from a first bitstream representation of the entire input image using a decoder in an E2E NIC framework. In some embodiments, a loss function is generated based on the output image and the input image.
[0073] In operation S540, method 500 includes optimizing the encoder in the E2E NIC framework by reducing the distortion loss between the input and output images. In some embodiments, an online training process is used to optimize the encoder in the E2E NIC framework to find one or more optimal parameters in the trainable module. For example, a set of training images can be collected from the Internet to train the parameters in the trainable module, while the parameters in the rest of the E2E NIC framework remain unchanged.
[0074] In some embodiments, the online training process includes adjusting the trainable module according to the learning rate of the input image when at least a predefined number of iterations are performed by the online training process, and / or when the distortion loss between the input image and the output image is less than a predefined threshold. In other words, the encoder includes at least one fixed module whose parameters are not affected by the online training process.
[0075] In operation S550, method 500 includes processing the input image using an optimized encoder in the E2E NIC framework to obtain a second bitstream representation of the input image. In some embodiments, the second bitstream and / or a compressed representation of the input image may be transmitted to, for example, a decoder and / or a receiving device.
[0076] Method 500 may further include segmenting the input image into one or more blocks. In this case, operations 520 to 550 are performed on each block, rather than the entire input image. That is, method 500 further includes training a model based on the E2E NIC framework to determine the best trainable module for each segmented block, thereby encoding the block to generate a block bitstream. The segmented blocks may have the same size or different sizes, and each block may have a different learning rate.
[0077] Method 500 can use an artificial neural network based on a pre-trained image coding model, where the parameters of the artificial neural network are fixed and gradients are used to update the input image.
[0078] although Figure 5 An example box of the method is shown, but in some implementations, the method may include... Figure 5 The boxes depicted may be fewer, different, or arranged differently compared to additional boxes. Alternatively, the method may be applied to two or more boxes in parallel.
[0079] Figure 6 This is a block diagram of an apparatus for optimizing end-to-end (E2E) neural image compression (NIC) according to an embodiment.
[0080] like Figure 6As shown, the device 600 includes a receiving code 610, an encoding code 620, a decoding code 630, and an optimization code 640.
[0081] Receive code 610 is configured to enable at least one processor to receive the input image.
[0082] Encoding code 620 is configured to cause at least one processor to: (i) process the entire input image using an encoder in the E2E NIC framework to obtain a first bitstream representation of the entire input image, and (ii) process the input image using an optimized encoder in the E2E NIC framework to obtain a second bitstream representation of the input image.
[0083] Decoding code 630 is configured to enable at least one processor to use a decoder in the E2E NIC framework to reconstruct the output image from the first bitstream representation of the entire input image.
[0084] Optimization code 640 is configured to enable at least one processor to optimize the encoder in the E2E NIC framework by reducing distortion loss between the input and output images.
[0085] The apparatus 600 may further include segmentation code configured to cause at least one processor to segment the input image into one or more blocks. In this case, encoding code 620, decoding code 630, and optimization code 640 are performed using each of the segmented blocks instead of the entire input image.
[0086] Additionally, the device 600 can use an artificial neural network based on a pre-trained image coding model, wherein the parameters of the artificial neural network are fixed and gradients are used to update the input image.
[0087] although Figure 6 An example block of the device is shown, but in some embodiments, the device may include [the following]. Figure 6 The frames depicted may be additional frames, fewer frames, different frames, or frames with different arrangements compared to other frames. Alternatively or optionally, two or more frames of the device may be combined.
[0088] The embodiments described in this paper illustrate an E2E image compression method. This method improves the coding efficiency of E2E NIC by introducing a trainable image processing module through a flexible and universal framework that adapts to various types of quality metrics.
[0089] The E2E image compression method according to one or more embodiments can be used alone or in any order. Furthermore, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuitry system (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-volatile computer-readable medium.
[0090] The above disclosure provides illustrations and descriptions, but is not intended to be exhaustive or to limit the implementation to the precise form disclosed. Modifications and variations can be made based on the above disclosure, or modifications and variations can be derived from the practice of implementation.
[0091] As used in this application, the term "component" is intended to be interpreted broadly as hardware, firmware, or a combination of hardware and software.
[0092] Obviously, the systems and / or methods described in this application can be implemented in different forms of hardware, firmware, or combinations of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods is not a limitation on the implementation method. Therefore, this application describes the operation and behavior of the systems and / or methods without referring to any specific software code—it should be understood that software and hardware can be designed to implement the systems and / or methods based on the descriptions in this application.
[0093] Even if specific combinations of features are listed in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically stated in the claims and / or not disclosed in the specification. While each dependent claim listed below may directly depend on only one claim, the disclosure of possible implementations includes every dependent claim in combination with all other claims in the claim set.
[0094] Unless expressly stated otherwise, no element, action, or instruction used in this application should be construed as critical or essential. Furthermore, as used in this application, the term "group" is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with "one or more". Additionally, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on".
Claims
1. An image compression method, characterized in that, The end-to-end E2E neural image compression NIC framework includes a trainable module, an encoder, and a decoder, and the method includes: Receive input image; Using the trainable module, the input image is mapped to a first image, and the first image is input to the encoder to obtain a first bitstream representation of the input image; Using the decoder, the output image is reconstructed from the first bitstream representation; The input image is segmented into one or more blocks, and an online training process is used to determine the optimal parameters in the trainable modules for each block, wherein... The parameters of the encoder and the decoder are fixed, and the gradient is used to update the trainable module itself. The hyperparameters of the trainable module include a step size, wherein the step size is selected based on the red-green-blue variance of the block.
2. The method according to claim 1, characterized in that, The trainable module is an artificial neural network.
3. An image compression device, characterized in that, The end-to-end E2E neural image compression NIC framework includes a trainable module, an encoder, and a decoder, and the device includes: The receiving module is used to receive input images; An encoding module is used to map the input image to a first image using the trainable module, and input the first image into the encoder to obtain a first bitstream representation of the input image; A decoding module is used to reconstruct an output image from the first bitstream representation using the decoder; An optimization module is used to segment the input image into one or more blocks and, using an online training process, determine the optimal parameters in the trainable module for each block, wherein the parameters of the encoder and the decoder are fixed, and gradients are used to update the trainable module itself; the hyperparameters of the trainable module include a step size, wherein the step size is selected based on the red-green-blue variance of the block.
4. The apparatus according to claim 3, characterized in that, The trainable module is an artificial neural network.
5. A non-volatile computer-readable storage medium, characterized in that, It stores instructions that, when image compression is performed by at least one processor, cause the at least one processor to perform the method as described in claim 1 or 2.
6. A computer device, characterized in that, It includes a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the method as described in claim 1 or 2.