System, method, and computer program for iterative content-adaptive online training in neural image compression
The iterative content-adaptive online training method for neural image compression addresses the optimization challenges of conventional frameworks by fine-tuning and enhancing the neural image compression framework based on input images, resulting in improved performance and reduced computational costs.
Patent Information
- Application Number
- JP2023558148
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-09-26
- Filing Date
- 2022-10-20
- Publication Date
- 2025-06-23
- Estimated Expiration
- 2042-10-20
AI Technical Summary
Conventional hybrid video codecs and neural network-based image/video compression frameworks face challenges in optimizing overall performance, leading to increased computational costs and rate-distortion losses when adapting to various compression frameworks.
A method for iterative content-adaptive online training in neural image compression, which involves fine-tuning an end-to-end neural image compression framework based on input images, calculating parameter updates, enhancing the framework with extended neural networks, and generating an updated framework to improve compression performance.
This approach enhances the overall performance of neural image compression by adaptively optimizing the compression framework for specific input images, reducing computational costs and rate-distortion losses, and improving visual quality.
Smart Images

Figure 0007697033000025 
Figure 0007697033000026 
Figure 0007697033000027
Abstract
Description
Technical Field
[0001] Cross - reference to related applications This application is based on and claims priority to U.S. Provisional Patent Application No. 63 / 289,055, filed on December 13, 2021, and U.S. Patent Application No. 17 / 952,865, filed on September 26, 2022, the entire contents of which are incorporated herein by reference.
Background Art
[0002] Conventional hybrid video codecs are difficult to optimize as a whole. Improving only a single module may not improve the overall performance of the coding gain. Recently, standardization groups and companies have been actively exploring potential needs for the standardization of future video coding technologies. These standardization groups and companies have established the JPEG - AI group, which focuses on AI - based end - to - end neural image compression using deep neural networks (DNNs). China's Audio - Video Coding Standard (AVS) has also established an AVS - AI special group to work on neural image and video compression technologies. Due to the success of recent approaches, the industry's interest in advanced neural image and video compression techniques is increasing.
[0003] However, in related technologies, neural network - based video or image coding frameworks are limited to specific types of compression frameworks. To adapt to various types of frameworks, conventional systems may result in increased computational memory / cost and increased rate - distortion loss, and as a result, the overall performance of the image or video framework / process decreases.
[0004] Therefore, there is a need for a method to optimize the coding framework and improve the overall performance.
Summary of the Invention
[0005] According to an embodiment, a method for iterative content-adaptive online training in neural image compression is provided.
[0006] According to one aspect of the present disclosure, a content-adaptive online training method for end-to-end (E2E) neural image compression (NIC) using a neural network executed by at least one processor is provided. The method includes receiving an input image into the E2E NIC framework; fine-tuning the E2E NIC framework based on the input image; calculating a parameter update using a first neural network of the fine-tuned E2E NIC framework; enhancing the fine-tuned E2E NIC framework based on a second neural network which is the extended network; and generating an updated E2E NIC framework based on the enhanced E2E NIC framework and the parameter update.
[0007] The method may further include encoding the input image and the parameter update to generate a compressed representation of the input image and the parameter update; decoding the compressed representation of the parameter update to generate a decoded parameter update; updating the E2E NIC framework based on the decoded parameter update; and decoding the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image.
[0008] The method may further include determining a distortion loss of the reconstructed image based on a consumption amount of the compressed representation of the input image and the parameter update, a trade-off hyperparameter, and a distortion of a decoded block residual of the reconstructed image.
[0009] The method may further include dividing the input image into one or more blocks.
[0010] In some embodiments, parameter update includes a learning rate and a number of steps, and the learning rate and the number of steps are selected based on the characteristics of the input image. Further, the characteristics of the input image are one of the RGB variance of the input image and the RD performance of the input image.
[0011] In some embodiments, the extended network is a set of convolutional neural networks or a layer of convolutional neural networks.
[0012] In some embodiments, one or more extended networks are used to extend the fine-tuned E2E NIC framework.
[0013] According to another aspect of the present disclosure, there is provided an apparatus for a content E2E NIC using a neural network, the apparatus including at least one memory configured to store computer program code, and at least one processor configured to read the computer program code and operate according to the instructions of the computer program code. The computer program code includes reception code configured to cause at least one processor to receive an input image to the E2E NIC framework; fine-tuning code configured to cause at least one processor to fine-tune the E2E NIC framework based on the input image; calculation code configured to cause at least one processor to calculate parameter update using a first neural network of the fine-tuned E2E NIC framework; extension code configured to cause at least one processor to extend the fine-tuned E2E NIC framework based on a second neural network that is an extended network; and generation code configured to cause at least one processor to generate an updated E2E NIC framework based on the extended E2E NIC framework and the parameter update.
[0014] The apparatus can further include code configured to cause at least one processor to encode an input image and a parameter update to generate a compressed representation of the input image and the parameter update; decode the compressed representation of the parameter update to generate a decoded parameter update; update an E2E NIC framework based on the decoded parameter update; and decode the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image.
[0015] The apparatus can further include code configured to cause at least one processor to determine a distortion loss of the reconstructed image based on a consumption amount of the compressed representation of the input image and the parameter update, a trade-off hyperparameter, and a distortion of a decoded block residual of the reconstructed image.
[0016] The apparatus can further include code configured to cause at least one processor to divide the input image into one or more blocks.
[0017] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable medium storing instructions that are executed by at least one processor of an apparatus for content-adaptive online training for an E2E NIC using a neural network. The instructions cause the at least one processor to receive an input image to an E2E NIC framework; fine-tune the E2E NIC framework based on the input image; calculate a parameter update using a first neural network of the fine-tuned E2E NIC framework; expand the fine-tuned E2E NIC framework based on a second neural network that is an expanded network; and generate an updated E2E NIC framework based on the expanded E2E NIC framework and the parameter update.
[0018] A non-transitory computer-readable medium can further include instructions for causing at least one processor to divide an input image into one or more blocks and compress the one or more blocks individually.
[0019] A non-transitory computer-readable medium can further include instructions for causing at least one processor to encode an input image and a parameter update to generate a compressed representation of the input image and the parameter update; decode the compressed representation of the parameter update to generate a decoded parameter update; update an E2E NIC framework based on the decoded parameter update; and decode the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image.
[0020] A non-transitory computer-readable medium can further include instructions for causing at least one processor to determine a distortion loss of a reconstructed image based on a consumption amount of a compressed representation of the input image and the parameter update, a trade-off hyperparameter, and a distortion of a decoded block residual of the reconstructed image.
[0021] Additional embodiments are described in the following description, become apparent in part from the description, and / or can be realized by practicing the embodiments presented in this disclosure.
Brief Description of the Drawings
[0022]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
[0023] The following detailed description of exemplary embodiments refers to the accompanying drawings. The same reference numerals in different drawings may identify the same or similar elements.
[0024] The above disclosure provides examples and explanations, but is not intended to be exhaustive or to limit the embodiments to the exact forms disclosed. Modifications and variations are possible in light of the above disclosure or can be obtained from the practice of the embodiments. Further, one or more features or components of one embodiment can be incorporated into or combined with another embodiment (or one or more features of another embodiment). Additionally, in the flowcharts and descriptions of operations provided below, it is understood that one or more operations may be omitted, one or more operations may be added, one or more operations may be executed simultaneously (at least partially), and the order of one or more operations may be changed.
[0025] The systems and / or methods described herein will clearly be implementable in different forms of hardware, software, or combinations of hardware and software. It is understood that the actual specific control hardware or software code used to implement these systems and / or methods does not limit the embodiments. Thus, the operations and behaviors of the systems and / or methods have been described herein without reference to specific software code. It is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0026] Even if specific combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible embodiments. In fact, many of these features can be combined in ways not particularly recited in the claims and / or not disclosed in the specification. Each of the dependent claims listed below may depend directly on only one claim, but the disclosure of possible embodiments includes each dependent claim combined with all other claims in the claim set.
[0027] The proposed features discussed below may be used individually or in any order of combination. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0028] Elements, acts, or instructions used in this specification should not be construed as important or essential unless expressly described as such. Also, as used in this specification, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more." When only one item is intended, the term "one" or a similar term is used. Also, as used in this specification, terms such as "has," "have," "having," "include," "including," etc. are intended to be open-end terms. Further, the phrase "based on" is intended to mean "at least partially based on" unless expressly stated otherwise. Further, expressions such as "at least one of [A] and [B]" or "at least one of [A] or [B]" should be understood to include only A, only B, or both A and B.
[0029] Exemplary embodiments of the present disclosure provide a method and apparatus for block-wise content-adaptive online training with post-filtering for an end-to-end (E2E) neural image compression (NIC) optimization network. The E2E optimization network may be, for example, an artificial neural network (ANN)-based image coding framework. In an ANN-based video coding framework, various modules are jointly optimized from input to output by performing a machine learning process to improve the final goal (e.g., rate-distortion performance), resulting in an E2E-optimized NIC.
[0030] FIG. 1 is a flowchart showing an overall overview of an iterative content-adaptive online training process for an E2E NIC, such as a content-adaptive online training NIC framework or a content-adaptive online training system, according to an embodiment.
[0031] First, an input image (or video sequence) is received (S110). The input image can be divided into blocks, for example. Block-based image coding can be performed to compress the blocks. In S120, the online training process fine-tunes the NIC framework and generates parameter updates. The NIC framework may be a pre-trained framework. In S130, based on the generated parameter updates, the parameters of the NIC framework are updated. The parameter updates may include, for example, a step size (i.e., learning rate) and the number of steps, but are not limited thereto. In S140, the NIC framework can be further enhanced using the enhanced network. The enhanced network is used to improve the visual quality of the image. Next, the input image and the generated parameter updates are encoded, for example, by a DNN encoder (S150), and then decoded, for example, by a DNN decoder (S160). The updated NIC framework is generated using the decoded parameter updates to update the NIC framework (S170). Finally, the final image is generated by decoding using the decoder of the updated NIC framework. That is, in S180, a decoded image is generated based on the updated NIC framework.
[0032] Figure 2 is a diagram of an environment 200 that can implement the methods, apparatuses, and systems described herein, according to an embodiment.
[0033] As shown in Figure 2, the environment 200 can include a user device 210, a platform 220, and a network 230. The devices in the environment 200 can be interconnected via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0034] The user device 210 includes one or more devices that can receive, generate, store, process, and / or provide information related to the platform 220. For example, the user device 210 can include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses, or a smartwatch), or a similar device. In some embodiments, the user device 210 can receive information from the platform 220 and / or transmit information to the platform 220.
[0035] The platform 220 includes one or more devices, as described elsewhere in this specification. In some embodiments, the platform 220 can include a cloud server or a group of cloud servers. In some embodiments, the platform 220 can be modularly designed such that software components can be swapped in or out. Thus, the platform 220 can be easily and / or quickly reconfigured for various applications.
[0036] In some embodiments, as shown, the platform 220 may be hosted in a cloud computing environment 222. In particular, in the embodiments described herein, the platform 220 is described as being hosted in the cloud computing environment 222, but in some embodiments, the platform 220 may not be cloud-based (i.e., can be implemented outside of a cloud computing environment), or may be partially cloud-based.
[0037] The cloud computing environment 222 includes an environment that hosts the platform 220. The cloud computing environment 222 can provide services such as computing, software, data access, storage, etc. without requiring the end user (e.g., user device 210) knowledge regarding the physical location and configuration of the system and / or device that hosts the platform 220. As shown, the cloud computing environment 222 can include a group of computing resources 224 (collectively referred to as "computing resources 224" and individually referred to as "computing resource 224").
[0038] The computing resources 224 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some embodiments, the computing resources 224 can host the platform 220. The cloud resources can include computing instances executed within the computing resources 224, storage devices provided within the computing resources 224, data transfer devices provided by the computing resources 224, etc. In some embodiments, the computing resources 224 can communicate with other computing resources 224 via a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0039] Furthermore, as shown in FIG. 2, the computing resources 224 include a group of cloud resources such as one or more applications (APP) 224-1, one or more virtual machines (VM) 224-2, virtualized storage (VS) 224-3, or one or more hypervisors (HYP) 224-4, etc.
[0040] Application 224-1 includes one or more software applications that can be provided to or accessed by user device 210 and / or platform 220. Application 224-1 can eliminate the need to install and execute software applications on user device 210. For example, Application 224-1 can include software associated with platform 220 and / or any other software that can be provided via cloud computing environment 222. In some embodiments, one Application 224-1 can transmit / receive information between one or more other Applications 224-1 via virtual machine 224-2.
[0041] Virtual machine 224-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 224-2 can be either a system virtual machine or a process virtual machine, depending on the degree of use and correspondence of virtual machine 224-2 to a real machine. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system (OS). A process virtual machine can execute a single program and support a single process. In some embodiments, virtual machine 224-2 can execute on behalf of a user (e.g., user device 210) and manage the infrastructure of cloud computing environment 222, such as data management, synchronization, or long-term data transfer.
[0042] The virtualized storage 224-3 includes one or more storage systems and / or one or more devices that use virtualization technology within a storage system or device of computing resources 224. In some embodiments, within the context of a storage system, types of virtualization can include block virtualization and file virtualization. Block virtualization refers to the abstraction (or separation) of logical storage from physical storage so that the storage system can be accessed regardless of the physical storage or heterogeneous structure. This separation allows the administrator of the storage system to gain flexibility in how to manage the end-user's storage. File virtualization can eliminate the dependency between the data accessed at the file level and the location where the files are physically stored. This can enable optimization of storage usage, server integration, and / or the performance of non-disruptive file migration.
[0043] The hypervisor 224-4 can provide hardware virtualization technology that enables multiple operating systems (e.g., "guest operating systems") to run simultaneously on a host computer such as computing resources 224. The hypervisor 224-4 can present a virtual operating platform to the guest operating systems and manage the execution of the guest operating systems. Multiple instances of various operating systems can share the virtualized hardware resources.
[0044] Network 230 includes one or more wired and / or wireless networks. For example, network 230 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, an optical fiber-based network, etc., and / or a combination of these networks or other types of networks.
[0045] An example is provided of the number and arrangement of the devices and networks shown in FIG. 2. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with different arrangements other than those shown in FIG. 2. Further, two or more of the devices shown in FIG. 2 may be implemented within a single device, and a single device shown in FIG. 2 may be implemented as a plurality of distributed devices. Additionally, or alternatively, a set of devices (e.g., one or more devices) in environment 200 may perform one or more functions described as being performed by another set of devices in environment 200.
[0046] FIG. 3 is a block diagram of an example of components of one or more of the devices of FIG. 2.
[0047] Device 300 may correspond to user device 210 and / or platform 220. As shown in FIG. 3, device 300 may include a bus 310, a processor 320, a memory 330, a storage component 340, an input component 350, an output component 360, and a communication interface 370.
[0048] Bus 310 includes components that permit communication among components of apparatus 300. Processor 320 is implemented in hardware, software, or a combination of hardware and software. Processor 320 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some embodiments, processor 320 includes one or more processors that can be programmed to execute functions. Memory 330 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions used by processor 320.
[0049] Storage component 340 stores information and / or software related to the operation and use of apparatus 300. For example, storage component 340, along with a corresponding drive, can include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disk (CD), a digital versatile disk (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium.
[0050] The input component 350 includes components that enable the device 300 to receive information via user input (e.g., touch screen, keyboard, keypad, mouse, button, switch, and / or microphone), etc. Additionally, or alternatively, the input component 350 may include sensors for sensing information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, and / or actuator). The output component 360 includes components that provide output information from the device 300 (e.g., a display, speaker, and / or one or more Light Emitting Diodes (LEDs)).
[0051] The communication interface 370 includes a transceiver-like component (e.g., a transceiver and / or separate receiver and transmitter) that enables the device 300 to communicate with other devices via a wired connection, wireless connection, or a combination of wired and wireless connections, etc. The communication interface 370 may enable the device 300 to receive information from another device and / or provide information to another device. For example, the communication interface 370 may include an Ethernet interface, optical interface, coaxial interface, infrared interface, Radio Frequency (RF) interface, Universal Serial Bus (USB) interface, Wi-Fi interface, or cellular network interface, etc.
[0052] The device 300 can execute one or more processes described herein. The device 300 can execute these processes in response to the processor 320 executing software instructions stored by a non-transitory computer-readable medium such as the memory 330 and / or the storage component 340. The computer-readable medium is defined herein as a non-transitory storage device. The memory device includes a memory space within a single physical storage device or a memory space spanning multiple physical storage devices.
[0053] Software instructions may be read into memory 330 and / or storage component 340 from another computer-readable medium or from another device via communication interface 370. When executed, the software instructions stored in memory 330 and / or storage component 340 can cause processor 320 to execute one or more processes described herein. Additionally, or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to execute one or more processes described herein. Thus, the embodiments described herein are not limited to a particular combination of hardware circuitry and software.
[0054] The number and arrangement of components shown in FIG. 3 are provided as an example. In practice, device 300 can include additional components, fewer components, different components, or components in a different arrangement than those shown in FIG. 3. Additionally, or alternatively, a set of components (e.g., one or more components) of device 300 can perform one or more functions described as being performed by another set of components of device 300.
[0055] In an embodiment, any one of the operations or processes of FIGS. 4-9 can be implemented by or using any one of the elements shown in FIGS. 2 and 3.
[0056] According to some embodiments, a general process for neural network-based image compression can be as follows. Given an image or video sequence x, the goal of the NIC is to use the image x as input to a DNN encoder and compute a compact compressed representation
Number
Number
Number
Number
Number
Number
Number
Number
Number
[0057]
Number
[0058] Embodiments relate to iterative content - adaptive E2E online training of the NIC framework. First, the input image x is segmented by optimizing the rate - distortion performance for input blocks. During online training, some (or all) of the parameters of a pre - trained network can be updated. The pre - trained network may be based on a neural network. The updated partial (or all) parameters are encoded into a bitstream along with the encoded input image (i.e., the compressed representation
Number
[0059] Next, a detailed description of the pre - processing of the iterative content - adaptive online training NIC framework according to one or more embodiments will be described.
[0060] As described above, the pre - trained NIC framework is fine - tuned based on the input image. Next, parameter updates are obtained using the fine - tuned NIC framework, and the NIC framework is updated thereby. In this way, the NIC framework can be adapted to the target image content. When the NIC framework is fine - tuned, one or more network parameters can be updated.
[0061] In some embodiments, the parameters may be updated wholly or partially. For example, the parameters may be updated only in one module (such as a context model or a hyperdecoder) of the NIC framework. As another example, the parameters may be updated for multiple or all modules of the NIC framework.
[0062] In some embodiments, only the bias term is optimized and updated. In another exemplary embodiment, the coefficient (weight) term is optimized. Alternatively, for example, all parameters may be optimized.
[0063] In some embodiments, the NIC framework is fine-tuned, and the updated NIC framework is generated based on a single input image. In some embodiments, the NIC framework is fine-tuned, and the fine-tuned NIC framework is used to generate an updated framework based on a set of input images.
[0064] The fine-tuning process includes a plurality of epochs in which the parameters are updated in this iterative online training process. Fine-tuning is stopped when the training loss (e.g., determined based on the target loss function of Equation 1) has flattened or is currently flattening. The iterative content-adaptive online training NIC framework has two main hyperparameters: the step size and the number of steps. The step size indicates the "learning rate" of the online training NIC framework. Images containing various types of content may correspond to various step sizes to achieve the best optimization results. The number of steps indicates the number of updates to be made. Along with the target loss function (Equation 1), the hyperparameters are used in the online learning process. For example, the step size can be used in the gradient descent algorithm or backpropagation calculations performed in the learning process. The number of iterations can be used as a threshold for the maximum number of iterations to control when the learning process can end. In some embodiments, during the iterative online training process, the learning rate (i.e., the step size) can be changed at each step by a scheduler. The scheduler determines the value of the learning rate, which can increase, decrease, or remain the same over several intervals. There may be a single scheduler or multiple (different) schedulers for different input images. Multiple parameter updates can be generated based on multiple learning rate schedulers, and a scheduler with better compression performance for each parameter update can be selected. At the end of the fine-tuning process, the parameter updates are calculated. In some embodiments, next, the parameter updates are compressed at the end of the fine-tuning process. For example, a compression algorithm (such as LZMA2) can be used to compress the parameter updates. In another exemplary embodiment, the compression of the parameter updates is not performed.
[0065] In some embodiments, the parameter update is calculated as the difference between the fine-tuned parameters and the pre-trained parameters. In some embodiments, the parameter update is the fine-tuned parameters. In another exemplary embodiment, the parameter update is some transformation of the fine-tuned parameters.
[0066] FIG. 4 shows an example of block-based image coding. In the iterative content-adaptive online training NIC framework according to an embodiment, instead of directly encoding the entire input image, a block-based coding mechanism can be used to compress the image frame. Using a block-based coding mechanism, the entire input image is first partitioned into blocks of the same (or various) sizes, and the blocks are compressed individually.
[0067] As shown in FIG. 4, the image 400 is first divided into blocks (indicated by the dashed lines in FIG. 4), and instead of the image 400 itself, the divided blocks are compressed. The compressed blocks are shown shaded in FIG. 4, and the blocks to be compressed are not shaded. The divided blocks may or may not be of the same size. The step size of each block may be different. For this purpose, different step sizes can be assigned to the image 400 to achieve better compression results. Block 410 is an example of one of the divided blocks having a height h and a width w. The blocks go through a block-based image coding process, and a bitstream of encoded information is generated.
[0068] In some embodiments, the image can be compressed without being divided into blocks and the entire image can be used as the input to the E2E NIC model. To achieve optimized compression results, the step size may be different if the images are different.
[0069] The step size (i.e., the learning rate of the content-adaptive online training NIC framework) can be selected based on the characteristics of the image (or block). For example, the characteristics of the image can be based on the red-green-blue (RGB) color model and the RGB variance of the image. Further, in some embodiments, the step size can be selected based on the RD performance of the image (or block). Thus, according to that embodiment, a plurality of parameter updates can be generated based on different step sizes, and a step size with better compression performance can be selected for each parameter update.
[0070] To achieve better compression results, a plurality of learning rate schedulers can be assigned to different blocks. In some embodiments, all blocks share the same learning rate schedule. The selection of the learning rate scheduler can also be based on the characteristics of the block, such as the RGB variance of the block or the RD performance of the block.
[0071] According to an embodiment, different blocks can update different parameters within different modules (e.g., a context module or a hyper decoder), or different types of parameters (biases or weights) of the iterative content-adaptive online training NIC framework. In some embodiments, all blocks share the same parameter update. The parameters (to be updated) can be selected based on the characteristics of the block, such as the RGB variance of the block or the RD performance of the block.
[0072] If the blocks are different, different methods can be selected to convert the parameter update. For example, in some embodiments, one block can choose to update the parameters of the NIC framework based on the difference between the fine-tuned parameters and the pre-trained parameters. Another block may choose to directly update the parameters. In some embodiments, all block parameters are updated in the same way. The method of converting the parameter update can be selected based on the characteristics of the block, such as the RGB variance of the block or the RD performance of the block.
[0073] If the blocks are different, different methods can be selected to compress the parameter update. For example, one block can compress the parameter update using the LZMA2 algorithm. Another block can compress the parameter update using the bzip2 algorithm. The embodiments are not limited to this, and any compression algorithm suitable for compressing the parameters can be used. In some embodiments, all blocks use the same method to compress (or not compress) the parameter update. The compression method can be selected based on the characteristics of the block, such as the RGB variance of the block or the RD performance of the block.
[0074] Each of the compressed images or blocks may use the extended network to improve the visual quality. The extension process can be the same as the process implemented to update the NIC framework parameters. In some embodiments, the extended network includes a set of convolutional neural networks. In some embodiments, the extended network consists of only one convolutional neural network layer. The extended network can be pre-trained by the training dataset. In another embodiment, the extended network is not pre-trained.
[0075] In some embodiments, each image / block to be compressed uses one enhanced network to improve visual quality. In another embodiment, each image / block to be compressed uses multiple enhanced networks to iteratively improve visual quality. That is, the image / block is enhanced one after another until no gain can be obtained.
[0076] The coding process of the iterative content-adaptive online training NIC framework applied to the image / block to generate the reconstructed image will be described with reference to FIG. 5.
[0077] FIG. 5 is an example flowchart of the coding process according to an embodiment.
[0078] First, at S510, the NIC framework encodes the input image and parameter updates. Subsequently, the encoded input and the encoded parameter updates are decoded (S520). When the parameter updates are compressed (yes at S530), the parameter updates obtained from the online training process are first decompressed (S540). When the parameter updates are not compressed (no at S530), the process proceeds to S550. At S550, the NIC framework is updated on the decoder side using the decoded parameter updates from S520 or the decompressed decoded parameter updates from S540. Finally, at S560, the updated NIC framework decoder is used to perform image decoding (to generate the reconstructed image). Based on how the parameters are transformed, the original (pre-trained) bias term is updated with the updated parameter values.
Number
[0079] The embodiments impose no limitations on any method used for, for example, neural encoders, encoders, decoders, and neural decoders. According to the embodiments, an iterative content-adaptive online training method can be compatible with different types of NIC frameworks. For example, this process can be executed using different types of encoding and decoding DNNs.
[0080] FIG. 6 shows an exemplary block diagram 600 of an E2E NIC framework using iterative content-adaptive online training according to an embodiment.
[0081] As shown in FIG. 6, the E2E NIC framework includes a main encoder 610, a main decoder 620, a hyper encoder 630, a hyper decoder 640, and a context model 650. The E2E NIC framework may include one or more such modules. The E2E NIC framework further includes quantizers 660 / 661, arithmetic coders 670 / 671, and arithmetic decoders 680 / 681. The same or similar modules are denoted by the same reference numerals. The E2E NIC framework can include one or more modules not shown in FIG. 6.
[0082] The E2E NIC framework can use any DNN-based image compression method, such as a scale-hyperprior encoder / decoder framework (or Gaussian mixture likelihood framework) and its variants, an RNN-based recursive compression method and its variants, etc.
[0083] According to the embodiments of the present disclosure, the E2E NIC framework can utilize block diagram 600 as follows. When an input image or video sequence x is provided, the main encoder 610, when compared with the input image x, provides a compact compressed representation for storage and transmission purposes
Number
number
number
number
number
number
number
[0084] According to some embodiments, the E2E NIC framework can include a hyper prior model and a context model to further improve compression performance during the online training phase. The hyper prior model can be used to capture the spatial dependencies of the latent representations generated between layers of the neural network. According to some embodiments, side information can be used by the hyper prior model, and the side information is typically generated by motion-compensated temporal interpolation of adjacent reference frames on the decoder side. This side information can be used for training and inference of the E2E NIC framework. The hyper encoder 630 can use a hyperprior neural network-based encoder to encode the compression representation
Number
[0085] According to the embodiment, the E2E NIC framework is self-trained. The goal of the training process is to learn DNN encoding and DNN decoding (i.e., the main encoder 610 and the main decoder 620). In the training process, the weight coefficients of the DNN (i.e., the main encoder 610 and the main decoder 620) are first initialized, for example, by using a pre-trained corresponding DNN model or by setting them (the models) to random numbers. Next, when the input training image x is given, the input training image x undergoes the encoding process described in FIG. 4 to generate encoded information into a bitstream, which then undergoes the decoding process described in FIG. 6 to calculate and reconstruct the image
Number
Number
Number
[0086]
Number
[0087] Here, E measures the distortion of the decoded block residual compared to the original block residual before encoding, which serves as the regularization loss for the residual encoding / decoding DNN and the encoding / decoding DNN. β is a hyperparameter for balancing the importance of the regularization loss.
[0088] In some embodiments, the encoding DNN and the decoding DNN can be updated together based on the backpropagation gradients in the E2E framework.
[0089] FIG. 7 is a flowchart showing a method 700 for iterative E2E NIC content-adaptive online training using a neural network according to an embodiment.
[0090] In some embodiments, one or more processing blocks of FIG. 7 can be executed by the platform 220. In some embodiments, one or more processing blocks of FIG. 7 may be executed by a device or a group of devices separate from the platform 220, such as the user device 210, or including the platform 220.
[0091] As shown in FIG. 7, at operation 710, the method can include receiving an input image into the E2E NIC framework. In some embodiments, the method includes splitting the input image into one or more blocks.
[0092] At operation 720, the method 700 can include fine-tuning the E2E NIC framework based on the input image.
[0093] At operation 730, the method 700 can include calculating a parameter update using the first neural network of the fine-tuned E2E NIC framework. The parameter update can include a learning rate and a number of steps, which are selected based on the characteristics of the input image. The characteristics of the input image can be one of the RGB variance of the input image and the RD performance of the input image.
[0094] In operation 740, method 700 can include enhancing a fine-tuned E2E NIC framework based on a second neural network that is an extended network. The extended network can be a set of convolutional neural networks or layers of convolutional neural networks. Further, one or more extended networks can be used to enhance the fine-tuned E2E NIC framework.
[0095] In operation 750, method 700 can include generating an updated E2E NIC framework based on the enhanced E2E NIC framework and parameter updates.
[0096] In some embodiments, the method includes encoding an input image and parameter updates to generate a compressed representation of the input image and parameter updates, decoding the compressed representation of the parameter updates to generate a decoded parameter update, updating the E2E NIC based on the decoded parameter update, and decoding the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image.
[0097] In some embodiments, the method includes determining a distortion loss of the reconstructed image based on the consumption of the compressed representation of the input image and parameter updates, a trade-off hyperparameter, and the distortion of the decoded block residual of the reconstructed image.
[0098] FIG. 7 shows exemplary blocks of the method, but in some embodiments, the method can include additional blocks, fewer blocks, different blocks, or blocks arranged differently than those shown in FIG. 7. Additionally, or alternatively, two or more blocks of the method can be executed in parallel.
[0099] FIG. 8 is a block diagram of an example of computer code 800 for iterative content adaptive online training of an E2E NIC using a neural network according to an embodiment. According to an embodiment of the present disclosure, an apparatus / device including at least one processor having a memory storing computer code may be provided. The computer code may be configured to perform any number of aspects of the present disclosure when executed by the at least one processor.
[0100] As shown in FIG. 8, computer code 800 includes received code 810, fine-tuning code 820, calculation code 830, expansion code 840, and generation code 850.
[0101] Received code 810 is configured to cause at least one processor to receive an input image into the E2E NIC framework. The computer code 800 may further include code configured to cause at least one processor to divide / partition the input image into one or more blocks.
[0102] Fine-tuning code 820 is configured to cause at least one processor to fine-tune the E2E NIC framework based on the input image.
[0103] Calculation code 830 is configured to cause at least one processor to calculate parameter updates using a first neural network of the fine-tuned E2E NIC framework. The parameter updates may include a learning rate and a number of steps, which are selected based on characteristics of the input image. The characteristics of the input image may be one of the RGB variance of the input image and the RD performance of the input image.
[0104] The extension code 840 is configured to cause at least one processor to extend the fine-tuned E2E NIC framework based on a second neural network that is an extended network. The extended network can be a set of convolutional neural networks or layers of convolutional neural networks. Further, one or more extended networks can be used to extend the fine-tuned E2E NIC framework.
[0105] The generation code 850 is configured to cause at least one processor to generate an updated E2E NIC framework based on the extended E2E NIC framework and parameter updates.
[0106] The computer code 800 may further include code configured to cause at least one processor to encode an input image and parameter updates to generate a compressed representation of the input image and parameter updates, decode the compressed representation of the parameter updates to generate a decoded parameter update, update the E2E NIC framework based on the decoded parameter update, and decode the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image.
[0107] The computer code 800 may further include code configured to cause at least one processor to determine a distortion loss of the reconstructed image based on the consumption of the compressed representation of the input image and parameter updates, a trade-off hyperparameter, and the distortion of the decoded block residual of the reconstructed image.
[0108] FIG. 8 shows exemplary blocks of code, but in some embodiments, the device / apparatus may include additional blocks, fewer blocks, different blocks, or differently arranged blocks than those shown in FIG. 8. Additionally, or alternatively, two or more blocks of the device may be combined. In other words, although FIG. 8 shows separate blocks of code, various code instructions need not be separate and may be combined.
[0109] The methods and processes for iterative content adaptive online training of the E2E NIC framework described in this disclosure provide flexibility to the adaptive online training mechanism, improve NIC coding efficiency, and support various types of learning-based quantization methods, including DNN-based methods or conventional model-based methods. The described methods also provide a flexible and general framework that can accommodate different DNN architectures and multiple quality metrics.
[0110] The above-described techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media or realized by one or more specially configured hardware processors. For example, FIG. 2 shows an environment 200 suitable for implementing various embodiments. In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0111] As used herein, the term component is intended to be broadly construed as hardware, software, or a combination of hardware and software.
[0112] It will be apparent that the systems and / or methods described herein can be implemented in different forms of hardware, software, or a combination of hardware and software. The actual specific control hardware or software code used to implement these systems and / or methods does not limit the embodiments. Thus, although the operation and behavior of the systems and / or methods have been described herein without reference to specific software code, it is understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0113] Computer software can be coded using any suitable machine code or computer language and be subject to processing by mechanisms such as assembly, compilation, linking, or the like, to create code that includes instructions that can be executed directly by a computer central processing unit (CPU), a graphics processing unit (GPU), etc., or executed through the interpretation of microcode.
[0114] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.
[0115] Although several exemplary embodiments of the present disclosure have been described, there are changes, substitutions, and various alternative equivalents that are within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that, although not explicitly illustrated or described herein, many systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure can be envisioned.
Claims
1. A method for content - adaptive online training for end - to - end (E2E) neural image compression (NIC) using a neural network executed by at least one processor, the method comprising: receiving an input image into the E2E NIC framework; fine - tuning the E2E NIC framework based on the input image; calculating a parameter update using a first neural network of the fine - tuned E2E NIC framework; expanding the fine - tuned E2E NIC framework based on a second neural network which is the expanded network; generating an updated E2E NIC framework based on the expanded E2E NIC framework and the parameter update, wherein the parameter update includes a learning rate and a number of steps, and the learning rate and the number of steps are selected based on characteristics of the input image.
2. encoding the input image and the parameter update to generate a compressed representation of the input image and the parameter update; decoding the compressed representation of the parameter update to generate a decoded parameter update; updating the E2E NIC framework based on the decoded parameter update; further comprising decoding the compressed representation of the input image based on the updated E2E NIC framework to generate a reconstructed image, the method according to claim 1.
3. further comprising determining a distortion loss of the reconstructed image based on a consumption amount of the compressed representation of the input image and the parameter update, a trade - off hyper - parameter, and a distortion of a decoding block residual of the input image, the method according to claim 2.
4. The method according to claim 1, further comprising the step of dividing the input image into one or more blocks.
5. The method according to claim 1, wherein the characteristic of the input image is one of the RGB variance of the input image and the RD performance of the input image.
6. The extended network is a set of convolutional neural networks or layers of convolutional neural networks, and one or more extended networks are used to extend the fine-tuned E2E NIC framework. The method according to claim 1.
7. An apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, the apparatus comprising: At least one memory configured to store computer program code; At least one processor configured to read the computer program code and operate according to the instructions of the computer program code, When the instructions are executed by the at least one processor, the at least one processor is caused to: Execute the method according to any one of claims 1 to 6. Apparatus.
8. A non-transitory computer-readable medium storing instructions, wherein when the instructions are executed by at least one processor of an apparatus for content-adaptive online training for end-to-end (E2E) neural image compression (NIC) using a neural network, the at least one processor is caused to: Execute the method according to any one of claims 1 to 6. Non-transitory computer-readable medium.
Citation Information
Patent Citations
Substitutional end-to-end video coding
US20210360259A1
A method, an apparatus and a computer program product for image compression
WO2020008104A1
An apparatus, a method and a computer program for video coding and decoding
WO2020165493A1