End-to-end neural image compression methods, devices, computer equipment, and media
By replacing the neural image compression (NIC) method with an end-to-end (E2E) approach, the traditional codec is difficult to optimize as a whole by receiving the input image, determining the learning rate and the replacement image, and encoding the bitstream, thus achieving a more efficient image compression effect.
Patent Information
- Application Number
- CN202180020844.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-10-13
- Filing Date
- 2021-10-14
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2041-10-14
AI Technical Summary
Traditional hybrid video codecs are difficult to optimize as a whole, and improvements to individual modules cannot generate overall coding performance gains. The performance of existing neural image compression schemes urgently needs to be improved, especially in terms of rate-distortion performance.
An end-to-end (E2E) alternative neural image compression (NIC) method is adopted. By receiving the input image, the step size of the learning rate is determined, an alternative image is determined based on the trained model, and the image is encoded to generate a bitstream, which is then mapped to a compressed representation to optimize rate-distortion performance.
Joint optimization from input to output was achieved, improving the rate-distortion performance of neural image compression and enhancing coding efficiency and quality.
Smart Images

Figure CN115715463B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to image processing techniques, and more specifically to end-to-end neural image compression. Background Technology
[0002] Traditional hybrid video codecs are difficult to optimize holistically. Improvements to individual modules may not yield overall coding performance gains. Recently, standards groups and companies have been actively exploring potential needs for standardization in future video coding technologies. These groups have established the JPEG-AI group, focusing on AI-based end-to-end neural image compression using deep neural networks (DNNs). The Chinese AVS standard has also established the AVS-AI dedicated group to research neural image and video compression technologies. Recent successes have generated increasing industrial interest in advanced neural image and video compression methods.
[0003] The goal of neural image compression is to compute a compressed representation using images as input, which is compact for storage and transmission purposes. The performance of existing neural image compression schemes needs improvement; for example, there is a pressing need for neural image compression with improved rate-distortion performance. Summary of the Invention
[0004] According to an exemplary embodiment, a method for using an end-to-end (E2E) alternative neural image compression (NIC) with a neural network, executed by at least one processor, includes: receiving an input image into an E2E NIC framework; determining a step size of the input image that indicates the learning rate of a training model; determining an alternative image based on the training model; encoding the alternative image in place of the input image to generate a bitstream; and mapping the alternative image to the bitstream to generate a compressed representation.
[0005] According to an exemplary embodiment, an apparatus for using an end-to-end (E2E) neural network to replace neural image compression (NIC) includes: at least one memory configured to store program code; and at least one processor configured to read the program code and operate according to the instructions of the program code. The program code includes: receiving code configured to cause the at least one processor to receive an input image into an E2E NIC framework; step size determination code configured to cause the at least one processor to determine a step size of the input image indicating a learning rate for a training model; first determination code configured to cause the at least one processor to determine a replacement image based on the training model; first encoding code configured to cause the at least one processor to encode the replacement image replacing the input image to generate a bitstream; and mapping code configured to cause the at least one processor to map the replacement image to the bitstream to generate a compressed representation.
[0006] According to an exemplary embodiment, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor for end-to-end (E2E) replacement of neural image compression (NIC), cause at least one processor to: receive an input image into an E2E NIC framework; determine a step size of the input image indicating a learning rate for a training model; determine a replacement image based on the training model; encode the replacement image replacing the input image to generate a bitstream; and map the replacement image to the bitstream to generate a compressed representation.
[0007] The scheme provided in this disclosure may include receiving an image and determining an alternative representation of the image by performing an optimization process to optimize the rate-distortion performance of encoding the alternative representation of the image based on an end-to-end (E2E) optimization framework. The alternative representation of the image may be encoded to generate a bitstream. The scheme provided in this disclosure may perform joint optimization from input to output to improve the end goal (e.g., rate-distortion performance), thereby achieving end-to-end (E2E) optimized neural image compression (NIC). Attached Figure Description
[0008] Figure 1 The diagram shows an environment in which the methods, apparatus and systems described herein can be implemented according to an embodiment.
[0009] Figure 2 yes Figure 1 A block diagram of example components of one or more devices.
[0010] Figure 3 This is an example block diagram of a general alternative NIC framework according to the implementation method.
[0011] Figure 4This is an example diagram illustrating a substitutional learning-based image coding preprocessing model.
[0012] Figure 5 This is a flowchart of an end-to-end (E2E) alternative neural image compression (NIC) method according to an implementation.
[0013] Figure 6 This is a block diagram of an apparatus for an end-to-end (E2E) alternative to neural image compression (NIC) according to an embodiment. Detailed Implementation
[0014] Implementation may include receiving an image and determining alternative representations of the image by performing an optimization process to tune elements of the alternative representations and optimize rate-distortion performance of encoding these representations based on an end-to-end (E2E) optimization framework. The E2E optimization framework may be a pre-trained artificial neural network (ANN)-based image or video coding framework. Alternative representations of the image can be encoded to generate a bitstream. In an ANN-based video coding framework, different modules can jointly optimize from input to output by performing a machine learning process to improve the final objective (e.g., rate-distortion performance), thereby achieving end-to-end (E2E) optimized neural image compression (NIC).
[0015] Figure 1 This is a diagram of an environment 100 in which the methods, apparatus and systems described herein can be implemented according to an embodiment.
[0016] like Figure 1 As shown, environment 100 may include user equipment 110, platform 120, and network 130. The various devices in environment 100 may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections.
[0017] User equipment 110 includes one or more devices capable of receiving, generating, storing, processing, and / or providing information associated with platform 120. For example, user equipment 110 may include computing devices (e.g., desktop computers, laptop computers, tablet computers, handheld computers, smart speakers, servers, etc.), mobile phones (e.g., smartphones, cordless phones, etc.), wearable devices (e.g., smart glasses or smartwatches), or similar devices. In some implementations, user equipment 110 may receive information from platform 120 and / or send information to platform 120.
[0018] Platform 120 includes one or more devices as described elsewhere herein. In some implementations, platform 120 may include a cloud server or a group of cloud servers. In some implementations, platform 120 may be designed to be modular, allowing software components to be swapped in or out. This allows platform 120 to be easily and / or quickly reconfigured for different purposes.
[0019] In some implementations, as shown, platform 120 may be hosted in a cloud computing environment 122. It is worth noting that while the implementations described herein depict platform 120 as hosted in a cloud computing environment 122, in some implementations, platform 120 may not be cloud-based (i.e., it may be implemented outside of a cloud computing environment) or may be partially cloud-based.
[0020] The cloud computing environment 122 includes the environment of the hosting platform 120. The cloud computing environment 122 can provide services such as computing, software, data access, and storage, without requiring end users (e.g., user equipment 110) to know the physical location and configuration of the systems and / or devices of the hosting platform 120. As shown, the cloud computing environment 122 may include a set of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resources 124").
[0021] Computing resource 124 includes one or more personal computers, workstations, server devices, or other types of computing and / or communication devices. In some implementations, computing resource 124 may host platform 120. Cloud resources may include: computing instances executing in computing resource 124, storage devices provided in computing resource 124, data transmission devices provided by computing resource 124, etc. In some implementations, computing resource 124 may communicate with other computing resources 124 via wired connections, wireless connections, or a combination of wired and wireless connections.
[0022] For example, further Figure 1 As shown, computing resources 124 include a set of cloud resources, such as one or more applications (“Application, APP”) 124-1, one or more virtual machines (“Virtual Machine, VM”) 124-2, virtualized storage devices (“Virtualized Storage, VS”) 124-3, one or more hypervisors (“Hypervisor, HYP”) 124-4, etc.
[0023] Application 124-1 includes one or more software applications that can be provided to or accessed by user device 110 and / or platform 120. Application 124-1 can eliminate the need to install and execute software applications on user device 110. For example, application 124-1 may include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, an application 124-1 may send information to or receive information from one or more other applications 124-1 via virtual machine 124-2.
[0024] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 can be a system virtual machine or a process virtual machine, depending on the extent to which virtual machine 124-2 uses and corresponds to any real machine. A system virtual machine can be a complete system platform that supports the execution of a full operating system (OS). A process virtual machine can execute a single program and can support a single process. In some implementations, virtual machine 124-2 can execute on behalf of a user (e.g., user device 110) and can manage the infrastructure of the cloud computing environment 122, such as data management, synchronization, or long-duration data transfer.
[0025] Virtualized storage device 124-3 includes one or more storage systems and / or one or more devices that utilize virtualization technology within the storage system or device of computing resource 124. In some implementations, the type of virtualization within the context of the storage system may include block virtualization and file virtualization. Block virtualization may refer to the extraction (or separation) of logical storage relative to physical storage, enabling access to the storage system regardless of physical storage or heterogeneous architecture. Separation allows storage system administrators flexibility in how they manage storage for end users. File virtualization eliminates the dependency between data accessed at the file level and the location where the files are physically stored. This enables performance optimization for storage usage, server consolidation, and / or non-disruptive file migration.
[0026] Hypervisor 124-4 can provide hardware virtualization technology that allows multiple operating systems (e.g., "guest operating systems") to run simultaneously on a host computer such as computing resource 124. Hypervisor 124-4 can present a virtual operating platform to the guest operating system and manage the execution of the guest operating system. Multiple instances of various operating systems can share virtualized hardware resources.
[0027] Network 130 includes one or more wired and / or wireless networks. For example, network 130 may include cellular networks (e.g., fifth-generation (5G) networks, long-term evolution (LTE) networks, third-generation (3G) networks, code division multiple access (CDMA) networks, etc.), public land mobile networks (PLMN), local area networks (LAN), wide area networks (WAN), metropolitan area networks (MAN), telephone networks (e.g., public switched telephone networks (PSTN)), private networks, self-organizing networks, intranets, the Internet, fiber-optic networks, etc., and / or combinations of these or other types of networks.
[0028] Figure 1 The number and arrangement of devices and networks shown are provided as examples. In practice, with... Figure 1 Compared to the devices and / or networks shown, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown can be implemented as multiple distributed devices. Alternatively or alternatively, a group of devices in environment 100 (e.g., one or more devices) can perform one or more functions described as being performed by another group of devices in environment 100.
[0029] Figure 2 yes Figure 1 A block diagram of example components of one or more devices.
[0030] Device 200 may correspond to user device 110 and / or platform 120. For example... Figure 2 As shown, device 200 may include bus 210, processor 220, memory 230, storage unit 240, input unit 250, output unit 260 and communication interface 270.
[0031] Bus 210 includes components that allow communication between parts of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a Central Processing Unit (CPU), Graphics Processing Unit (GPU), Accelerated Processing Unit (APU), microprocessor, microcontroller, Digital Signal Processor (DSP), Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), or another type of processing unit. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes Random Access Memory (RAM), Read Only Memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 220.
[0032] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240 may include hard disks (e.g., magnetic disks, optical disks, magneto-optical disks, and / or solid-state drives), compact discs (CDs), digital versatile discs (DVDs), floppy disks, cassette tapes, magnetic tapes, and / or other types of non-transitory computer-readable media and corresponding drives.
[0033] Input component 250 includes components that allow device 200 to receive information, such as via user input (e.g., touchscreen display, keyboard, keypad, mouse, buttons, switches, and / or microphone). Alternatively or additionally, input component 250 may include sensors for sensing information (e.g., Global Positioning System (GPS) components, accelerometers, gyroscopes, and / or actuators). Output component 260 includes components that provide output information from device 200 (e.g., display, speaker, and / or one or more light-emitting diodes (LEDs)).
[0034] Communication interface 270 includes transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 270 may allow device 200 to receive information from and / or provide information to another device. For example, communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.
[0035] Device 200 can perform one or more of the processes described herein. Device 200 can perform these processes in response to processor 220 executing software instructions stored in non-transitory computer-readable media such as memory 230 and / or storage unit 240. Computer-readable media are defined herein as non-transitory memory devices. Memory devices include storage space within a single physical storage device or storage space distributed across multiple physical storage devices.
[0036] Software instructions can be read from another computer-readable medium or from another device into memory 230 and / or storage unit 240 via communication interface 270. When executed, the software instructions stored in memory 230 and / or storage unit 240 can cause processor 220 to perform one or more processes described herein. Alternatively or alternatively, hard-wired circuitry can be used in place of or in combination with software instructions to perform one or more processes described herein. Therefore, the implementations described herein are not limited to any particular combination of hardware circuitry and software.
[0037] Figure 2 The number and arrangement of the components shown are provided as an example. In practice, with... Figure 2 Compared to the components shown, device 200 may include additional components, fewer components, different components, or components arranged differently. Alternatively or additionally, a group of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another group of components of device 200.
[0038] Given an input image x, the goal of the NIC is to use image x as input to a DNN encoder to compute a compressed representation. This compression representation It is compact for storage and transmission purposes. Then, compression is used for representation. Used as input to the DNN decoder to reconstruct the image Some NIC methods may employ a variational autoencoder (VAE) architecture, where the DNN encoder directly uses the entire image x as its input, which is passed through a set of network layers that operate like a black box to compute the output representation (i.e., a compressed representation). Accordingly, the DNN decoder will compress the entire representation. As its input, the entire compressed representation The image is reconstructed using another set of network layers that act like another black box. Use the following target loss function To optimize the rate-distortion (RD) loss to obtain a reconstructed image with a trade-off hyperparameter λ. Distortion loss Compressed representation The trade-off between bit consumption R:
[0039]
[0040] One implementation related to preprocessing objectives proposes that, for each input image to be compressed, online training can be used to find the optimal alternative and compress that optimal alternative instead of the original image. Figure 3 This is an example block diagram of a general alternative NIC framework 300 according to an implementation method. For example... Figure 3 As shown, the general alternative NIC framework 300 includes an alternative module 310, an encoding module 320, and a decoding module 330.
[0041] The input image x is passed through the substitution module 310 to generate a substitute image with minimum target loss according to equation (1). By using this substitution, the encoding module 320 can achieve better compression performance. The compressed image can be decoded using the decoding module 330 to generate the reconstructed output z. t This method serves as a preprocessing step for improving the compression performance of any E2E NIC framework. It does not require any training or fine-tuning of the pre-trained compressed model itself or any training data. Detailed methods and apparatus for preprocessing models, according to one or more embodiments, will now be described.
[0042] Figure 4 This is an example diagram illustrating an alternative learning-based image coding preprocessing model.
[0043] Learning-based image compression can be viewed as a two-step mapping process. For example... Figure 4 As shown, the original image x0 in the high-dimensional space is mapped to a bitstream of length R(x0) (encoded mapping 400), which is then mapped back with a distortion loss of of The original space at that location (decoding map 410).
[0044] In the example implementation, such as Figure 4 As shown, if there exists an alternative image x′0 such that it is mapped to a bitstream of length R(x′0), this bitstream is then mapped to a bitstream with a distortion loss of... A space closer to the original image x0 Given a distance measurement or loss function, better compression can be achieved using an alternative image. According to Equation (1), optimal compression performance is achieved at the global minimum of the target loss function. In another example implementation, an alternative can be found at any intermediate step of the ANN to reduce the difference between the decoded image x1 and the original image x0.
[0045] Unlike the model training phase, where gradients are used to update model parameters, in the preprocessed model, the model parameters are fixed and the gradients can be used to update the input image itself. The entire model becomes differentiable by replacing non-differentiable parts with differentiable parts (e.g., replacing quantization with noise injection), allowing gradients to propagate back. Therefore, the above optimization can be solved iteratively using gradient descent.
[0046] This preprocessing model has two key hyperparameters: step size and number of steps. Step size indicates the "learning rate" of online training. Images with different content types may correspond to different step sizes to obtain optimal optimization results. Number of steps indicates the number of times the operation is updated. Hyperparameters are related to the target loss function. Both are used in the learning process. For example, the step size can be used for backpropagation calculations or gradient descent algorithms performed during the learning process. The number of iterations can be used as a threshold for the maximum number of iterations to control when the learning process can be terminated.
[0047] In the example implementation, during iterative online training, the learning rate (i.e., the step size) can be changed in each step by a scheduler. The scheduler determines the learning rate value, which can be increased or decreased. Alternatively, the learning rate can remain the same for one or more iterations of online training.
[0048] According to the implementation, a single scheduler or multiple different schedulers can be used to determine the learning rate for different input images. That is, multiple substitutions are generated based on multiple schedulers. The scheduler with the best compression performance is selected for each substitution. Furthermore, according to the implementation, the image can be compressed by dividing it into patches. For this purpose, multiple learning rate schedulers can be assigned to the image to obtain better compression results.
[0049] Figure 5This is a flowchart of an end-to-end (E2E) alternative neural image compression (NIC) method 500 according to an implementation method.
[0050] In some implementations, Figure 5 One or more process blocks can be executed by platform 120. In some implementations, Figure 5 One or more process blocks may be executed by another device or a group of devices (such as user equipment 110) that is separate from or includes platform 120.
[0051] like Figure 5 As shown, in operation 510, method 500 includes receiving an input image into an E2E NIC framework.
[0052] In operation 520, method 500 includes determining the step size of the input image that indicates the learning rate of the training model.
[0053] In operation 530, method 500 includes determining a substitute image based on a trained model. The substitute image can be determined through an optimization process of the trained model. This is done by adjusting the elements of the input image to generate a substitute representation and selecting the element with the minimum distortion loss between the input image and the substitute representation as the substitute image. Furthermore, the trained model can be trained based on a determined step size, the number of updates to the input image, and the distortion loss. For one or more iterations of the trained model, the step size can be increased, decreased, or kept the same. One or more substitute images can be generated based on different step sizes. The step size value corresponding to different step sizes can be determined based on a scheduler. A substitute image is selected based on a step size that results in better compression performance. Alternatively, the input image can be segmented into blocks, and each block is assigned a different scheduler.
[0054] An alternative image exists such that it maps to the input image, and the distance between the alternative image and the reconstructed image of the input image is shorter than the distance between the alternative image and the input image as measured by a distance metric or a loss function.
[0055] In operation 540, method 500 includes encoding a substitute image for the input image to generate a bitstream.
[0056] In operation 550, method 500 includes mapping an alternative image to a bitstream to generate a compressed representation. In implementations, one or more of the bitstream or compressed representation may be sent to, for example, a decoder and / or a receiving device.
[0057] When the input image is segmented into blocks, alternative blocks are determined for each block according to operations 530 to 550, and each alternative block is encoded and compressed.
[0058] Method 500 can use an artificial neural network based on a pre-trained image coding model, where the parameters of the artificial neural network are fixed and gradients are used to update the input image.
[0059] although Figure 5 An example block of this method is shown, but in some implementations, it differs from... Figure 5 Compared to the described method, this method may include additional blocks, fewer blocks, different blocks, or blocks arranged differently. Alternatively or concurrently, two or more blocks of this method may be executed in parallel.
[0060] Figure 6 This is a block diagram of an end-to-end (E2E) alternative to neural image compression (NIC) according to an embodiment.
[0061] like Figure 6 As shown, the device 600 includes a receiving code 610, a step size determination code 620, a determination code 630, an encoding code 640, and a mapping code 650.
[0062] Receive code 610 is configured to cause at least one processor to receive the input image into the E2E NIC framework.
[0063] Step size determination code 620 is configured to enable at least one processor to determine the step size of the input image that indicates the learning rate of the training model.
[0064] Code 630 is configured to enable at least one processor to determine alternative images based on a trained model.
[0065] Encoding code 640 is configured to cause at least one processor to encode a substitute image in place of the input image to generate a bitstream.
[0066] Mapping code 650 is configured to cause at least one processor to map an alternative image to a bitstream to generate a compressed representation.
[0067] The alternative image determined by determination code 630 can be determined through an optimization process of training the model. This is accomplished by adjusting and selecting codes, wherein the adjusting code is configured to cause at least one processor to adjust elements of the input image to generate an alternative representation, and the selection code is configured to cause at least one processor to select the element with the minimum distortion loss between the input image and the alternative representation as the alternative image. Furthermore, the training model can be trained based on the determined step size, the number of updates to the input image, and the distortion loss. For one or more iterations of training the model, the step size can be increased, decreased, or kept the same. One or more alternative images can be generated based on different step sizes. Step size values corresponding to multiple step sizes are determined based on multiple schedulers. Additionally, the input image can be segmented into blocks, where each block is assigned a different scheduler and encoded.
[0068] An alternative image exists such that it maps to the input image, and the distance between the alternative image and the reconstructed image of the input image is shorter than the distance between the alternative image and the input image as measured by a distance metric or a loss function.
[0069] In addition, device 600 can use an artificial neural network based on a pre-trained image coding model, wherein the parameters of the artificial neural network are fixed and gradients are used to update the input image.
[0070] although Figure 6 An example block of the device is shown, but in some implementations, it differs from... Figure 6 Compared to the depicted version, the device may include additional blocks, fewer blocks, different blocks, or blocks arranged differently. Additionally or alternatively, two or more blocks of the device may be combined.
[0071] This paper describes E2E image compression methods. These methods utilize alternative mechanisms to improve NIC coding efficiency by using a flexible and general framework that adapts to various types of quality metrics.
[0072] The E2E image compression method according to one or more embodiments can be used individually or in any order. Furthermore, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuitry system (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-transitory computer-readable medium.
[0073] The foregoing disclosure provides explanations and descriptions, but is not intended to be exhaustive or to limit the implementation to the precise form disclosed. Modifications and variations can be made based on the foregoing disclosure, or modifications and variations can be derived from practical implementation.
[0074] As used herein, the term “component” is intended to be interpreted broadly as hardware, firmware, or a combination of hardware and software.
[0075] It will be apparent that the systems and / or methods described herein can be implemented in various forms of hardware, firmware, or a combination of hardware and software. The actual dedicated control hardware or software code used to implement these systems and / or methods does not limit the implementation method. Therefore, this document describes the operation and behavior of the systems and / or methods without reference to specific software code—it should be understood that software and hardware can be designed to implement the systems and / or methods based on the description herein.
[0076] Even if combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of possible implementations. In fact, many of these features can be combined in ways not specifically recited in the claims and / or not disclosed in the specification. While each dependent claim listed below may directly refer to only one claim, the disclosure of possible implementations includes every dependent claim combined with every other claim in the group of claims.
[0077] Unless explicitly stated otherwise, no element, action, or instruction used herein should be construed as critical or necessary. Furthermore, as used herein, the articles “a” and “an” are intended to include one or more items and may be used interchangeably with “one or more.” Additionally, as used herein, the term “group” is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be used interchangeably with “one or more.” The term “an” or similar language is used where only one item is intended. Furthermore, as used herein, the terms “having,” “possessing,” “with,” etc., are intended to be open-ended terms. Additionally, unless explicitly stated otherwise, the phrase “based on” is intended to mean “at least partially based on.”
Claims
1. An end-to-end (E2E) neural image compression NIC method, characterized in that, The method includes: The input image is received by the E2E NIC framework; Determine the step size of the input image, whereby the step size indicates the learning rate of the training model; Based on the training model and the step size, a replacement image is determined to replace the input image; The alternative image is programmed to generate a bitstream; and the alternative image is mapped to the bitstream to generate a compressed representation; The training model is trained based on a determined step size, the number of updates to the input image, and a distortion loss; and for one or more iterations of the training model, the step size can be increased, decreased, or kept the same. Multiple alternative images are determined based on multiple step sizes; The step size value corresponding to the multiple step sizes is determined based on multiple schedulers, and an alternative image with the highest compression performance is selected for encoding.
2. The method according to claim 1, characterized in that, The alternative image is determined by performing an optimization process on the trained model, including: Adjusting the elements of the input image to generate an alternative representation; and The element with the least distortion loss between the input image and the alternative representation is selected as the alternative image.
3. The method according to claim 1, characterized in that, The alternative image is mapped to the input image; and Wherein, the distance between the alternative image and the reconstructed image of the input image is shorter than the distance between the alternative image and the input image as measured by distance measurement or loss function.
4. The method according to claim 1, characterized in that, The method further includes: The input image is segmented into one or more blocks. Each of the one or more blocks is assigned a scheduler from the plurality of schedulers.
5. The method according to claim 1, characterized in that, The training model is based on an artificial neural network with pre-trained image encoding, and The parameters of the artificial neural network are fixed, and gradients are used to update the input image.
6. An apparatus for end-to-end E2E neural image compression NIC, the apparatus comprising: The receiving module is configured to receive the input image into the E2E NIC framework; A step size determination module is configured to determine the step size of the input image, the step size indicating the learning rate of the training model, wherein the training model is trained based on the determined step size, the number of updates to the input image, and a distortion loss; and the step size can be increased, decreased, or kept the same for one or more iterations of the training model. The determination module is configured to determine a replacement image for the input image based on the training model and the step size; An encoding module is configured to encode the alternative image to generate a bitstream; and A mapping module is configured to map the alternative image to the bitstream to generate a compressed representation; The apparatus is further configured to determine multiple alternative images based on multiple step sizes; wherein multiple schedulers determine step size values corresponding to the multiple step sizes, and select the alternative image with the highest compression performance for encoding.
7. A computer device, comprising: At least one memory is configured to store program code; as well as At least one processor is configured to read the program code and operate according to the instructions of the program code to implement the method as described in any one of claims 1-5.
8. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform the method as described in any one of claims 1-5.