Quantization of weights in neural network-based compression scheme

By quantizing and encoding the maximum absolute value of weights in the neural network layer, the problem of insufficient image compression efficiency and reconstruction quality in the prior art is solved, and efficient image encoding and decoding is achieved.

CN120019651APending Publication Date: 2025-05-16INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070570.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-04
Filing Date
2023-09-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing image and video compression techniques are difficult to surpass traditional methods in terms of peak signal-to-noise ratios for single-image compression, and end-to-end trainable depth models have shortcomings in compression efficiency and reconstruction quality.

Method used

By obtaining the weights of the neural network, the maximum absolute value representing the weights in the neural network layer is quantized and encoded with the quantized weights into the bitstream to achieve encoding and decoding of the image.

Benefits of technology

This method can effectively encode and decode images, improve compression efficiency and reconstruction quality, and can surpass traditional methods in terms of peak signal-to-noise ratio of single-image compression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019651A_ABST
    Figure CN120019651A_ABST
Patent Text Reader

Abstract

An encoding method is disclosed. First, a weight of a neural network representing an input image is acquired. At least one value representative of a maximum absolute value of a weight in a layer of the neural network is then obtained. A weight of the layer is quantized in response to the at least one value. Finally, the at least one value and the quantized weight are encoded into a bitstream. These encoded weights may be provided to a decoder configured to reconstruct the image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of European application number 22306480.9, filed on October 4, 2022, which is incorporated herein by reference in its entirety. Technical Field

[0003] At least one of the present embodiments generally relates to a method and apparatus for encoding (respectively decoding) weights of a neural network, the weights representing an image. Background Art

[0004] Image and video compression is a fundamental task in image processing, which has become crucial in an era of massive popularity and increase in video streaming. Thanks to decades of great efforts by the community, traditional methods have reached state-of-the-art rate-distortion performance and dominate current industrial codec solutions. End-to-end trainable deep models have recently emerged as an alternative and achieved promising results. Now, they beat the best traditional compression methods (VVC, Versatile Video Coding) even in terms of peak signal-to-noise ratio for single image compression. Summary of the invention

[0005] In one embodiment, a coding method is disclosed, comprising:

[0006] Obtaining weights of a neural network, wherein the weights represent an input image;

[0007] obtaining at least one value representing a maximum absolute value of a weight in a layer of the neural network;

[0008] quantizing a weight of the layer in response to the at least one value; and

[0009] The at least one value and the quantized weight are encoded into a bitstream.

[0010] A coding device is disclosed, the coding device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the method just disclosed above.

[0011] In another embodiment, a decoding method is disclosed, comprising:

[0012] Obtaining a bitstream comprising at least one value representing a maximum absolute value of a weight in a layer of a neural network and a quantized weight of the layer;

[0013] decoding the at least one value and the quantized weights of the neural network from the bitstream;

[0014] inverse quantizing the quantized weights of the layer in response to the at least one value to obtain inverse quantized weights; and

[0015] The image is reconstructed using a neural network parameterized by the dequantized weights.

[0016] A decoding device is disclosed, the decoding device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the method disclosed above.

[0017] Additional embodiments are described herein that may be used alone or in combination.

[0018] One or more embodiments also provide a computer program, the computer program including instructions, which, when executed by one or more processors, cause the one or more processors to perform the method for encoding / decoding image or video data according to any embodiment described herein. One or more embodiments of the present embodiment also provide a non-transitory computer-readable medium and / or a computer-readable storage medium, on which instructions for encoding / decoding image or video data according to the method described herein are stored.

[0019] One or more embodiments also provide a computer-readable storage medium having stored thereon coded data generated according to the methods described herein, such as a bitstream. One or more embodiments also provide a method and apparatus for transmitting or receiving coded data (e.g., a bitstream) generated according to the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A block diagram showing a system in which aspects of the present embodiment may be implemented;

[0021] Figure 2 An example of an end-to-end neural network based compression system 200 for encoding an image using a deep neural network is shown;

[0022] Figure 3 An example of an end-to-end implicit neural network based compression system for encoding an image is shown;

[0023] Figure 4 An example of a flow chart of a method for encoding according to an embodiment is shown;

[0024] Figure 5 An example of an image decoder according to at least one embodiment is shown;

[0025] Figure 6shows an example of a flow chart of a method for decoding according to an embodiment;

[0026] Figure 7 A method for training aware quantized INR according to an embodiment is shown;

[0027] Figure 8 An example of a flow chart of a method for encoding according to an embodiment is shown;

[0028] Fig. 9 A model for entropy coding according to an embodiment is shown;

[0029] Fig.10 An example of an image decoder according to at least one embodiment is shown;

[0030] Fig.11 An example of a flowchart of a method for decoding according to an embodiment is shown; and

[0031] Figure 12 to Figure 14 Experimental results obtained using a method according to at least one embodiment are shown. DETAILED DESCRIPTION

[0032] The application describes various aspects, including tools, features, embodiments, models, methods, etc. Many aspects in these aspects are described as having specificity, and at least in order to show each feature, are usually described in a manner that may sound restrictive. However, this is for the purpose of describing clarity, and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide other aspects. In addition, these aspects can also be combined and interchanged with the aspects described in previous files.

[0033] The aspects described and contemplated in this application can be implemented in many different forms. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, devices, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the methods described, and / or computer-readable storage media having stored thereon a bitstream generated according to any of the methods described.

[0034] In this application, the terms "reconstruction" and "decoding" are used interchangeably, the terms "encoding" or "coding" are used interchangeably, the terms "pixel" and "sample" are used interchangeably, and the terms "image", "picture" and "frame" are used interchangeably. Usually, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0035] Various methods are described herein, and each method in these methods includes one or more steps or actions for realizing the method.Unless the correct operation of the method requires a specific order of steps or actions, the order and / or use of specific steps and / or actions can be modified or combined.Additionally, in various embodiments, terms such as "first", "second" can be used to modify elements, parts, steps, operations, etc., such as, for example, "first decoding" and "second decoding".Unless specifically required, the use of such terms does not imply that the modified operations are sorted.Therefore, in this example, the first decoding does not need to be performed before the second decoding, but can occur in a time period overlapping with the second decoding, such as before, during, or during the second decoding.

[0036] Figure 1 A block diagram of an example of a system in which various aspects and embodiments can be implemented is shown. System 100 can be embodied as a device including various components described below, and is configured to perform one or more aspects described in this application. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 100 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 100 are distributed over multiple ICs and / or discrete components. In various embodiments, system 100 is coupled to other systems or other electronic devices via, for example, a communication bus or by a dedicated input and / or output port communication. In various embodiments, system 100 is configured to implement one or more aspects described in this application.

[0037] The system 100 includes at least one processor 110 configured to execute instructions loaded therein to implement, for example, various aspects described in the present application. The processor 110 may include embedded memory, input-output interfaces, and various other circuits known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device and / or a non-volatile memory device). The system 100 includes a storage device 140, which may include non-volatile memory and / or volatile storage, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive, and / or optical disk drive. As a non-limiting example, the storage device 140 may include an internal storage device, an attached storage device, and / or a network-accessible storage device.

[0038] The system 100 includes an encoder / decoder module 130, which is configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. In addition, the encoder / decoder module 130 may be implemented as a separate element of the system 100, or may be incorporated into the processor 110 as a combination of hardware and software known to those skilled in the art.

[0039] Program code to be loaded onto the processor 110 or the encoder / decoder module 130 to perform various aspects described in this application may be stored in the storage device 140 and subsequently loaded onto the memory 120 for execution by the processor 110. According to various embodiments, one or more of the processor 110, the memory 120, the storage device 140, and the encoder / decoder module 130 may store one or more of various items during the execution of the processes described in this application. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.

[0040] In some embodiments, memory internal to the processor 110 and / or encoder / decoder module 130 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory can be memory 120 and / or storage device 140, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations.

[0041] As indicated at block 105, input may be provided to the elements of system 100 via various input devices. Such input devices include, but are not limited to: (i) an RF section that receives a radio frequency (RF) signal, such as transmitted over the air by a broadcaster, (ii) a component (COMP) input terminal (or a collection of COMP input terminals), (iii) a universal serial bus (USB) input terminal, and / or (iv) a high-definition multimedia interface (HDMI) input terminal. Figure 1 Other examples not shown include composite video.

[0042] In various embodiments, the input device of block 105 has associated corresponding input processing elements known in the art. For example, the RF section can be associated with elements suitable for the following operations: (i) selecting a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower frequency band to select a signal frequency band that (for example) can be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select a target data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a demultiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or a near-baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the target frequency band by filtering, down-conversion and again to perform frequency selection. Various embodiments rearrange the order of above-mentioned (and other) elements, remove some and / or add other elements of similar or different functions in these elements. Adding element can include inserting element between existing element, for example, inserting amplifier and analog-to-digital converter. In various embodiments, the RF part includes antenna.

[0043] In addition, the USB and / or HDMI terminals may include respective interface processors for connecting the system 100 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented as needed, for example, within a separate input processing IC or within the processor 110. Similarly, various aspects of USB or HDMI interface processing may be implemented as needed within a separate interface IC or within the processor 110. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 110 and an encoder / decoder module 130 operating in conjunction with memory and storage elements, to process the data stream as needed for presentation on an output device.

[0044] The various components of the system 100 may be disposed within an integrated housing. Within the integrated housing, the various components may be interconnected and transmit data therebetween using a suitable connection arrangement 115, such as an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.

[0045] The system 100 includes a communication interface 150 that can communicate with other devices via a communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to send and receive data through the communication channel 190. The communication interface 150 may include, but is not limited to, a modem or a network card, and the communication channel 190 may be implemented in, for example, a wired and / or wireless medium.

[0046] In various embodiments, a Wi-Fi network (such as IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers)) is used to stream data to the system 100. The Wi-Fi signal of these embodiments is received through a communication channel 190 and a communication interface 150 suitable for Wi-Fi communication. The communication channel 190 of these embodiments is usually connected to an access point or router, which provides access to an external network including the Internet to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to the system 100, and the set-top box transmits data through the HDMI connection of the input block 105. Similarly, other embodiments use the RF connection of the input block 105 to provide streaming data to the system 100. As indicated above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use a wireless network other than Wi-Fi, such as a cellular network or a Bluetooth network.

[0047] The system 100 can provide output signals to various output devices, including a display 165, a speaker 175, and other peripheral devices 185. The display 165 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 165 can be used for a television, a tablet computer, a laptop computer, a mobile phone, or other devices. The display 165 can also be integrated with other components (for example, as in a smartphone), or separate (for example, an external monitor of a laptop computer). In various examples of embodiments, other peripheral devices 185 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, for both terms), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 185 that provide functions based on the output of the system 100. For example, a disc player performs the function of playing the output of the system 100.

[0048] In various embodiments, control signals are transmitted between the system 100 and the display 165, speaker 175, or other peripheral device 185 using signaling such as AV link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output device can be communicatively coupled to the system 100 via a dedicated connection through the corresponding interfaces 160, 170, and 180. Alternatively, the output device can be connected to the system 100 using a communication channel 190 via the communication interface 150. The display 165 and the speaker 175 can be integrated into a single unit with other components of the system 100 in an electronic device (e.g., a television). In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (TCon) chip.

[0049] For example, if the RF portion of input 105 is part of a stand-alone set-top box, the display 165 and speaker 175 may alternatively be separate from one or more of the other components. In various embodiments in which the display 165 and speaker 175 are external components, the output signal may be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.

[0050] The embodiments may be performed by computer software implemented by the processor 110 or by hardware or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. As a non-limiting example, the memory 120 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as an optical memory device, a magnetic memory device, a semiconductor-based memory device, a fixed memory, and a removable memory. As a non-limiting example, the processor 110 may be of any type suitable for the technical environment and may cover one or more microprocessors, general-purpose computers, special-purpose computers, and processors based on multi-core architectures.

[0051] Figure 2 An example of an end-to-end neural network based compression system 200 for encoding an image using a deep neural network is shown. An input image I to be encoded is first processed by a deep neural network encoder 210 (hereinafter referred to as a deep encoder). The output y of the encoder is referred to as the embedding of the image. The embedding is encoded into a bitstream 220 by passing through a quantizer Q and then through an entropy encoder (e.g., an arithmetic encoder AE). The resulting bitstream 220 is decoded by passing through an entropy decoder (e.g., an arithmetic decoder AD) to reconstruct the quantized embedding The reconstructed quantized embedding can be processed by a deep neural network decoder 230 (hereinafter referred to as deep decoder or decoder) to obtain a decompressed image

[0052] Deep encoders and decoders consist of multiple neural layers such as convolutional layers. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called bias, and then applies a nonlinear function to the resulting value. The values ​​of the tensors and biases are denoted by the term "weights". The parameters of the weights and (if applicable) the nonlinear functions are called parameters of the network. In such compression systems, the encoder and decoder are fixed, based on a predetermined model that should be known at the time of encoding and decoding. The encoder and decoder neural networks are trained, for example, simultaneously so that they are compatible. In practice, in order to learn the weights of the encoder and decoder, the neural network is trained on a massive image database D. Together they are sometimes referred to as "autoencoders", which encode the input and then reconstruct it. The architecture of the decoder is usually the opposite of the encoder, although some layers or their order can be slightly different.

[0053] Figure 3 An example of an end-to-end implicit neural network (INR) based compression system 300 for encoding an image is shown. The system includes an encoder 312 and a decoder 332, which generates encoded data (e.g., in the form of a bitstream 320). The encoder 312 includes the INR 310, and the decoder 332 includes the INR 330. Compared to autoencoder-based image compression that uses latent points to control the rate-distortion target, in the INR-based compression system, the rate (R)-distortion (D) trade-off is controlled by the number of weights or the size of the neural network. Therefore, for different rates, the INR has different neural network architectures and different numbers of weights. As shown in FIG. Figure 3 As shown, the INR 310 or 330 maps pixel coordinates (x, y) to pixel values, such as (R, G, B) values ​​or other values, such as YCbCr, YUV or any other color values ​​of a given color space, for example, in the case of considering the RGB color space. θ (x, y) = (r, g, b), where f θ () is the INR function. INR is designed using a multilayer perceptron (MLP), where “L” is the number of layers and each layer contains the required number of hidden neurons. Each layer can be described as a function that first multiplies the input value by a tensor, adds a bias, and finally transforms the result through a nonlinear activation function. The values ​​of the tensor and bias are represented by the term “weights” and are denoted as θ. These weights are unknown and are to be estimated on the encoder side.

[0054] Using the INR function θTo compress the image I is equivalent to determining these weights for storage or transmission. To this end, the image I is first processed by an INR 310 responsible for determining the weights θ from the image I. The weights θ are encoded into a bit stream 320 by passing through a quantizer Q and then through an encoder ENC (for example, an entropy encoder such as an arithmetic encoder). The resulting bit stream is decoded by passing through a decoder DEC to reconstruct the quantized weights, which are dequantized by an inverse quantizer IQ (also called an inverse quantizer). The pixel coordinates of the image to be reconstructed are then input into the INR 330 parameterized by the dequantized weights to obtain the reconstructed image

[0055] In contrast to the autoencoder, the weights θ can be determined by learning the images to be encoded I. Therefore, each image to be encoded has its own associated weights. The weights θ can be determined by minimizing the following loss function:

[0056]

[0057] In equation (1), the sum of all pixels at coordinates (x, y) in an image of size M×N is calculated, and d is the measured reconstructed pixel value (also called the predicted pixel value, denoted as f θ (x, y)) and the actual pixel values ​​of image I (denoted as I(x, y)). Therefore, d can be any differentiable distortion measure, such as the mean squared error. Perceptual metrics such as LPIPS (Learning Perceptual Patch Similarity) can also be used. In this case, the loss is the mean squared error between the activations of the neural network. The weights θ can be determined by batch gradient descent or stochastic gradient descent. The nonlinear activation function used in INR plays a vital role in overfitting the high-frequency signals in the underlying image. The sinusoidal activation function can be used to capture high-frequency details and better overfit the image I.

[0058] For each image I, there is a specific INR function f θ Overfitting to a given image I. By f θ The quality of the reconstructed image depends on the size of the neural network. Since the weights serve as descriptors of the image, the larger the size of the neural network, the higher the bit length. On the other hand, limiting the number of weights will reduce the bit length at the expense of distortion.

[0059] Some existing methods for quantizing weights perform simple quantization of weights by quantizing 32-bit precision weights to 16-bit precision weights. Post-training quantization or raw quantization-aware training methods can also be used. However, the compression efficiency of these methods is not optimal because INR is unaware of the distortion from post-training quantization or quantization methods, and the entropy model is not efficient in existing quantization-aware training processes.

[0060] The embodiments described below aim to improve quantization and possibly entropy coding to increase compression efficiency, i.e. to reduce the file size of the weights with negligible or minimal loss in reconstruction quality. The principle can also be applied to the encoding / decoding of images (i.e. frames) of a video sequence. Furthermore, the decoding method disclosed below makes it possible to decode an image step by step, for example by first decoding a portion or a low-resolution image of the image, simply by evaluating the function f at each pixel position (e.g. one of two pixels) θ Parts of the image are difficult to decode using an autoencoder.

[0061] Figure 4 An example of a flow chart of a method for encoding an image I according to an embodiment is shown. The method may be Figure 3 The encoder 312 operates, and for example Figure 1 is implemented in the system 100. is a collection of tensor values, and is the aggregate bias of all layers with full precision, and θ = [θ w ,θ b ]. These weights can be obtained at step S100 by training the neural network with full precision (eg, 32-bit floating point weights) by minimizing the loss function of equation (1).

[0062] In step S110, the maximum absolute value is obtained in the current layer of index l in one type of weight, for example, in tensor value or in bias. The maximum absolute value is calculated for this type of weight (for example, for tensor value) as follows:

[0063]

[0064] In step S120, the weights in the current layer are quantized in response to the obtained maximum absolute value to obtain quantized weights. The quantized weights include dividing the weights by the obtained maximum absolute value to obtain normalized weights, as shown below:

[0065]

[0066] The normalized weights are then quantized using a fixed-bit quantizer. Let q be the fixed number of bits used to quantize the weights. In one example, for 8-bit quantization, q = 8, and The weights for regaining quantization are as follows:

[0067]

[0068] In other words, the quantized weights are obtained directly as follows:

[0069]

[0070] The fixed number of bits (q) may be the same for the entire data set. In a variation, the value of q may be selected based on any incoming image to be encoded, and may therefore vary from image to image. In this case, the value of q may be encoded into the bitstream for decoding at the decoder side.

[0071] In step S130, n bits (eg, n=16 bits) are used to obtain the maximum absolute value Encode and quantize the weights The encoding is in, for example, a bitstream 400, which can be stored on a storage medium or transmitted to another device, such as a decoder. The quantized weights can be written directly into the bitstream using q bits. In a variation, the quantized weights can be entropy encoded, for example, using an arithmetic encoder. Those skilled in the art will appreciate that the elements 410 (the encoded maximum absolute value) and 420 (the encoded quantized weights) in the bitstream 400 can be arranged in any order, or can even be interleaved in the bitstream.

[0072] In one example, the above steps S110 to S130 may be repeated for another layer. In one example, the above steps S110 to S130 are repeated for all remaining layers, and the fixed bit quantization weights of all layers are expressed as Encoding the maximum absolute value of the weights of all layers costs L×n bits in addition to the fixed-bit quantized weights.

[0073] In a similar manner, the above steps S110 to S130 may be repeated to quantize and encode another type of weight (e.g., bias) for a current layer or more than one layer (e.g., all layers). The quantized bias of layer 1 is represented as And the maximum absolute value is expressed as The deviation of fixed bit quantization of all layers is expressed as

[0074] In addition to the network weights, it takes 2×L×n bits to encode the maximum absolute value of the tensor and bias of the current layer. In this case,

[0075] In one example, only a subset of weights may be quantized, e.g., only biases, only tensor values, and / or only certain layers. In this case, only a subset of the largest absolute values ​​is thus signaled in the bitstream, e.g., only the

[0076] In one example, the above quantization may be performed on all quantized weights at once, rather than in an iterative process over the layers.

[0077] The above iterative process is not performed layer by layer, but can be performed on any subset of weights, such as weight by weight, neuron by neuron, neuron group by neuron, or any combination of these subsets, including, for example, quantizing some weights of some / all layers at each iteration.

[0078] Figure 5 An example of an image decoder 432 according to at least one embodiment is shown. The image decoder 432 is, for example, Figure 1 The system 100 is implemented in a system 100 and is suitable for decoding encoded data, such as data arranged as a bit stream 400, including an encoded maximum absolute value 410 and an encoded quantized weight 420. The encoded maximum absolute value 410 is decoded (dec) from the bit stream. The encoded quantized weight 420 is decoded (DEC) in response to the decoded maximum absolute value and inverse quantized (also called dequantized). The pixel coordinates of the image to be reconstructed are then input into the INR 430 parameterized by the dequantized weight to obtain a reconstructed image.

[0079] Figure 6 An example of a flowchart of a method for decoding according to an embodiment is shown. The method may be Figure 3 Decoder 332 or Figure 5 The decoder 432 operates, and for example in Figure 1 is implemented in the system 100.

[0080] In step S600, the decoder obtains coded data in the form of a bitstream, for example, received from another device or read from a storage medium. The coded data, for example, the bitstream 400, includes at least one maximum absolute value w max 410 and quantized weights For example Figure 5 shown.

[0081] In step S610, the quantized weights are decoded from the bitstream and at least one maximum value w max This step is the inverse step of step S130 on the encoder side. Therefore, in case the quantized weights are entropy encoded, they are entropy decoded at S610.

[0082] In step S620, the decoded quantized weights are inversely quantized. As an example, for the current layer l and the tensor value, the inversely quantized weights are obtained as follows: The same principle can apply to all layers and all types of weights, such as biases, or a subset of biases depending on what is being encoded.

[0083] In step S630, the pixel coordinates of the image to be reconstructed are then input into the INR 330 parameterized by the inverse quantization weights to obtain the reconstructed image

[0084] Figure 7 A method for training quantization-aware INR according to an embodiment is shown. The method may be used to obtain weights to be encoded at step S100.

[0085] Quantization-aware training can be done from the weights θ of the trained model with full precision. * In other words, at step S100-1, the initial weights, such as weight θ * In a variation, one can instead obtain a default random initialization of the weights.

[0086] In step S100-2, by applying Figure 4 Steps S110 to S120 of the method quantize these weights into quantized weights (Right now, Where x=w or x=b and is used as the initial value of the parameter.

[0087] In step S100-3, the quantized weights are dequantized as follows:

[0088]

[0089] The dequantized weights are expressed as

[0090] In step S100-4, the reconstruction loss is calculated as follows:

[0091]

[0092] The loss function is given by (called the prediction of the quantized model, i.e., from the weights using dequantization The distortion d() between the reconstructed image (i) and the original input image I) of the parameterized neural network INR is defined.

[0093] In step S100 - 5 , the weights are updated using a batch gradient descent method or a stochastic gradient descent method in response to the reconstruction loss.

[0094] These steps S100 - 2 to S100 - 5 are repeated until a stop criterion is reached ( S100 - 6 ). The stop criterion may be a convergence criterion (eg, loss<threshold) or reaching a certain number K of iterations, for example, K=10000.

[0095] In a first variant, based on predictions from a quantitative model The quantization-aware training of the loss function defined by the distortion with respect to the original input image I is modified to include a regularization term T with a hyperparameter λ. Therefore, during training, the following loss function is minimized instead of the loss of Eq. (2):

[0096]

[0097] The regularization term T can have various definitions.

[0098] In the first example, T is the prediction of the quantitative model Predictions with a fixed (throughout training) full precision model In other words, T is obtained by dequantizing the weights The parameterized neural network INR reconstructs the image from the full-precision weights θ * The distortion between the images reconstructed by the parameterized neural network INR. Therefore, during training, the following loss function is minimized:

[0099]

[0100] In the second example, T is the prediction of the quantized model at the current iteration Compared with the prediction of the unquantized model f θ In other words, T is the weight obtained by inverse quantization. The distortion between the image reconstructed from the parameterized neural network INR and the image reconstructed from the neural network INR parameterized with unquantized weights θ. Therefore, during training, the following loss function is minimized:

[0101]

[0102] Using the regularization term T during training has at least two advantages. First, it smooths the noise in the gradients introduced by quantization during the forward pass. In the neural network literature, the forward pass means the flow direction from "input" to "output". The backward pass means the flow direction from "output" to "input", after which the gradients are propagated backwards.

[0103] Secondly, in the quantitative model In the case where it cannot converge to the high-frequency components in the original image, at least it will try to converge to the predictions of a full-precision model that has fewer high-frequency components than the original image. Therefore, this regularization term helps optimization, especially optimization for higher quality.

[0104] In order to have faster encoding, it is sufficient to minimize the regularization term in equation (4), i.e. The hyperparameter λ can be selected once and used for the entire dataset, or it can be adjusted per specific image. During encoding, each hyperparameter in the set of hyperparameters can be trained in multiple devices, for example to use a faster encoding, rather than encoding on a single device. The weights corresponding to the lower loss are the quantized and encoded weights. Having image-specific hyperparameters λ leads to better performance.

[0105] During the backward pass, due to the non-differentiable nature of quantization, the gradients are calculated using a straight-through estimator (STE) and the weights are updated using an arbitrary optimizer. Finally, once determined, the weights (tensor values ​​and / or biases) are quantized to q bits and encoded. For this purpose, as in the aforementioned embodiment, steps S110 to S130 are applied to the weights obtained by the above-mentioned training method. Therefore, the maximum absolute value of each layer and each type of weight (tensor, bias, etc.) is determined. The weights obtained by the above-mentioned training method are quantized in response to the obtained maximum absolute value. The obtained maximum absolute value is encoded using n bits (e.g., n=16 bits), and the quantized weights are encoded, for example, by an entropy encoder. to encode.

[0106] Figure 8 An example of a flow chart of a method for encoding an image I according to another embodiment is shown. The method may be Figure 3 The encoder 312 operates, and for example Figure 1 is implemented in the system 100. Figure 4 The same steps as the encoding method described above are performed in Fig. 9 The same numerical references are used above. Specifically, the method includes steps S100, S110 and S120. As described with reference to the previous embodiment, at step S130, the quantized weights can be written directly into the bitstream using q bits. However, they can also be encoded using various methods, such as entropy coding methods, and more specifically arithmetic coding methods, to obtain additional compression efficiency. Entropy coding can utilize the weight distribution shape, and the q-bit quantized weights can be encoded using the q-bit quantized weights. Modeled to follow an explicit univariate probability distribution, i.e., fixed probabilities P for the boundary values border (For 8-bit quantization, it is -127 and +127, or more generally, for q-bit quantization, it is –(2 q-1 -1) and +(2 q-1 -1)), and the Gaussian distribution G of the remaining symbols, such as Fig. 9 As shown. In fact, in each layer, there is at least one symbol with the maximum absolute value (positive or negative). In the case of 8-bit quantization, this symbol can be -127 or +127, and their probabilities cannot be well fitted by any Gaussian distribution. A number, and At least L of the tensor values ​​and There are L deviations (they are quantized to -127 or +127 with equal probability), so the equal probability can be defined as follows The remaining symbols can follow a truncated Gaussian distribution with support [-126+126] and a total probability of The parameters of the Gaussian distribution can be calculated from the statistics of the symbols encoded with values ​​other than -127 or +127. Therefore, if the weight to be encoded is not -127 or +127, then Definition, then we can Estimate the mean of a Gaussian distribution and variance Therefore, in N(.;μ,σ 2 ) is a function with given parameters μ, σ 2 In the case of a Gaussian distribution, the probability of each symbol can be defined as follows.

[0107]

[0108] In step S132, the quantized weights are entropy encoded, for example, by an arithmetic encoder using the above-mentioned probability distribution (also called probability model), which is defined as a truncated Gaussian distribution (also called normal distribution) with parameters μ and σ 2 And further by fixing the probability boundary value P border to define.

[0109] The rate (expected bit length) can be calculated as follows:

[0110]

[0111] In this embodiment, at step S134, in addition to the maximum absolute value of the weight and the quantized weight In addition, for example, the mean μ and variance σ of the Gaussian distribution are each calculated using 16-bit floating point in a bitstream (such as bitstream 500). 2 Or standard deviation σ is encoded.

[0112] In another embodiment, different values ​​may be used. and To define the probability of the boundary value. In this case, L1 can be encoded in the bitstream.

[0113] The probabilities of the boundary values ​​can also include terms from a Gaussian distribution, defined as follows:

[0114]

[0115] The probability of the boundary value can be defined as a fixed value instead of

[0116] In case the data is adjusted for each picture, additional information may be included in the bitstream, eg, L1 or one or more bits signaling the selection made for each picture.

[0117] Fig.10 An example of an image decoder 532 according to at least one embodiment is shown. The decoder 532 is, for example, Figure 1 The system 100 is implemented in the embodiment of the present invention and is suitable for decoding encoded data, such as data arranged as a bit stream 500, including entropy model parameters 510 mean μ and standard deviation σ (or variance σ 2 ), the encoded maximum absolute value 515 and the encoded quantized weight 520. Decode (dec) the encoded maximum absolute value from the bitstream. Decode (D) the parameters of the entropy model. The encoded quantized weight 520 is entropy decoded by an entropy decoder AD, whose probability model is composed of parameters μ and σ 2 parameterized, and further by a fixed probability boundary value P border The decoded quantized weights are inversely quantized (also called dequantized) in response to the decoded maximum absolute value. The pixel coordinates of the image to be reconstructed are then input into the INR 530 parameterized by the dequantized weights to obtain the reconstructed image.

[0118] Fig.11 An example of a flowchart of a method for decoding according to an embodiment is shown. The method may be Figure 3 Decoder 332 or Fig.10 The 532 operation, and for example in Figure 1 is implemented in the system 100.

[0119] In step S900, the decoder obtains coded data in the form of a bit stream, for example, received from another device or read from a storage medium. Fig. 9 As depicted, the encoded data includes at least one maximum absolute value w of the probability model max , quantized weights Mean μ and standard deviation σ (or variance σ 2 ).

[0120] In step S910, the mean μ and standard deviation σ (or variance σ) are decoded from the bit stream. 2 ) and at least one maximum absolute value w max .

[0121] In step S920, the entropy-encoded quantized weights are entropy decoded using a probability model defined as a truncated Gaussian distribution, where the parameters of the probability model are μ and σ. 2 And further by fixing the probability boundary value P border Definition. This step is the inverse of the entropy encoding step.

[0122] In step S930, the decoded quantized weight is inversely quantized in response to the decoded maximum absolute value. As an example, for the current layer l and the tensor value, the inversely quantized weight is obtained as follows: The same principle can apply to all layers and all types of weights, such as biases, or a subset of biases depending on what is encoded.

[0123] In step S940, the pixel coordinates of the image to be reconstructed are then input into the INR 330 parameterized by the inverse quantization weights to obtain the reconstructed image

[0124] The figure below shows the experimental results obtained on the Kodak test set using the above method (using quantization, entropy coding and quantization-aware training). Fig.12 The rate distortion curves averaged over all images in the Kodak dataset are shown and show that the proposed method 600 has significant gains over competitors, called Coin 610 and Coin++ 620. To quantify the gain in %, the BD rate gain is calculated. Fig.13 shows that the average gain of the disclosed method 700 over the coin method 710 is 41.8%, and Fig.14 It is shown that the disclosed method 800 has an average gain of 31.5% over COIN++ 810. The regularization term T brings a gain of about 10% over 8-bit quantization using only entropy coding. In addition, the disclosed method is general and can be applied to any INR-based image / video codec.

[0125] Unless otherwise stated or technically excluded, the aspects described in this application may be used alone or in combination.

[0126] Various numerical values ​​are used in this application. The specific values ​​are for illustrative purposes, and the described aspects are not limited to these specific values.

[0127] Various implementations involve decoding. "Decoding" as used in this application may include, for example, all or part of a process performed on a received coded sequence to produce a final output suitable for display. In various embodiments, such a process includes one or more of the processes typically performed by a decoder, such as entropy decoding and inverse quantization. Based on the context of the specific description, whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refers to a broader decoding process will be clear, and it is believed that those skilled in the art will understand it well.

[0128] Various implementations involve encoding.In a manner similar to the discussion above regarding "decoding," "encoding" as used in this application may include, for example, all or part of the processing performed on an input video sequence to produce an encoded bitstream.

[0129] The implementations and aspects described herein may be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if only discussed in the context of a single implementation form (e.g., discussed only as a method), the implementation of the features discussed may be implemented in other forms (e.g., a device or program). The device may be implemented in, for example, appropriate hardware, software, and firmware. The method may be implemented in, for example, a device, such as a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0130] References to "one embodiment" or "an embodiment" or "an implementation" or "implementations" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this specification are not necessarily all referring to the same embodiment.

[0131] Additionally, this application or its claims may refer to "determining" various information. Determining information may include one or more of: for example, estimating the information, calculating the information, predicting the information, or retrieving the information from a memory.

[0132] Additionally, the application or its claims may refer to "accessing" various information. Accessing information may include one or more of: for example, receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.

[0133] Furthermore, the application or its claims may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of: for example, accessing information or retrieving information (e.g., from a memory or optical media storage device). Furthermore, "receiving" is often involved in one way or another during operations, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0134] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", any of the following uses of " / ", "and / or", and "at least one of" are intended to include selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). It will be readily appreciated by those of ordinary skill in this and related arts that this can be extended to as many listed items as possible.

[0135] In addition, as used herein, the word "signaling" refers in particular to indicating something to a corresponding decoder. For example, in some embodiments, the encoder signals a quantization matrix for inverse quantization, or at least one value representing the maximum absolute value of the weights in the layer of the neural network, quantizing the weights, mean and standard deviation of the Gaussian distribution. In this way, in an embodiment, the same parameters are used on both the encoder side and the decoder side. Thus, for example, the encoder can transmit (explicit signaling) specific parameters to the decoder so that the decoder can use the same specific parameters. On the contrary, if the decoder already has specific parameters as well as other parameters, signaling can be used without transmitting (implicit signaling) to simply allow the decoder to know and select specific parameters. By avoiding the transmission of any actual function, bit savings are achieved in various embodiments. It should be understood that signaling can be implemented in a variety of ways. For example, in various embodiments, one or more syntactic elements, flags, etc. are used to send information to the corresponding decoder. Although the verb form of the word "signaling" is mentioned above, the word "signal" can also be used as a noun in this article.

[0136] It will be apparent to those skilled in the art that implementations may generate a variety of signals formatted to carry information that may be stored or transmitted, for example. Information may include, for example, instructions for executing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is well known, a signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.

[0137] A number of embodiments have been described above. The features of these embodiments may be provided alone or in any combination across various claim categories and types.

[0138] A coding method is disclosed, the method comprising:

[0139] Obtaining weights of a neural network, wherein the weights represent an input image;

[0140] obtaining at least one value representing a maximum absolute value of a weight in a layer of the neural network;

[0141] quantizing a weight of the layer in response to the at least one value; and

[0142] The at least one value and the quantized weight are encoded into a bitstream.

[0143] In one example, quantizing the weight of the layer in response to the at least one value includes:

[0144] dividing the weight by the at least one value to obtain a normalized weight; and

[0145] The normalized weights are quantized using a fixed-bit quantizer.

[0146] In one example, encoding the quantized weights includes entropy encoding the quantized weights using a probability model defined by a fixed probability of boundary symbol values ​​and a Gaussian distribution of remaining symbols.

[0147] In one example, the mean and standard deviation of the Gaussian distribution are encoded into the bitstream.

[0148] In one example, obtaining weights of the neural network includes minimizing distortion between the input image and an image reconstructed from the neural network parameterized by the dequantized weights.

[0149] In one example, obtaining weights of a neural network includes minimizing a loss function that is a weighted sum between a first distortion and a second distortion, wherein the first distortion is a distortion between the input image and an image reconstructed from the neural network parameterized by dequantized weights, and the second distortion is a distortion between an image reconstructed from the neural network parameterized by fixed weights with full precision and an image reconstructed from the neural network parameterized by dequantized weights.

[0150] In one example, obtaining weights of a neural network includes minimizing a loss function that is a weighted sum between a first distortion and a second distortion, wherein the first distortion is a distortion between the input image and an image reconstructed from the neural network parameterized by inverse quantized weights, and the second distortion is a distortion between an image reconstructed from the neural network parameterized by non-quantized weights and an image reconstructed from the neural network parameterized by inverse quantized weights.

[0151] In one example, a weight belongs to a weight set, which includes a bias and a tensor value.

[0152] A decoding method is disclosed, the method comprising:

[0153] Obtaining a bitstream comprising at least one value representing a maximum absolute value of a weight in a layer of a neural network and a quantized weight of the layer;

[0154] decoding the at least one value and the quantized weights of the neural network from the bitstream;

[0155] inverse quantizing the quantized weights of the layer in response to the at least one value to obtain inverse quantized weights; and

[0156] The image is reconstructed using a neural network parameterized by the dequantized weights.

[0157] In one example, inverse quantizing the weights of the layer in response to the at least one value includes:

[0158] Inverse quantizing the quantized weights using a fixed bit quantizer; and

[0159] The inverse quantized weight is multiplied by the at least one value to obtain an inverse quantized weight.

[0160] In one example, decoding the quantized weights includes entropy decoding the quantized weights using a probability model defined by a fixed probability of boundary symbol values ​​and a Gaussian distribution of remaining symbols.

[0161] In one example, the mean and standard deviation of the Gaussian distribution are decoded from the bitstream.

[0162] In one example, the weight belongs to a weight set, which includes a bias and a tensor value.

[0163] A coding device is disclosed, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the coding method described in any of the previously disclosed examples.

[0164] A decoding device is disclosed, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to execute the decoding method described in any of the previously disclosed examples.

[0165] A computer program is disclosed, comprising program code instructions for implementing the encoding or decoding method when executed by a processor.

[0166] A computer-readable storage medium is disclosed, on which instructions for implementing the encoding or decoding method are stored.

Claims

1. A coding method, comprising: Obtaining weights of a neural network, the weights representing an input image; obtaining at least one value representing a maximum absolute value of a weight in a layer of the neural network; quantizing a weight of the layer in response to the at least one value; as well as The at least one value and the quantized weight are encoded into a bitstream.

2. The method of claim 1 , wherein quantizing the weight of the layer in response to the at least one value comprises: dividing the weight by the at least one value to obtain a normalized weight; as well as The normalized weights are quantized using a fixed-bit quantizer.

3. The method of claim 1 or 2, wherein encoding the quantized weights comprises entropy encoding the quantized weights using a probability model defined by a fixed probability of boundary symbol values ​​and a Gaussian distribution of remaining symbols.

4. The method of claim 3, wherein the mean and standard deviation of the Gaussian distribution are encoded into the bitstream.

5. The method of any one of claims 1 to 4, wherein obtaining the weights of the neural network comprises minimizing the distortion between the input image and an image reconstructed from the neural network parameterized by the inverse quantized weights.

6. The method of any one of claims 1 to 4, wherein obtaining the weights of the neural network comprises minimizing a loss function that is a weighted sum between a first distortion and a second distortion, wherein the first distortion is a distortion between the input image and an image reconstructed from the neural network parameterized by inverse quantized weights, and the second distortion is a distortion between an image reconstructed from the neural network parameterized by fixed weights with full precision and an image reconstructed from the neural network parameterized by inverse quantized weights.

7. The method of any one of claims 1 to 4, wherein obtaining the weights of the neural network comprises minimizing a loss function that is a weighted sum between a first distortion and a second distortion, wherein the first distortion is a distortion between the input image and an image reconstructed from the neural network parameterized by inverse quantized weights, and the second distortion is a distortion between an image reconstructed from the neural network parameterized by non-quantized weights and an image reconstructed from the neural network parameterized by inverse quantized weights.

8. The method of any one of claims 1 to 7, wherein the weights belong to a weight set, the weight set comprising a bias and a tensor value.

9. A decoding method, comprising: Obtaining a bitstream comprising at least one value representing a maximum absolute value of a weight in a layer of a neural network and a quantized weight of the layer; decoding the at least one value and the quantized weights of the neural network from the bitstream; inverse quantizing the quantized weights of the layer in response to the at least one value to obtain inverse quantized weights; as well as The image is reconstructed using a neural network parameterized by the dequantized weights.

10. The method of claim 9, wherein inverse quantizing the weights of the layer in response to the at least one value comprises: Inverse quantizing the quantized weights using a fixed bit quantizer; as well as The inverse quantized weight is multiplied by the at least one value to obtain an inverse quantized weight.

11. The method of claim 9 or 10, wherein decoding the quantized weights comprises entropy decoding the quantized weights using a probability model defined by a fixed probability of boundary symbol values ​​and a Gaussian distribution of remaining symbols.

12. The method of claim 11, wherein a mean and a standard deviation of the Gaussian distribution are decoded from the bitstream.

13. A method as claimed in any one of claims 8 to 12, wherein the weights belong to a weight set, the weight set comprising a bias and a tensor value.

14. An encoding device comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 1 to 8.

15. A decoding device, comprising one or more processors and at least one memory coupled to the one or more processors, wherein the one or more processors are configured to perform the method according to any one of claims 9 to 13.

16. A computer program comprising program code instructions for implementing the method according to any one of claims 1 to 13 when the program code instructions are executed by a processor.

17. A computer-readable storage medium having instructions stored thereon, the instructions being used to implement the method according to any one of claims 1 to 13 when executed by a processor.