Instance adaptive image and video compression using machine learning systems
By using machine learning systems and neural network training models, efficient image and video data compression bitstreams are generated, solving the problems of large data volume and artifacts in traditional technologies, and achieving efficient and low-burden data transmission and storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-30
- Publication Date
- 2026-04-07
AI Technical Summary
Existing image and video data compression technologies, while meeting high-quality requirements, overburden communication networks and equipment, and traditional encoding and decoding technologies may produce artifacts after decoding.
Machine learning systems, particularly neural networks, are used to generate bitstreams of compressed and decompressed image/video data by training and fine-tuning model parameters. Potential priors and model priors are used to reduce data volume and bit rate, and training is performed in conjunction with rate and distortion losses to optimize compression efficiency.
It achieves efficient compression of high-quality image and video data, reduces the burden on communication networks and equipment, and improves the performance and efficiency of compression and decompression, while avoiding artifact problems in traditional technologies.
Smart Images

Figure CN116250236B_ABST
Abstract
Description
Technical Field
[0001] In general, this disclosure relates to data compression, and more specifically, to the use of machine learning systems to compress image and / or video content. Background Technology
[0002] Many devices and systems allow image / video data to be processed and output for consumption. Digital image / video data comprises vast amounts of data to meet ever-increasing demands for image / video quality, performance, and features. For example, consumers of video data typically expect high-quality video with high fidelity, high resolution, and high frame rates. The large amounts of video data often required to meet these demands place a heavy burden on communication networks and devices that process and store video data. Video encoding and decoding technologies can be used to compress video data. An example goal of video encoding and decoding is to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. As evolving video services become available and the demand for large amounts of video data continues to increase, there is a need for encoding and decoding technologies with better performance and efficiency. Summary of the Invention
[0003] In some examples, systems and techniques for using one or more machine learning systems to compress and / or decompress data are described. In some examples, machine learning systems for compressing and / or decompressing image / video data are provided. According to at least one illustrative example, a method for compressing and / or decompressing image / video data is provided. In some examples, the method may include: receiving input data via a neural network compression system for compression via the neural network compression system; determining a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generating a first bitstream comprising a compressed version of the input data via the neural network compression system using a latent prior; generating a second bitstream comprising a compressed version of the updated model parameters via the neural network compression system using the latent prior and the model prior; and outputting the first bitstream and the second bitstream for transmission to a receiver.
[0004] According to at least one illustrative example, a non-transitory computer-readable medium is provided for compressing and / or decompressing image / video data. In some aspects, the non-transitory computer-readable medium may include instructions that, when executed by one or more processors, cause the one or more processors to: receive input data via a neural network compression system for compression via the neural network compression system; determine a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generate a first bitstream comprising a compressed version of the input data via the neural network compression system using a latent prior; generate a second bitstream comprising a compressed version of the updated model parameters via the neural network compression system using the latent prior and the model prior; and output the first bitstream and the second bitstream for transmission to a receiver.
[0005] According to at least one illustrative example, an apparatus for compressing and / or decompressing image / video data is provided. In some aspects, the apparatus may include a memory having computer-readable instructions stored thereon and one or more processors configured to: receive input data via a neural network compression system for compression via the neural network compression system; determine a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generate a first bitstream comprising a compressed version of the input data via the neural network compression system using latent priors; generate a second bitstream comprising a compressed version of the updated model parameters via the neural network compression system using the latent priors and model priors; and output the first bitstream and the second bitstream for transmission to a receiver.
[0006] According to another illustrative example, an apparatus for compressing and / or decompressing image / video data may include units for performing the following operations: receiving input data through a neural network compression system for compression by the neural network compression system; determining a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generating a first bitstream comprising a compressed version of the input data using the neural network compression system with latent priors; generating a second bitstream comprising a compressed version of the updated model parameters using the neural network compression system with the latent priors and model priors; and outputting the first bitstream and the second bitstream for transmission to a receiver.
[0007] In some aspects, the methods, apparatus, and computer-readable media described above can generate a concatenated bitstream comprising the first bitstream and the second bitstream; and transmit the concatenated bitstream to the receiver.
[0008] In some examples, the second bitstream also includes a compressed version of the potential prior and a compressed version of the model prior.
[0009] In some cases, generating the second bitstream may include: entropy encoding the latent prior using the model prior via the neural network compression system; and entropy encoding the updated model parameters using the model prior via the neural network compression system.
[0010] In some examples, the updated model parameters include one or more updated parameters of the decoder model. In some cases, the one or more updated parameters may be tuned using the input data.
[0011] In some examples, the updated model parameters include one or more updated parameters of the encoder model. In some cases, the one or more updated parameters may be tuned using the input data. In some cases, the first bitstream is generated by the neural network compression system using the one or more updated parameters.
[0012] In some examples, generating the second bitstream may include: encoding the input data into a latent spatial representation of the input data using the neural network compression system with one or more updated parameters; and entropy encoding the latent spatial representation into the first bitstream using the latent priors through the neural network compression system.
[0013] In some aspects, the methods, apparatus, and computer-readable media described above can generate model parameters of the neural network compression system based on a training dataset used to train the neural network compression system; use the input data to tune the model parameters of the neural network compression system; and determine the set of updates based on the differences between the model parameters and the tuned model parameters.
[0014] In some examples, the model parameters are tuned based on the input data, the bit size of the compressed version of the input data, the set of updated bit sizes, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0015] In some examples, the model parameters are tuned based on the input data and the ratio of the cost of sending the set of updates to the distortion between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
[0016] In some examples, tuning the model parameters may include determining to include one or more parameters in the tuned model parameters based on the following: including the one or more parameters in the tuned model parameters is accompanied by a reduction in at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0017] In some examples, determining the set of updates for the neural network compression system may include: processing the input data at the neural network compression system; determining one or more losses for the neural network compression system based on the processed input data; and tuning model parameters of the neural network compression system based on the one or more losses, the tuned model parameters including the set of updates for the neural network compression system.
[0018] In some cases, the one or more losses include: rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream, distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
[0019] In some examples, the receiver includes an encoder. In some aspects, the methods, apparatus, and computer-readable media described above can receive data including a first bitstream and a second bitstream via the encoder; decode the compressed version of the updated model parameters based on the second bitstream via the decoder; and generate a reconstructed version of the input data based on the compressed version of the input data in the first bitstream using the set of updated parameters via the decoder.
[0020] In some respects, the methods, apparatus, and computer-readable media described above can train the neural network compression system by reducing rate distortion and model rate loss, wherein the model rate reflects the length of the bit stream used to send model updates.
[0021] In some examples, the model priors include independent Gaussian network priors, independent Laplacian network priors, and / or independent Spike and Slab network priors.
[0022] In some aspects, the device may be or be a subset of the following: a camera (e.g., an IP camera), a mobile device (e.g., a mobile phone or so-called "smartphone" or other type of mobile device), a smart wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, a 3D scanner, a multi-camera system, or other devices. In some aspects, the device includes a camera or multiple cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the aforementioned device may include one or more sensors.
[0023] This invention is not intended to identify key or essential features of the claimed inventive subject matter, nor is it intended to be used alone to define the scope of the claimed inventive subject matter. The inventive subject matter should be understood by referring to the appropriate portions of the entire specification, any or all of the drawings, and each claim.
[0024] The foregoing and other features and embodiments will become more apparent from the following description, claims and drawings. Attached Figure Description
[0025] The illustrative embodiments of this application are described in detail below with reference to the following accompanying drawings, wherein:
[0026] Figure 1 This is a diagram illustrating an example image processing system according to some examples of this disclosure;
[0027] Figure 2A This is a diagram illustrating examples of fully connected neural networks according to some examples of this disclosure;
[0028] Figure 2B This is a diagram illustrating an example of a locally connected neural network according to some examples of this disclosure;
[0029] Figure 2C This is a diagram illustrating examples of convolutional neural networks according to some examples of this disclosure;
[0030] Figure 2D This is a diagram illustrating some examples of a deep convolutional network (DCN) for recognizing visual features from an image, according to this disclosure;
[0031] Figure 3 This is a block diagram illustrating an example deep convolutional network (DCN) according to some examples of this disclosure;
[0032] Figure 4 This is a diagram illustrating examples of systems according to this disclosure, including a transmitting device for compressing video content and a receiving device for decompressing a received bitstream into video content;
[0033] Figure 5A and Figure 5B This is a diagram illustrating some examples of an example rate distortion automatic encoder system according to this disclosure;
[0034] Figure 6 This is a diagram illustrating an example neural network compression system for instance adaptive data compression, based on some examples of this disclosure;
[0035] Figure 7 This is a diagram illustrating an example architecture of a neural network compression system that uses model prior fine-tuning (e.g., instance adaptation) according to some examples of this disclosure;
[0036] Figure 8 This is a diagram illustrating an example inference process implemented by an example neural network compression system using model prior fine-tuning, according to some examples of this disclosure;
[0037] Figure 9 This is a diagram illustrating, according to some examples, the encoding and decoding tasks performed by an example neural network compression system fine-tuned using model priors, based on the present disclosure;
[0038] Figure 10 Examples of rate distortion autoencoder models, fine-tuned at data points to be sent to a receiver and without fine-tuning at data points to be sent to a receiver, are shown in accordance with this disclosure.
[0039] Figure 11 This is a flowchart illustrating an example process 1100 for adaptive compression using an instance of a neural network compression system adapted to the compressed input data (e.g., fine-tuning for the compressed input data), according to some examples of the present disclosure.
[0040] Figure 12 This is a flowchart illustrating an example of a process for compressing one or more images, according to some examples of this disclosure;
[0041] Figure 13 This is a flowchart illustrating examples of a process for decompressing one or more images, according to some examples of this disclosure; and
[0042] Figure 14 An example computing system is shown, which is based on some examples of this disclosure. Detailed Implementation
[0043] Certain aspects and embodiments of this disclosure are provided below. It will be apparent to those skilled in the art that some of these aspects and embodiments can be applied independently, and some can be applied in combination. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be implemented without using these specific details. The drawings and description are not intended to be limiting.
[0044] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the subsequent description of exemplary embodiments will provide those skilled in the art with the ability to implement the descriptions for carrying out the exemplary embodiments. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0045] As mentioned above, digital image and video data can encompass large amounts of data, especially with the continued growth in demand for high-quality video data. For example, consumers of image and video data typically desire increasingly higher video quality, such as high fidelity, resolution, and frame rate. However, the large amounts of data required to meet such demands can place a heavy burden on communication networks (e.g., high bandwidth and network resource requirements) and on the devices that process and store video data. Therefore, compression algorithms (also known as encoding / decoding algorithms or tools) are advantageous for reducing the amount of data required to store and / or transmit image and video data.
[0046] Various techniques can be used to compress image and video data. Image data compression is accomplished using algorithms such as the Joint Picture Experts Group (JPEG) and Better Portable Graphics (BPG). In recent years, neural network-based compression methods have shown great promise in compressing image data. Video encoding and decoding can be performed according to specific video codec standards. Example video codec standards include High Efficiency Video Codec (HEVC), Basic Video Codec (EVC), Advanced Video Codec (AVC), Moving Picture Experts Group (MPEG) codec, and Universal Video Codec (VVC). However, such traditional image and video codec techniques can introduce artifacts in the reconstructed image after decoding.
[0047] In some respects, this document describes systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for performing data compression and decompression (also known as encoding and decoding, collectively referred to as encoding and decoding) using one or more machine learning systems. One or more machine learning systems can be trained as described herein and used to perform data compression and / or decompression, such as image, video, and / or audio compression and decompression. The machine learning systems described herein can perform training and compression / decompression techniques that produce high-quality data output.
[0048] The systems and techniques described herein can perform compression and / or decompression of any type of data. For example, in some cases, the systems and techniques described herein can perform compression and / or decompression of image data. As another example, in some cases, the systems and techniques described herein can perform compression and / or decompression of video data. As another example, in some cases, the systems and techniques described herein can perform compression and / or decompression of audio data. For purposes of simplicity, illustration, and explanation, the systems and techniques described herein are discussed with reference to the compression and / or decompression of image data (e.g., images, videos, etc.). However, as stated above, the concepts described herein can also be applied to other modalities, such as audio data and any other type of data.
[0049] In some examples, machine learning systems used for data compression and / or decompression can be trained on training datasets (e.g., images, videos, audio, etc.) and can be further fine-tuned (e.g., trained, fitted) for data that will be sent to and decoded by a receiver. In some cases, the encoder of the machine learning system can send updated parameters of a compressed model, fine-tuned using data that will be sent to and decoded by the decoder, to the decoder. In some examples, the encoder can send updated model parameters (and / or, instead of sending the complete set of model parameters) without other model parameters to reduce the amount of data and / or bit rate sent to the decoder. In some cases, model priors can be used to quantize and compress the updated model parameters to reduce the amount of data and / or bit rate sent to the decoder.
[0050] The compression model used by the encoder and / or decoder can be generalized to different types of data. Furthermore, by fine-tuning the compression model for the data being transmitted and decoded, machine learning systems can improve the compression and / or decompression performance, quality, and / or efficiency for that specific data. In some cases, the model of a machine learning system can be trained using rate and distortion losses, as well as an additional rate loss that reflects and / or explains the additional overhead and bit rate required to transmit updated model parameters. The model can be trained to minimize rate (e.g., the size / length of the bitstream), distortion (e.g., the distortion between the input and the reconstructed output), and model rate loss (e.g., the loss reflecting the cost of transmitting updated model parameters). In some examples, the machine learning system may make trade-offs between rate, distortion, and model rate (e.g., the size / length of the bitstream required to transmit updated model parameters).
[0051] In some examples, a machine learning system may include one or more neural networks. Machine learning (ML) is a subset of artificial intelligence (AI). A machine learning system includes algorithms and statistical models that a computer system can use to perform various tasks through dependency patterns and inferences without explicit instructions. An example of an ML system is a neural network (also known as an artificial neural network), which can consist of groups of interconnected artificial neurons (e.g., neuron models). Neural networks can be used in a variety of applications and / or devices, such as image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and more.
[0052] Individual nodes in a neural network can mimic biological neurons by taking input data and performing simple operations on that data. The results of these simple operations on the input data are selectively passed to other neurons. Weights are associated with each vector and node in the network, and these values constrain how the input data relates to the output data. For example, the input data for each node can be multiplied by a corresponding weight, and the products can be summed. The sum of the products can be tuned with optional biases, and activation functions can be applied to the result, producing the node's output signal or "output activation" (sometimes called an activation map or feature map). The weights can initially be determined by an iterative stream of training data through the network (e.g., weights are established during training phases where the network learns how to identify specific classes based on the characteristics of its typical input data).
[0053] There are different types of neural networks, such as deep generative neural network models (e.g., generative adversarial networks (GANs)), recurrent neural network (RNN) models, multilayer perceptron (MLP) neural network models, convolutional neural network (CNN) models, autoencoders (AEs), etc. For example, a GAN is a generative neural network that learns patterns in input data so that the neural network model can generate new synthetic outputs that reasonably likely come from the original dataset. A GAN may include two neural networks operating together. One neural network (called the generative neural network or generator, denoted as G(z)) generates the synthetic output, while the other neural network (called the discriminative neural network or discriminator, denoted as D(X)) evaluates the output for realism (whether the output comes from the original dataset, such as the training dataset, or was generated by the generator). Training inputs and outputs may include images as illustrative examples. The generator is trained to try and fool the discriminator into determining that the synthetic images generated by the generator are real images from the dataset. The training process continues, and the generator becomes better at generating synthetic images that look like real images. The discriminator continues to look for defects in the synthesized images, and the generator identifies what the discriminator is looking at to determine the defects in the image. Once the network is trained, the generator is able to produce realistic images that the discriminator cannot distinguish from real images.
[0054] RNNs work on the principle of storing the output of a layer and feeding that output back to the input to help predict the outcome of that layer. In MLP neural networks, data can be fed into the input layer, and one or more hidden layers provide a level of abstraction for the data. Predictions can then be made on the output layer based on this abstract data. MLPs are particularly well-suited for classification prediction problems where the input is assigned a category or label. Convolutional Neural Networks (CNNs) are a type of feedforward artificial neural network. A CNN can consist of an ensemble of artificial neurons, each with a corresponding receptive field (e.g., a localized region of the input space) that together tile the input space. CNNs have many applications, including pattern recognition and classification.
[0055] In a hierarchical neural network architecture (called a deep neural network when there are multiple hidden layers), the output of the first layer's artificial neurons becomes the input of the second layer's artificial neurons, the output of the second layer's artificial neurons becomes the input of the third layer's artificial neurons, and so on. Convolutional neural networks can be trained to recognize hierarchical structures of features. The computation in a convolutional neural network architecture can be distributed across a set of processing nodes, which can be configured in one or more chains of operations. These multi-layered architectures can be trained layer by layer and can be fine-tuned using backpropagation.
[0056] Autoencoders (AEs) can learn effective data encoding and decoding in an unsupervised manner. In some examples, an AE can learn a representation of a set of data (e.g., data encoding and decoding) by training a network to ignore signal noise. An AE can include an encoder and a decoder. The encoder maps input data to code, while the decoder maps code to a reconstruction of the input data. In some examples, a rate-distortion autoencoder (RD-AE) can be trained to minimize the average rate-distortion loss over a dataset of data points (e.g., image and / or video data points). In some cases, an RD-AE can perform forward passes during speculation to encode new data points.
[0057] In some cases, the RD-AE can be fine-tuned on the data to be sent to a receiver (e.g., a decoder). In some examples, by fine-tuning the RD-AE at data points, high compression (e.g., rate / distortion) performance can be achieved. The encoder associated with the RD-AE can send the RD-AE model or a portion of the RD-AE model to the receiver (e.g., a decoder) for the receiver to decode the bitstream comprising the compressed data sent by the encoder.
[0058] In some cases, the AE model may be large, which can increase the bit rate and / or reduce the rate distortion gain. In some examples, the RD-AE model can be fine-tuned using a model prior (e.g., an RDM-AE prior). A model prior can be defined and used to generate model updates for a receiver (e.g., a decoder) to implement the model for decompressing transmitted data. A model prior can reduce the amount of data sent to the decoder by generating model updates under that prior. In some examples, a model prior can be designed to reduce the cost of sending model updates. For example, a model prior can be used to reduce and / or limit the bit rate overhead of model updates and / or produce smaller model updates. In some cases, the bit rate overhead of model updates may increase as more parameters are fine-tuned. In some examples, the number of parameters being fine-tuned can be reduced, which can also reduce or limit the bit rate overhead.
[0059] In some cases, the RDM-AE loss can be used to fine-tune the model. A loss term can be added to the bit rate used for model updates. The added loss term can compensate for the bits "spent" on model updates. For example, any bits "spent" on model updates during fine-tuning can be compensated for by improvements in rate distortion (R / D). In some examples, any bits "spent" on model updates during fine-tuning can be compensated for by at least as much improvement in rate distortion.
[0060] In some cases, the design of the model prior can be improved as further described herein. In some illustrative examples, the model prior design may include an independent Gaussian model prior. In other illustrative examples, the model prior design may include an independent Laplace model prior. In other illustrative examples, the model prior design may include independent Spike and Slab priors. In some illustrative examples, the model prior and the global AE model can be trained jointly. In some illustrative examples, the model prior may include complex dependencies learned by a neural network.
[0061] Figure 1 This is a diagram illustrating an example of an image processing system 100 according to some examples of the present disclosure. In some cases, the image processing system 100 may include a central processing unit (CPU) 102 or a multi-core CPU configured to perform one or more functions described herein. Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., a weighted neural network), latency, frequency bin information, task information, and other information may be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from program memory associated with the CPU 102 and / or memory block 118.
[0062] The image processing system 100 may include additional processing blocks tailored for specific functions, such as a GPU 104; a DSP 106; a connectivity block 110, which may include fifth-generation (5G) connectivity, fourth-generation Long Term Evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.; and / or a multimedia processor 112 that can, for example, detect and recognize features. In one embodiment, an NPU 108 is implemented within the CPU 102, DSP 106, and / or GPU 104. The image processing system 100 may also include a sensor processor 114, one or more image signal processors (ISPs) 116, and / or a storage device 120. In some examples, the image processing system 100 may be based on the ARM instruction set.
[0063] Image processing system 100 may be part of a computing device or a plurality of computing devices. In some examples, image processing system 100 may be part of an electronic device (or apparatus), such as a camera system (e.g., a digital camera, IP camera, camcorder, security camera, etc.), a telephone system (e.g., a smartphone, cellular phone, conferencing system, etc.), a desktop computer, an XR device (e.g., a head-mounted display, etc.), a smart wearable device (e.g., a smartwatch, smart glasses, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a digital media player, a game console, a video streaming device, a drone, an in-vehicle computer, a system-on-a-chip (SoC), an Internet of Things (IoT) device, or any other suitable electronic device.
[0064] Although the image processing system 100 is shown to include certain components, those skilled in the art will understand that the image processing system 100 may include more than Figure 1 The components shown may include more or fewer components. For example, in some cases, the image processing system 100 may also include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more network interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or Figure 1 Other hardware or processing devices not shown in the diagram. The following refers to… Figure 14 Illustrative examples of computing devices and hardware components that can be implemented using the image processing system 100.
[0065] Image processing system 100 and / or its components may be configured to perform compression and / or decompression (also referred to as encoding and / or decoding, collectively known as image encoding / decoding) using the machine learning systems and techniques described herein. In some cases, image processing system 100 and / or its components may be configured to perform image or video compression and / or decompression using the techniques described herein. In some examples, the machine learning system may leverage a deep learning neural network architecture to perform compression and / or decompression of image, video, and / or audio data. By using a deep learning neural network architecture, the machine learning system can improve the efficiency and speed of compression and / or decompression of content on a device. For example, a device using the described compression and / or decompression techniques may efficiently compress one or more images using machine learning-based techniques, send one or more compressed images to a receiving device, and the receiving device may more efficiently decompress the one or more compressed images using the machine learning-based techniques described herein. As used herein, an image may refer to a still image and / or video frame associated with a frame sequence (e.g., video).
[0066] As mentioned above, neural networks are examples of machine learning systems. A neural network can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes in the input layer, processed by hidden nodes in one or more hidden layers, and output is produced by output nodes in the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network can include feature maps or activation maps, which can include artificial neurons (or nodes). Feature maps can include filters, kernels, etc. Nodes can include one or more weights used to indicate the importance of nodes in one or more layers. In some cases, a deep learning network can have a series of many hidden layers, where earlier layers are used to determine simple and low-level features of the input, while later layers build a hierarchy of more complex and abstract features.
[0067] Deep learning architectures can learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to recognize spectral power at a specific frequency. The second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Higher layers can learn to recognize common visual objects or spoken phrases.
[0068] Therefore, deep learning architectures can perform particularly well when applied to problems with natural hierarchical structures. For example, classifying motor vehicles can benefit from first learning to identify features such as wheels, windshields, and others. These features can then be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0069] Neural networks can be designed with a variety of connection patterns. In feedforward networks, information is passed from lower layers to higher layers, where each neuron in a given layer communicates with neurons in higher layers. As mentioned above, hierarchical representations can be built into successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be fed to another neuron in the same layer. Recurrent architectures can help identify patterns across more than one block of input data that is sequentially passed to the neural network. Connections from neurons in a given layer to neurons in lower layers are called feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognizing higher-level concepts might help distinguish specific lower-level features of the input.
[0070] The connections between layers in a neural network can be fully connected or locally connected. Figure 2A An example of a fully connected neural network 202 is shown. In the fully connected neural network 202, neurons in the first layer can transmit their outputs to each neuron in the second layer, such that each neuron in the second layer receives input from each neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is shown. In the locally connected neural network 204, neurons in a first layer can connect to a limited number of neurons in a second layer. More generally, the locally connected layers of the locally connected neural network 204 can be configured such that each neuron in the layer will have the same or similar connection pattern, but the connection strength can have different values (e.g., 210, 212, 214, and 216). The connection patterns of locally connected networks may produce spatially different receptive fields in higher layers because neurons in higher layers in a given region can receive inputs that are trained on the properties of a restricted portion of the total input to the network.
[0071] An example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is shown. The convolutional neural network 206 can be configured such that the connection strength associated with the input of each neuron in the second layer is shared (e.g., 208). Convolutional neural networks may be well-suited for problems where the spatial location of the input is meaningful. According to aspects of this disclosure, the convolutional neural network 206 can be used to perform one or more aspects of video compression and / or decompression.
[0072] One type of convolutional neural network is the deep convolutional network (DCN). Figure 2D A detailed example of DCN 200 is shown, which is designed to recognize visual features from image 226 input by image capture device 230 (e.g., an in-vehicle camera). The DCN 200 in this example can be trained to recognize traffic signs and the numbers provided on them. Of course, DCN 200 can be trained for other tasks, such as recognizing lane markings or traffic lights.
[0073] The DCN 200 can be trained using supervised learning. During training, an image, such as image 226 of a speed limit sign, can be presented to the DCN 200, and the forward pass can then be computed to produce output 222. The DCN 200 may include a feature extraction part and a classification part. Upon receiving image 226, convolutional layer 232 may apply a convolutional kernel (not shown) to image 226 to generate a first set of feature maps 218. As an example, the convolutional kernel of convolutional layer 232 may be a 5x5 kernel that generates a 28x28 feature map. In this example, because four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at convolutional layer 232. A convolutional kernel may also be referred to as a filter or convolutional filter.
[0074] The first set of feature maps 218 can be second-sampled by a max-pooling layer (not shown) to generate the second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220, for example, 14x14, is smaller than the size of the first set of feature maps 218, for example, 28x28. The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved by one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).
[0075] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of the image 226, such as "sign", "60", and "100". A Softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Thus, the output 222 of the DCN 200 is the probability that the image 226 includes one or more features.
[0076] In this example, the probabilities of "symbol" and "60" in output 222 are higher than those of other values in output 222 (e.g., "30", "40", "50", "70", "80", "90", and "100"). Before training, output 222 generated by DCN 200 may be incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth value of image 226 (e.g., "symbol" and "60"). The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is closer to the target output.
[0077] To adjust the weights, the learning algorithm can compute the gradient vector of the weights. The gradient indicates how much the error will increase or decrease if the weights are adjusted. At the top layers, the gradient can directly correspond to the weight values connecting the activated neurons in the penultimate layer to the neurons in the output layer. In lower layers, the gradient can depend on the values of the weights and the error gradient computed in the higher layers. The weights can then be adjusted to reduce the error. This method of adjusting weights can be called "backpropagation" because it involves "passing backward" through the neural network.
[0078] In practice, the error gradient of the weights can be computed using a small number of examples to make the computed gradient approximate the true error gradient. This approximation method can be called stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, a new image can be presented to the DCN, and the forward pass of the network can produce an output that can be considered a speculation or prediction of the DCN.222
[0079] Deep Belief Networks (DBNs) are probabilistic models that include multiple layers of hidden nodes. DBNs can be used to extract hierarchical representations of training datasets. DBNs can be obtained by stacking layers of Restricted Boltzmann Machines (RBMs). An RBM is a type of artificial neural network that learns a probability distribution from a set of inputs. Because RBMs can learn a probability distribution without information about the class each input should be assigned to, they are often used for unsupervised learning. Using a hybrid unsupervised and supervised paradigm, the bottom RBM of a DBN can be trained unsupervised and used as a feature extractor, while the top RBM can be trained supervised (on the joint distribution of inputs from the previous layer and the target class) and used as a classifier.
[0080] Deep convolutional networks (DCNs) are networks of convolutional networks configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. DCNs can be trained using supervised learning, where both the input and output targets are known for many examples and are used to modify the network's weights using gradient descent.
[0081] DCNs can be feedforward networks. Furthermore, as mentioned above, the connections from neurons in the first layer of a DCN to groups of neurons in the next higher layer are shared among neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational burden of a DCN may be much smaller than that of a similarly sized neural network containing recurrent or feedback connections.
[0082] The processing at each layer of a convolutional network can be thought of as a spatially invariant template or fundamental projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be considered three-dimensional, with two spatial dimensions along the image's axes and a third dimension capturing color information. The outputs of the convolutional connections can be thought of as forming feature maps in subsequent layers, where each element of the feature map (e.g., 220) receives input from a series of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature maps can be further processed non-linearly, such as corrected, max(0,x). The values from neighboring neurons can be further pooled, which corresponds to downsampling and can provide additional local invariance and dimensionality reduction.
[0083] Figure 3 This is a block diagram illustrating an example of a deep convolutional network 350. The deep convolutional network 350 can include multiple layers of different types based on connectivity and weight sharing. For example... Figure 3 As shown, the deep convolutional network 350 includes convolutional blocks 354A and 354B. Each of the convolutional blocks 354A and 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max pooling layer (MAX POOL) 360.
[0084] Convolutional layer 356 may include one or more convolutional filters that can be applied to input data 352 to generate feature maps. Although only two convolutional blocks 354A and 354B are shown, this disclosure is not limited thereto, and any number of convolutional blocks (e.g., blocks 354A and 354B) may be included in the deep convolutional network 350 according to design preferences. Normalization layer 358 may normalize the output of the convolutional filters. For example, normalization layer 358 may provide whitening or lateral suppression. Max pooling layer 360 may provide spatial downsampling aggregation for local invariance and dimensionality reduction.
[0085] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the image processing system 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the image processing system 100. Furthermore, the deep convolutional network 350 can access other processing blocks that may exist on the image processing system 100, such as the sensor processor 114.
[0086] The deep convolutional network 350 may also include one or more fully connected layers, such as layer 362A (labeled "FC1") and layer 362B (labeled "FC2"). The deep convolutional network 350 may also include a logistic regression (LR) layer 364. Between each layer 356, 358, 360, 362, and 364 of the deep convolutional network 350 are weights to be updated (not shown). The output of each layer (e.g., 356, 358, 360, 362, and 364) can be used as input to the next layer in the deep convolutional network 350 (e.g., 356, 358, 360, 362, and 364) to learn hierarchical feature representations from the input data 352 (e.g., images, audio, video, sensor data, and / or other input data) provided at the first convolutional block 354A in convolutional block 354A. The output of the deep convolutional network 350 is a classification score 366 of the input data 352. The classification score 366 can be a set of probabilities, where each probability is the probability that the input data includes features from a set of features.
[0087] Image and video content can be stored on devices and / or shared between devices. For example, image and video content can be uploaded to media hosting services and sharing platforms and can be sent to various devices. Recording uncompressed image and video content typically results in large file sizes, which increase significantly with the resolution of the image and video content. For example, uncompressed 16-bit per channel video recorded in 1080p / 24 format (e.g., with a resolution of 1920 pixels wide and 1080 pixels high, where 24 frames per second are captured) can occupy 12.4 megabytes per frame, or 297.6 megabytes per second. Uncompressed 16-bit per channel video recorded at 4K resolution at 24 frames per second may occupy 49.8 megabytes per frame, or 1195.2 megabytes per second.
[0088] Because uncompressed image and video content can result in large files, which may require considerable memory for physical storage and significant bandwidth for transmission, several techniques can be used to compress such video content. For example, various compression algorithms can be applied to image and video content to reduce the size of image content, and thus reduce the amount of storage required to store the image content and the amount of bandwidth required to transmit the video content.
[0089] In some cases, a priori-defined compression algorithms can be used to compress image content, such as Joint Image Experts Group (JPEG) and Better Portable Graphics (BPG). For example, JPEG is a lossy compression form based on Discrete Cosine Transform (DCT). For instance, a device performing JPEG compression on an image can convert the image to an optimal color space (e.g., the YCbCr color space, including luminance (Y), chrominance-blue (Cb), and chrominance-red (Cr)), downsample the chrominance components by averaging groups of pixels together, and apply a DCT function to pixel blocks to remove redundant image data, thus compressing the image data. Compression is based on identifying similar regions within an image and converting these regions to the same color code (based on the DCT function). A priori-defined compression algorithms can also be used to compress video content, such as the Cinema Experts Group (MPEG) algorithm, H.264, or efficient video codecs.
[0090] These predefined compression algorithms may be able to preserve most of the information in the original image and video content and can be predefined based on ideas from signal processing and information theory. However, while these predefined compression algorithms may generally be applicable (e.g., applicable to any type of image / video content), the compression algorithms may not take into account the similarity of the content, the new resolution or frame rate of video capture and transmission, non-natural images (e.g., radar images or other images captured via various sensors), etc.
[0091] A priori defined compression algorithm is considered a lossy compression algorithm. In lossy compression of an input image (or video frame), the input image cannot be encoded and then decoded / reconstructed to reconstruct an accurate input image. Instead, in lossy compression, an approximate version of the input image is generated after decoding / reconstructing the compressed input image. Lossy compression results in a reduction in bit rate, but at the cost of distortion, leading to artifacts in the reconstructed image. Therefore, a rate-distortion tradeoff exists in lossy compression systems. For some compression methods (e.g., JPEG, BPG, etc.), distortion-based artifacts may take the form of blocky or other artifacts. In some cases, neural network-based compression can be used, and high-quality compression of image and video data is possible. In some cases, blurring and color shift are examples of artifacts.
[0092] When the bit rate is lower than the true entropy of the input data, it can be difficult or impossible to reconstruct accurate input data. However, the fact that distortion / loss occurs during data compression / decompression does not mean that the reconstructed image or frame will necessarily have artifacts. In fact, a compressed image can be reconstructed into another similar but different image with high visual quality.
[0093] As previously described, the systems and techniques described herein can use one or more machine learning (ML) systems to perform compression and decompression. In some examples, machine learning techniques can provide image and / or video compression that produces high-quality visual output. In some examples, these systems and techniques described herein can use deep neural networks such as rate-distortion autoencoders (RD-AEs) to perform compression and decompression of content (e.g., image content, video content, audio content, etc.). The deep neural network can include an autoencoder (AE) that maps images to a latent code space (e.g., including a set of codes z). The latent code space can include a code space used by the encoder and decoder, where content has been encoded into code z. The code (e.g., code z) can also be referred to as a latent quantity, latent variable, or latent representation. The deep neural network can include a probabilistic model (also referred to as a model prior or code model) that can perform lossless compression of code z from the latent code space. The probabilistic model can generate a probability distribution on the code set z, which can represent encoded data based on input data. In some cases, the probability distribution can be represented as (P(z)).
[0094] In some examples, a deep neural network may include an arithmetic encoder that generates a bitstream of compressed data to be output based on a probability distribution P(z) and / or a code set z. The bitstream containing the compressed data may be stored and / or sent to a receiving device. The receiving device may use, for example, an arithmetic decoder, a probabilistic (or code) model, and a decoder of the AE to perform the inverse process to decode or decompress the bitstream. The device that generated the bitstream containing the compressed data may also perform a similar decoding / decompression process when retrieving the compressed data from storage. Similar techniques can be used to compress / encode and decompress / decode updated model parameters.
[0095] In some examples, RD-AEs can be trained and operated to perform as multi-rate AEs (including high-rate and low-rate operations). For example, the latent code space generated by the encoder of a multi-rate AE can be divided into two or more blocks (e.g., code z is divided into blocks z1 and z2). In high-rate operation, the multi-rate AE can send a bitstream based on the entire latent space (e.g., code z, including z1, z2, etc.), which the receiving device can use to decompress data, similar to the operation described above for RD-AEs. In low-rate operation, the bitstream sent to the receiving device is based on a subset of the latent space (e.g., block z1 instead of z2). The receiving device can infer the remainder of the latent space based on the sent subset and can use the subset of the latent space and the inferred remainder of the latent space to generate reconstructed data.
[0096] Encoding and decoding mechanisms can be adapted to a variety of use cases by using RD-AE or multi-rate AE to compress (and decompress) content. Machine learning-based compression techniques can generate compressed content with high quality and / or reduced bitrate. In some examples, RD-AE can be trained to minimize the average rate distortion loss over a dataset of data points (e.g., image and / or video data points). In some cases, RD-AE can also be fine-tuned for specific data points to be sent to and decoded by a receiver. In some examples, high compression (rate / distortion) performance can be achieved by fine-tuning RD-AE on data points. The encoder associated with the RD-AE can send the AE model or a portion of the AE model to the receiver (e.g., decoder) to decode the bitstream.
[0097] In some cases, neural network compression systems can reconstruct input instances (e.g., input images, videos, audio, etc.) from (quantized) latent representations. Neural network compression systems can also use priors to perform lossless compression of the latent representation. In some cases, neural network compression systems can determine that the data distribution at test time is known and relatively low in entropy (e.g., a camera viewing a static scene, a dashcam in an autonomous vehicle, etc.) and can be fine-tuned or adapted to such a distribution. Fine-tuning or adaptation can result in improved rate / distortion (RD) performance. In some examples, the model of a neural network compression system can be adapted to a single input instance to be compressed. Neural network compression systems can provide model updates, and in some examples, these updates can be quantized and compressed using parameter space priors and latent representations.
[0098] Fine-tuning can take into account the effects of model quantization and the additional cost of sending model updates. In some examples, the neural network compression system can be fine-tuned using the RD loss and an additional model rate term M, which measures the number of bits required to send model updates given the model prior, resulting in a combined RDM loss.
[0099] Figure 4This is a diagram illustrating a system 400 including a transmitting device 410 and a receiving device 420, according to some examples of the present disclosure. Each of the transmitting device 410 and the receiving device 420 may be referred to as RD-AE in some cases. The transmitting device 410 can compress image content and can store the compressed image content and / or send the compressed image content to the receiving device 420 for decompression. The receiving device 420 can decompress the compressed image content and can output the decompressed image content on the receiving device 420 (e.g., for display, editing, etc.) and / or can output the decompressed image content to other devices connected to the receiving device 420 (e.g., a television, mobile device, or other device). In some cases, the receiving device 420 can become a transmitting device by compressing image content (using encoder 422) and storing and / or sending the compressed image content to another device, such as the transmitting device 410 (in which case the transmitting device 410 becomes a receiving device). Although system 400 is described herein for the purpose of image compression and decompression, those skilled in the art will understand that system 400 can use the techniques described herein to compress and decompress video content.
[0100] like Figure 4 As shown, transmitting device 410 includes an image compression pipeline, while receiving device 420 includes an image bitstream decompression pipeline. According to aspects of this disclosure, the image compression pipeline in transmitting device 410 and the bitstream decompression pipeline in receiving device 420 typically use one or more artificial neural networks to compress image content and / or decompress received bitstreams into image content. The image compression pipeline in transmitting device 410 includes an autoencoder 401, a code model 404, and an arithmetic encoder 406. In some embodiments, the arithmetic encoder 406 is optional and can be omitted in some cases. The image decompression pipeline in receiving device 420 includes an autoencoder 421, a code model 424, and an arithmetic decoder 426. In some embodiments, the arithmetic decoder 426 is optional and can be omitted in some cases. The autoencoder 401 and code model 404 of transmitting device 410 are... Figure 4 The machine learning system is shown as having been previously trained and thus configured to perform operations during the inference or operation of the trained machine learning system. The autoencoder 421, code model 424, and completion model 425 are also shown as previously trained machine learning systems.
[0101] The autoencoder 401 includes an encoder 402 and a decoder 403. The encoder 402 performs lossy compression on the received uncompressed image content by mapping pixels in one or more images of the uncompressed image content to a latent code space (including code z). Typically, the encoder 402 can be configured such that the code z representing the compressed (or encoded) image is discrete or binary. These codes can be generated based on random perturbation techniques, soft vector quantization, or other techniques that can generate different codes. In some aspects, the autoencoder 401 can map the uncompressed image to codes with a compressible (low-entropy) distribution. The cross-entropy of these codes can approximate a predefined or learned prior distribution.
[0102] In some examples, the autoencoder 401 can be implemented using a convolutional architecture. For instance, in some cases, the autoencoder 401 can be configured as a two-dimensional convolutional neural network (CNN) so that the autoencoder 401 learns spatial filters for mapping image content to a latent code space. In an example where system 400 is used to encode and decode video data, the autoencoder 401 can be configured as a three-dimensional CNN such that the autoencoder 401 learns spatiotemporal filters for mapping video to a latent code space. In such a network, the autoencoder 401 can encode the video based on keyframes (e.g., the initial frame marking the start of a frame sequence, where subsequent frames in the sequence are described as differences relative to the initial frame in the sequence), warpage (or differences) between the keyframes and other frames in the video, and residual factors. In other aspects, the autoencoder 401 can be implemented as a two-dimensional neural network conditioned on previous frames and residual factors between frames, and conditioned by stacking channels or including recurrent layers.
[0103] The encoder 402 of the automatic encoder 401 can receive the first image (in Figure 4 The encoder 402 takes an image (x) as input and maps it to a code z in the latent code space. As described above, the encoder 402 can be implemented as a two-dimensional convolutional network such that the latent code space has a vector at each (x, y) location describing a patch of image x centered at that location. The x-coordinate can represent the horizontal pixel position in that patch of image x, and the y-coordinate can represent the vertical pixel position in that patch of image x. When encoding and decoding video data, the latent code space can have a t variable or a position, where the t variable represents the timestamp in the patch of video data (in addition to the spatial x and y coordinates). By using the two dimensions of horizontal and vertical pixel positions, the vector can describe an image patch in image x.
[0104] The decoder 403 of the autoencoder 401 can then decompress the code z to obtain a reconstruction of the first image x. Typically, refactoring It can be an approximation of the uncompressed first image x, without needing to be an exact copy of the first image x. In some cases, the reconstructed image... It can be output as a compressed image file for storage on the sending device.
[0105] Code model 404 receives a code z representing an encoded image or a portion thereof and generates a probability distribution P(z) over a set of compressed codewords that can be used to represent code z. In some examples, code model 404 may include a probabilistic autoregressive generative model. In some cases, the code for which the probability distribution can be generated includes a learned distribution based on an arithmetic encoder-decoder 406 to control bit allocation. For example, using the arithmetic encoder 406, the compressed code for a first code z can be predicted individually; the compressed code for a second code z can be predicted based on the compressed code for the first code z; the compressed code for a third code z can be predicted based on the compressed codes for the first and second codes z, and so on. The compressed code typically represents different spatiotemporal blocks of a given image to be compressed.
[0106] In some respects, z can be represented as a three-dimensional tensor. The three dimensions of the tensor can include the feature channel dimension, as well as the height and width spatial dimensions (e.g., represented as code z). c,w,h Each code z c,w,h (The codes, representing those indexed by channel and horizontal and vertical positions,) can all be used for prediction based on previous codes, which can be in a fixed order or theoretically arbitrarily ordered. In some examples, codes can be generated by analyzing a given image file from beginning to end and analyzing each block in the image in raster scan order.
[0107] Code model 404 can use a probabilistic autoregressive model to learn the probability distribution of the input code z. The probability distribution can be conditional on its previous values (as described above). In some examples, the probability distribution can be represented by the following formula:
[0108]
[0109] Where c is the channel index of all image channels C (e.g., R, G, and B channels, Y, Cb, and Cr channels, or other channels), w is the width index of the total image frame width W, and h is the height index of the total image frame height H.
[0110] In some examples, a probability distribution P(z) can be predicted using a fully convolutional neural network with causal convolution. In some aspects, the kernels of each layer of the convolutional neural network can be masked, allowing the network to know previous values z when calculating the probability distribution. 0:c,0:w,0:hAnd it may be unaware of other values. In some respects, the last layer of a convolutional network may include a softmax function that determines the probability that a code in the latent space is suitable for the input values (e.g., the probability that a given code can be used to compress a given input).
[0111] Arithmetic encoder 406 uses the probability distribution P(z) generated by code model 404 to generate a bitstream 415 corresponding to the prediction of code z (in Figure 4 (This is illustrated as "0010011..."). The prediction of code z can be represented as the code with the highest probability score among the probability distribution P(z) generated over a set of possible codes. In some aspects, the arithmetic encoder 406 can output a variable-length bitstream based on the accuracy of the prediction of code z and the actual code z generated by the autoencoder 401. For example, if the prediction is accurate, the bitstream 415 can correspond to a short codeword, while as the magnitude of the difference between code z and the prediction of code z increases, the bitstream 415 can correspond to a longer codeword.
[0112] In some cases, bitstream 415 can be output by arithmetic encoder 406 for storage in a compressed image file. Bitstream 415 can also be output for transmission to a requesting device (e.g., receiving device 420, such as...). Figure 4 (As shown). Typically, the bitstream 415 output by the arithmetic encoder 406 can losslessly encode z, so that z can be accurately recovered during the decompression process applied to the compressed image file.
[0113] The bit stream 415 generated by the arithmetic encoder 406 and transmitted from the transmitting device 410 can be received by the receiving device 420. Transmission between the transmitting device 410 and the receiving device 420 can occur using any of a variety of suitable wired or wireless communication technologies. Communication between the transmitting device 410 and the receiving device 420 can be direct or can be performed through one or more network infrastructure components (e.g., base stations, relay stations, mobile stations, network hubs, routers, and / or other network infrastructure components).
[0114] As shown in the figure, receiving device 420 may include an arithmetic decoder 426, a code model 424, and an autoencoder 421. Autoencoder 421 includes an encoder 422 and a decoder 423. For a given input, decoder 423 may produce the same or similar output as decoder 403. Although autoencoder 421 is shown as including encoder 422, encoder 422 is not required during the decoding process to obtain the code z received from transmitting device 410. (For example, an approximation of the original image x compressed at the transmitting device 410).
[0115] The received bitstream 415 can be fed into an arithmetic decoder 426 to obtain one or more codes z from the bitstream. The arithmetic decoder 426 can extract the decompressed code z based on the probability distribution P(z) generated by the code model 424 over a set of possible codes and information associated with each generated code z in the bitstream. Given a received portion of the bitstream and a probability prediction of the next code z, the arithmetic decoder 426 can generate a new code z, as it is encoded by the arithmetic encoder 406 at the transmitting device 410. Using the new code z, the arithmetic decoder 426 can make probability predictions for consecutive codes z, read additional portions of the bitstream, and decode consecutive codes z until the entire received bitstream is decoded. The decompressed code z can be provided to the decoder 423 in the autoencoder 421. The decoder 423 decompresses the code z and outputs an approximation of the image content x. (This can be referred to as a reconstructed or decoded image.) In some cases, an approximation of content x can be stored. For later retrieval. In some cases, the content x is approximated. It can be recovered by the receiving device 420 and displayed on a screen that is communicatively coupled or integrated with the receiving device 420.
[0116] As described above, the automatic encoder 401 and code model 404 of the transmitting device 410 are in Figure 4 The image shown is a previously trained machine learning system. In some respects, the autoencoder 401 and code model 404 can be trained together using image data. For example, the encoder 402 of the autoencoder 401 can receive a first training image n as input and can map the first training image n to code z in the latent code space. The code model 404 can learn the probability distribution P(z) of code z using a probabilistic autoregressive model (similar to the technique described above). The arithmetic encoder 406 can generate an image bitstream using the probability distribution P(z) generated by the code model 404. Using the bitstream from the code model 404 and the probability distribution P(z), the arithmetic encoder 406 can generate code z and can output code z to the decoder 403 of the autoencoder 401. The decoder 403 can then decompress the code z to obtain a reconstruction of the first training image n. (where reconstruction) It is an approximation of the uncompressed first training image n).
[0117] In some cases, the backpropagation engine used during training of the transmitting device 410 can perform a backpropagation process to tune the parameters (e.g., weights, biases, etc.) of the neural networks of the autoencoder 401 and code model 404 based on one or more loss functions. In some cases, the backpropagation process can be based on stochastic gradient descent. Backpropagation can include forward propagation, one or more loss functions, backpropagation, and weight (and / or other parameter) updates. For a single training iteration, forward propagation, loss function, backpropagation, and parameter updates can be performed. This process can be repeated a certain number of iterations for each set of training data until the weights and / or other parameters of the neural network are accurately tuned.
[0118] For example, the auto encoder 401 can compare n and To determine the first training image n and the reconstructed first training image The loss between the two (e.g., represented by a distance vector or other difference). The loss function can be used to analyze the error in the output. In some examples, the loss can be based on maximum likelihood. This is used when using an uncompressed image n as input and reconstructing an image... As an illustrative example of the output, the loss function Loss = D + beta * R can be used to train a neural network system of autoencoder 401 and code model 404, where R is the rate, D is the distortion, * denotes the multiplication function, and beta is a trade-off parameter set to a value defining the bit rate. In another example, the loss function... Neural network systems that can be used to train autoencoder 401 and code model 404. In some cases, such as when using different training data, other loss functions can be used. An example of another loss function is mean squared error (MSE), defined as… MSE is calculated as half the sum of the actual answer and the square of the predicted (output) answer.
[0119] Based on a determined loss (e.g., distance vector or other difference) and using a backpropagation process, the parameters (e.g., weights, biases, etc.) of the neural network system of autoencoder 401 and code model 404 can be adjusted (effectively adjusting the mapping between the received image content and the latent code space) to reduce the loss between the input uncompressed image and the compressed image content generated as output by autoencoder 401.
[0120] For the first training image, the loss (or error) may be high because the actual output value (reconstructed image) may differ significantly from the input image. The goal of training is to minimize the loss of the predicted output. A neural network can perform backpropagation by determining which nodes (with corresponding weights) contribute most to the network's loss, and the weights (and / or other parameters) can be adjusted to reduce the loss and eventually minimize it. The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute most to the network's loss. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, w i η represents the initial weights, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate involves larger weight updates, while a lower value indicates smaller weight updates.
[0121] The neural network system of autoencoder 401 and code model 404 can continue to be trained in this manner until the desired output is obtained. For example, autoencoder 401 and code model 404 can repeat the backpropagation process to minimize or otherwise reduce the input image n and the reconstructed image resulting from the decompression of the generated code z. The difference between them.
[0122] The autoencoder 421 and code model 424 can be trained using techniques similar to those described above for training the autoencoder 401 and code model 404 of the transmitting device 410. In some cases, the autoencoder 421 and code model 424 can be trained using the same or different training datasets used for training the autoencoder 401 and code model 404 of the transmitting device 410.
[0123] exist Figure 4 In the example shown, the rate distortion autoencoder (transmitting device 410 and receiving device 420) is trained and operates in speculation based on the bit rate. In some implementations, the rate distortion autoencoder can be trained at multiple bit rates to allow the generation and output of high-quality reconstructed image or video frames (e.g., with little or no artifacts due to distortion relative to the input image) when different amounts of information are provided in the latent code z.
[0124] In some implementations, the latent code z can be divided into at least two blocks, z1 and z2. When using the RD-AE model at a high rate setting, both blocks are sent to the device for decoding. When using the rate distortion autoencoder model at a low rate setting, only block z1 is sent, and block z2 is inferred from z1 on the decoder side. Various techniques can be used to perform the inference of z2 from z1, as described in more detail below.
[0125] In some implementations, a set of continuous latents (e.g., those conveying a large amount of information) and corresponding quantized discrete latents (e.g., those containing less information) can be used. After training the RD-AE model, an auxiliary dequantization model can be trained. In some cases, when using RD-AE, only discrete latents are transmitted, and the continuous latents are inferred from the discrete latents at the decoder using the auxiliary dequantization model.
[0126] Although system 400 is shown to include certain components, those skilled in the art will understand that system 400 may include more than Figure 4 The components shown may include more or fewer components. For example, the transmitting device 410 and / or receiving device 420 of system 400 may, in some cases, also include one or more memory devices (e.g., RAM, ROM, cache, etc.), one or more network interfaces (e.g., wired and / or wireless communication interfaces, etc.), one or more display devices, and / or Figure 4 Other hardware or processing devices not shown. Figure 4 The components shown and / or other components of system 400 may be implemented using one or more computing or processing components. One or more computing components may include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), and / or an image signal processor (ISP). The following describes... Figure 14 This section describes illustrative examples of computing devices and hardware components that can be implemented using System 1400.
[0127] System 400 may be a single computing device or part of or implemented by multiple computing devices. In some examples, transmitting device 410 may be part of a first device and receiving device 420 may be part of a second computing device. In some examples, transmitting device 410 and / or receiving device 420 may be included as part of electronic devices (or devices) such as: telephone systems (e.g., smartphones, cellular phones, conferencing systems, etc.), desktop computers, laptops or notebooks, tablets, set-top boxes, smart TVs, display devices, game consoles, video streaming devices, SOCs, IoT (Internet of Things) devices, smart wearable devices (e.g., head-mounted displays (HMDs), smart glasses, etc.), camera systems (e.g., digital cameras, IP cameras, camcorders, security cameras, etc.) or any other suitable electronic devices. In some cases, system 400 may be comprised of... Figure 1 The image processing system 100 shown is used to implement this. In other cases, system 400 may be implemented by one or more other systems or devices.
[0128] Figure 5A This is a diagram illustrating an example neural network compression system 500. In some examples, the neural network compression system 500 may include an RD-AE system. Figure 5A In this system, the neural network compression system 500 includes an encoder 502, an arithmetic encoder 508, an arithmetic decoder 512, and a decoder 514. In some cases, encoder 502 and / or decoder 514 may be the same as encoder 402 and / or decoder 403, respectively. In other cases, encoder 502 and / or decoder 514 may be different from encoder 402 and / or decoder 403, respectively.
[0129] Encoder 502 can receive image 501 (image x) i ) can be used as input and image 501 (image x) can be used as input. i Mapping and / or transforming to latent code 504 (latent z) in the latent code space i Image 501 may represent a still image and / or video frame associated with a frame sequence (e.g., video). In some cases, encoder 502 may perform a forward pass to generate latent code 504. In some examples, encoder 502 may implement a learnable function. In some cases, encoder 502 may implement a function derived from... Parameterized learnable functions. For example, encoder 502 can implement functions. In some examples, the learnable function does not need to be shared with or known by the decoder 514.
[0130] Arithmetic encoder 508 can be based on latent code 504 (latent z). iThe latent prior 506 is used to generate the bitstream 510. In some examples, the latent prior 506 can implement a learnable function. In some cases, the latent prior 506 can implement a learnable function parameterized by ψ. For example, the latent prior 506 can implement the function p. ψ (z)). The latent prior 506 can be used to compress the latent code 504 (latent z) using lossless compression. i The potential prior 506 can be shared and / or available at both the sender (e.g., encoder 502 and / or arithmetic encoder 508) and the receiver (e.g., arithmetic decoder 512 and / or decoder 514).
[0131] The decoder 514 can receive the encoded bitstream 510 from the arithmetic encoder 508 and use the latent prior 506 to process the latent code 504 (latent z) in the encoded bitstream 510. i Decoding is performed. Decoder 514 can decode the latent code 504 (latent z). i ) decoded into an approximate reconstructed image 516 (reconstructed) In some cases, decoder 514 can implement learnable functions parameterized by θ. For example, decoder 514 can implement the function p. θ (x|z). The learnable function implemented by decoder 514 can be shared and / or available at both the sender (e.g., encoder 502 and / or arithmetic encoder 508) and the receiver (e.g., arithmetic decoder 512 and / or decoder).
[0132] A neural network compression system 500 can be trained to minimize rate distortion. In some examples, the rate reflects the length of the bitstream 510 (bitstream b), while the distortion reflects the image 501 (image x). i ) and reconstructed image 516 (reconstructed) The parameter β can be used to train a model for a specific rate-to-distortion ratio. In some examples, the parameter β can be used to define and / or achieve a trade-off between rate and distortion.
[0133] In some examples, the loss can be represented as follows Here, the function E is the expected value. Distortion (x|z; θ) can be determined based on a loss function such as mean squared error (MSE). In some examples, the term -logp... θ (x|z) can indicate and / or represent the distortion D(x|z; θ).
[0134] The rate used to transmit potential values can be expressed as R. z (z;ψ). In some examples, the term logp ψ (z) can indicate and / or represent the rate Rz (z;ψ). In some cases, the loss can be minimized over the entire dataset D, as shown below:
[0135] Figure 5B This is a diagram illustrating the inference process 530 performed by the neural network compression system 500. As shown, the encoder 502 can convert an image 501 into a latent code 504. In some examples, the image 501 may represent still images and / or video frames associated with a frame sequence (e.g., video).
[0136] In some examples, encoder 502 can use a single forward pass. Image 501 is encoded. Arithmetic encoder 508 can then be used to encode latent code 504 (latent z) under latent prior 506. i Arithmetic encoding of 520 to generate bitstream In some examples, the arithmetic encoder 508 can generate bitstream 520 as follows:
[0137] Arithmetic decoder 512 can receive bitstream 520 from arithmetic encoder 508 and execute latent code 504 (latent z) under latent prior 506. i Arithmetic decoding of 504 from bitstream 520. In some examples, arithmetic decoder 512 can decode the underlying code 504 from bitstream 520 as follows: Decoder 514 can handle latent code 504 (latent z). i Decode and generate reconstructed image 516 (reconstructed) In some examples, decoder 514 can use a single forward pass to decode latent code 504 (latent z). i The decoding is as follows:
[0138] In some examples, the RD-AE system can be trained using a set of training data and further fine-tuned on data points (e.g., image data, video data, audio data) that will be sent to a receiver (e.g., a decoder) and decoded by it. For example, during speculation, the RD-AE system can be fine-tuned on image data sent to the receiver. Since compressed models are typically large, sending the parameters associated with the model to the receiver can be very expensive in terms of resources such as network (e.g., bandwidth), storage, and computational resources. In some cases, the RD-AE system can be fine-tuned on a single compressed data point and sent to the receiver for decompression. This can limit the amount of information sent to the receiver (and associated costs) while maintaining and / or improving compression / decompression efficiency, performance, and / or quality.
[0139] Figure 6 This is a diagram illustrating an example neural network compression system 600 for instance-adaptive data compression. The neural network compression system 600 can be trained and further fine-tuned for the compressed data to provide compression adapted to / fine-tuned to the compressed data (e.g., instance adaptation). In this example, the neural network compression system 600 is shown as implementing an average-scale super-prior model architecture using a variable autoencoder (VAE) framework. In some cases, a shared superdecoder can be used to predict the mean and scale parameters of the average-scale super-prior model.
[0140] like Figure 6 As shown, the neural network compression system 600 can be trained using training dataset 602. Training dataset 602 can be processed by encoder 606 of codec 604 to generate a latent spatial representation 608(z2) of training dataset 602. Encoder 606 can provide the latent spatial representation 608(z2) to decoder 610 of codec 604 and superencoder 614 of supercoder 612.
[0141] The super encoder 614 can use the latent space representation 608(z2) and the latent prior 620 to generate a super latent space representation 616(z1) of the training dataset 602. In some examples, the super latent space representation 616(z1) and the super latent space representation 616(z1) can provide a hierarchical latent variable model for the latent space z = {z1, z2}.
[0142] The superdecoder 618 of the supercoder 612 can generate a super-prior model 624 using a super-latent space representation 616(z1). The superdecoder 618 can predict the mean and scaling parameters of the super-prior model 624. In some examples, the super-prior model 624 may include probability distributions over the parameters of the latent space representation 608(z2) and the super-latent space representation 616(z1). In some examples, the super-prior model 624 may include probability distributions over the parameters of the latent space representation 608(z2), the super-latent space representation 616(z1), and the superdecoder 618.
[0143] Decoder 610 can use the super-prior model 624 and the latent space representation 608(z2) to generate reconstructed data 626 for the training dataset 602. ).
[0144] In some examples, during training, the neural network compression system 600 can implement a hybrid quantization strategy, wherein the quantized latent space representation 608(z2) is used to compute the distortion loss, and noisy samples are used for the latent space representation 608(z2) and the hyperlatent space representation 616(z1) when computing the rate loss.
[0145] Figure 7 This is a diagram illustrating an example architecture of a neural network compression system 700 that uses model prior fine-tuning (e.g., instance fitting). In some examples, the neural network compression system 700 may include an RD-AE system that uses RDM-AE model prior fine-tuning. The neural network compression system 700 may include an encoder 702, an arithmetic encoder 706, an arithmetic decoder 716, and a decoder 718. In some cases, encoder 702 may be the same as or different from encoder 402, encoder 502, or encoder 606, and decoder 718 may be the same as or different from decoder 403, decoder 514, or decoder 610. Arithmetic encoder 706 may be the same as or different from arithmetic encoder / decoder 406 or arithmetic encoder 508, while arithmetic decoder 716 may be the same as or different from arithmetic decoder 426 or 508.
[0146] Encoder 702 can receive image 701 (image x) i ) can be used as input and image 701 (image x) can be used as input. i Mapping and / or transforming to latent code 704 (latent z) in the latent code space i In some examples, image 701 may represent a still image and / or video frame associated with a frame sequence (e.g., video). In some examples, image 701 may be derived from a training dataset (e.g., ...) before processing image 701. Figure 6 The encoder 702 is trained on the training dataset 602. Furthermore, the encoder 702 can be further trained or fine-tuned (e.g., instance adaptation) on image 701. For example, image 701 can be used to fine-tune the encoder 702. Fine-tuning the encoder 702 on image 701 can result in high compression performance of image 701. For example, fine-tuning the encoder 702 can allow the neural network compression system 700 to improve the rate distortion of the compressed image 701.
[0147] In some cases, encoder 702 can generate latent code 704 using a single forward pass. In some examples, encoder 702 can implement a learnable function. In some cases, the learnable function can be derived from... Parameterization. For example, encoder 702 can implement functions. In some examples, the learnable function does not need to be shared with or known by the arithmetic decoder 716 and / or decoder 718.
[0148] The arithmetic encoder 706 can entropy encode the latent code 704 using the latent prior 708 and generate a bitstream 710 (e.g., bitstream) that will be sent to the arithmetic decoder 716. Bitstream 710 may include compressed data representing latent code 704. Latent prior 708 can be used to compress latent code 704 (latent z) using lossless compression. i The latent prior 708 is converted to a bitstream 710. The latent prior 708 can be shared and / or available at both the sender (e.g., encoder 702 and / or arithmetic encoder 706) and the receiver (e.g., arithmetic decoder 716 and / or decoder 718). In some examples, the latent prior 708 can implement a learnable function. In some cases, the learnable function can be parameterized by ψ. For example, the latent prior 708 can implement the function p. ψ (z).
[0149] The neural network compression system 700 may also include a model prior 714. The model prior 714 may include a probability distribution over the parameters of a latent prior 708 and the decoder 718. In some examples, the model prior 714 may implement a learnable function. In some cases, the learnable function may be parameterized by ω. For example, the model prior 714 may implement a function p. ω (ψ|θ). The neural network compression system 700 can use model priors 714 to convert the fine-tuned parameters of the encoder 702 and latent priors 708 into a bitstream 712 (e.g., bitstream) to be sent to the arithmetic decoder 716. ).
[0150] The model prior 712 can be shared and / or available at both the sender (e.g., encoder 702 and / or arithmetic encoder 706) and the receiver (e.g., arithmetic decoder 716 and / or decoder 718). For example, arithmetic encoder 706 can send bitstream 712 to arithmetic decoder 716 for decoding bitstream 710 generated from latent code 704. Bitstream 712 may include compressed data representing fine-tuned parameters of encoder 702 and latent prior 708, which arithmetic decoder 716 can use to obtain the fine-tuned parameters of encoder 702 and latent prior 708. Arithmetic decoder 716 and decoder 718 can reconstruct image 701 based on latent code 704 obtained from bitstream 710 and the fine-tuned parameters of encoder 702 and latent prior 708 obtained from bitstream 712.
[0151] Arithmetic decoder 716 can use latent prior 708 and model prior 714 to convert bitstream 710 into latent code 704 (latent z). i For example, the arithmetic decoder 716 can use the latent prior 708 and the model prior 714 to decode the bitstream 710. The decoder 718 can use the latent code 704 (latent z) decoded by the arithmetic decoder 716. i To generate the reconstructed image 720 (reconstructed) For example, decoder 718 can decode latent code 704 (latest z). i ) decoded into an approximately reconstructed image 720 (reconstructed) In some examples, decoder 718 may use latent code 704, latent prior 708, and / or model prior 714 to generate reconstructed image 720.
[0152] In some cases, the decoder 718 can implement a learnable function parameterized by θ. For example, the decoder 718 can implement a learnable function p. θ (x|z). The learnable functions implemented by decoder 718 can be shared and / or available at both the sender (e.g., encoder 702 and / or arithmetic encoder 706) and the receiver (e.g., arithmetic decoder 716 and / or decoder 718).
[0153] In some examples, on the encoder side, the rate distortion model (RDM) loss can be used on image 701 (image x). i The model was fine-tuned as follows: In some examples, a single forward pass of the fine-tuned encoder 702 can be used to process image 701 (image x). i The encoding is as follows: In some cases, the finely tuned latent prior 708 can be entropy encoded as follows: Furthermore, the fine-tuned decoder 718 and / or arithmetic encoder 706 can be entropy encoded / decoded as follows: In some cases, the potential code 704 (potential z) i It can be encoded and decoded by entropy as follows:
[0154] On the decoder side, in some examples, the finely tuned latent prior 708 can be entropy encoded / decoded as follows: Furthermore, the fine-tuned decoder 718 and / or arithmetic decoder 716 can be entropy encoded / decoded as follows: In some cases, the latent code 704 (latent zi) can also be entropy encoded as follows:
[0155] In some examples, decoder 718 can use a single forward pass of a fine-tuned decoder (e.g., decoder 718) to decode latent code 704 (latest z). i ) decoded into an approximate reconstructed image 720 (reconstruction) )as follows
[0156] Figure 8This is a diagram illustrating an example inference process implemented by an example neural network compression system 800 fine-tuned using model priors. In some examples, the neural network compression system 800 may include an RD-AE system fine-tuned using an RDM-AE model prior. In some cases, the neural network compression system 800 may include an AE model fine-tuned using model priors.
[0157] In this illustrative example, the neural network compression system 800 includes an encoder 802, an arithmetic encoder 808, an arithmetic decoder 812, a decoder 814, a model prior 816, and a latent prior 806. In some cases, encoder 802 may be the same as or different from encoder 402, encoder 502, encoder 606, or encoder 702, and decoder 814 may be the same as or different from decoder 403, decoder 514, decoder 610, or decoder 718. Arithmetic encoder 808 may be the same as or different from arithmetic encoder / decoder 406, arithmetic encoder 508, or arithmetic encoder 706, while arithmetic decoder 812 may be the same as or different from arithmetic decoder 426, arithmetic decoder 508, or arithmetic decoder 716.
[0158] The neural network compression system 800 can generate latent code 804 (latent z) for image 801. i The neural network compression system 800 can use latent code 804 and latent prior 806 to compress image 801 (image x). i Encode the image and generate a reconstructed image 820 that can be used by the receiver. The bitstream 810. In some examples, image 801 may represent still images and / or video frames associated with a frame sequence (e.g., video).
[0159] In some examples, the neural network compression system 800 can be fine-tuned using the RDM-AE loss. The neural network compression system 800 can be trained by minimizing the rate-distortion-model rate (RDM) loss. In some examples, on the encoder side, the RDM loss can be used on image 801 (image x). i The AE model was fine-tuned as follows:
[0160] The finely tuned encoder 802 can process image 801 (image x) i The encoder 802 encodes the image 801 (image x) to generate the potential code 804. In some cases, the fine-tuned encoder 802 can use a single forward pass as follows to encode the image 801 (image x). i Encode Arithmetic encoder 808 can use latent prior 806 to convert latent code 804 into a bitstream 810 for arithmetic decoder 812. Arithmetic encoder 808 can entropy-encode the parameters of fine-tuned decoder 814 and fine-tuned latent prior 806 under model prior 816, and generate a bitstream 811 including compressed parameters of fine-tuned decoder 814 and fine-tuned latent prior 806. In some examples, bitstream 811 may include updated parameters of fine-tuned decoder 814 and fine-tuned latent prior 806. Updated parameters may include, for example, parameter updates relative to the baseline decoder and latent prior (e.g., decoder 814 and latent prior 806 before fine-tuning).
[0161] In some cases, the finely tuned latent prior 806 can be entropy encoded under the model prior 816 as follows: The finely tuned decoder 814 can perform entropy encoding under the model prior 816 as follows: And the potential code 804 (potential z) i Entropy encoding can be performed under a finely tuned latent prior of 806 as follows: In some cases, on the decoder side, the finely tuned latent prior 806 can be entropy encoded under the model prior 816 as follows: The finely tuned decoder 814 can perform entropy encoding under the model prior 816 as follows: And the potential code 804 (potential z) i Entropy encoding can be performed under a finely tuned latent prior of 806 as follows:
[0162] Decoder 814 can output latent code 804 (latent z). i ) decoded into an approximate reconstructed image 820 (reconstruction) In some examples, decoder 814 can use a single forward pass of a finely tuned decoder to decode the underlying code 804 as follows:
[0163] As previously mentioned, the neural network compression system 800 can be trained by minimizing the RDM loss. In some cases, the rate can reflect the length of the bitstream b (e.g., bitstreams 810 and / or 811), and the distortion can reflect the input image 801 (image x). i ) and reconstructed image 820 (reconstructed) The model-rate can reflect the length of the bitstream required to send model updates (e.g., updated parameters) to the receiver (e.g., to decoder 814) and / or for that operation. The parameter β can be used to train the model for a specific rate-distortion ratio.
[0164] In some examples, the loss for data point x can be minimized during inference, as shown below: In some examples, the RDM loss can be represented as follows: In some cases, the distortion D(x|z;θ) can be determined based on a loss function such as mean squared error (MSE).
[0165] Terminology - logp θ (x|z) can indicate and / or represent the distortion D(x|z; θ). The term βlogp ψ (z) can indicate and / or represent the use of sending potential R z The rate of (z; ψ), and the term βlogp ω (ψ,θ) can indicate and / or represent the parameters used to send fine-tuned model updates R. ψ,θ The rate of (ψ,θ;ω).
[0166] In some cases, the model prior 816 can reflect the length of the bit rate overhead used to send model updates. In some examples, the bit rate used to send model updates can be described as follows: In some cases, model priors can be chosen to make sending models without updates cheap, i.e., the bit length (model-rate-loss) is small:
[0167] In some cases, using the RDM loss function, the neural network compression system 800 can add bits for model updates only if the potential rate or distortion decreases by at least as many bits. This can improve rate-distortion (R / D) performance. For example, if the rate or distortion could decrease with at least the same number of bits, the neural network compression system 800 can increase the number of bits in the bitstream 811 used for sending model updates. In other cases, the neural network compression system 800 can add bits for model updates. In a bitstream, even if the potential rate or distortion does not decrease with at least the same number of bits.
[0168] The neural network compression system 800 can be trained end-to-end. In some cases, the RDM loss can be minimized during end-to-end inference. In some examples, a large amount of computation can be spent upfront (e.g., fine-tuning the model) and a high compression ratio can subsequently be obtained without additional costs on the receiver side. For example, a content provider might spend significant computation to train and fine-tune the neural network compression system 800 more broadly for a video that will be delivered to a large number of receivers. The highly trained and fine-tuned neural network compression system 800 can provide high compression performance for that video. After spending a large amount of computation, the video provider can store the updated parameters of the model prior and efficiently provide them to each receiver of the compressed video for decompression. The video provider can gain significant benefits in compression (and reduced network and computational resources) with each video transmission, which far outweighs the initial computational cost of training and fine-tuning the model.
[0169] Due to the large number of pixels in videos and high-resolution images, the learning and fine-tuning methods presented in this paper can be highly advantageous for video compression and / or high-resolution images. In some cases, complexity and / or decoder computation can be used as additional considerations for the overall system design and / or implementation. For example, very small networks that can be rapidly inferred can be fine-tuned. As another example, a cost term can be added to the receiver complexity, which can force and / or cause the model to remove one or more layers. In some examples, machine learning can be used to learn more complex model priors to obtain greater gains.
[0170] Model prior design can include various attributes. In some examples, the implemented model prior can include a model prior that assigns a high probability. Used to send models that are not being updated, and therefore has a low bit rate: In some cases, model priors can include model priors that provide for The surrounding values are assigned non-zero probabilities, thus allowing for the encoding of different instances of the fine-tuned model in practice. In some cases, model priors may include model priors that can be quantized at the time of speculation and used for entropy encoding / decoding.
[0171] Figure 9 This is a diagram illustrating the encoding and decoding tasks performed by an example neural network compression system 900 fine-tuned using model priors. In this example, encoder 904 receives image 902 and compresses it into a latent space 906(z). Arithmetic encoder 908 can use a latent prior 910 pre-trained on training set D, parameters of decoder 922, and global model parameters 914(θ). D To calculate the quantized model parameters.
[0172] In some examples, the quantized model parameters It can be represented as the global model parameters 914(θ) pre-trained on the training set D. D ) and quantized model parameter updates The sum of . In some examples, the quantized model parameter updates It can be generated by the quantizer (Q) t Based on the global model parameters 914(θ) pre-trained on the training set D D It is generated by the difference between (θ) and the model parameter update (θ).
[0173] Arithmetic encoder 908 can update quantized model parameters Entropy encoding is performed to generate a bitstream 918, which is used to transmit model parameter updates as a signal to a receiver (e.g., an arithmetic decoder 920). The arithmetic encoder 908 can use the quantized model prior 912. Update the quantized model parameters Entropy encoding is performed. In some examples, the arithmetic encoder 908 can implement a continuous model prior (p[δ]) to normalize the model parameter updates (δ), and use a quantized model prior 912. To update the quantized model parameters Perform entropy encoding.
[0174] The arithmetic encoder 908 can also entropy encode the latent space 906(z) to generate a bitstream 916 for signaling to a receiver (e.g., an arithmetic decoder 920). In some examples, the arithmetic encoder 908 can concatenate bitstreams 916 and 918 and send the concatenated bitstream to a receiver (e.g., an arithmetic decoder 920).
[0175] The arithmetic decoder 920 can use quantized model priors 912. Entropy decoding of bitstream 918 is performed to obtain quantized model parameter updates. Arithmetic decoder 920 can use fine-tuned latent prior 910 to entropy decode the latent space 906(z) from bitstream 916. Decoder 922 can use latent space 906 and model parameter updates to generate reconstructed image 924.
[0176] In some examples, the neural network compression system 900 may implement Algorithm 1 below during the encoding process and Algorithm 2 below during the decoding process.
[0177]
[0178]
[0179] In some examples, formula (2) cited above can be used to calculate the RDM loss as follows:
[0180]
[0181] Here, logp(δ) represents the model update M, and β is the trade-off parameter.
[0182] In some examples, the model prior 912 can be designed as a probability density function (PDF). In some cases, the parameter update δ = θ - θ D Modeling can be performed using a Gaussian distribution centered at zero updates. In some cases, the model prior can be defined on the updates as a multivariate zero-centered Gaussian with zero covariance, and a single shared (hyperparameter) σ representing the standard deviation. It can be equivalent to passing through Take the modulus of θ.
[0183] In some cases, when in The following is a quantitative update. When performing entropy encoding and decoding, even zero updates may not be free. In some cases, these initial static costs can be defined as... Since the defined model prior pattern is zero, these initial costs It can be equal to the minimum cost. Minimizing the above formula (2) ensures that after overcoming these static costs, any additional bits spent on model updates can be accompanied by a corresponding improvement in RD performance.
[0184] In some cases, the model prior design can be based on independent Gaussian network priors, where the independence between parameters can be assumed as follows: In the various examples in this article, methods for using a single parameter p are provided. ω (θ (i) Equations and examples are provided. However, other examples may involve equations with multiple parameters.
[0185] In some examples, the prior Gaussian model may include a Gaussian RDAE model centered on the global model, with a variable standard deviation σ as follows:
[0186] In some cases, a trade-off can be made between higher σ (e.g., higher cost (initial cost) for sending an untuned model) and lower σ (e.g., more limited probability quality for fine-tuning the model).
[0187] In some examples, the model prior design can be based on an independent Laplace model prior. In other examples, the Laplace model prior can include a Gaussian RDAE model centered on a global model, with a variable standard deviation σ as follows:
[0188] In some cases, a trade-off can be made between higher σ (e.g., higher cost (initial cost) for sending an untuned model) and lower σ (e.g., more limited probability quality for fine-tuning the model). In some examples, the model may implement sparsity in parameter updates.
[0189] In some cases, model prior design can be based on independent Spike and Slab model priors. In some examples, Spike and Slab model priors may include a Gaussian RDAE model centered on the global model, with variable standard deviations σ for the slab components. slab Variable standard deviation σ for spike components spike <<σ slab And the Spike / Slab ratio α, as follows:
[0190] In some examples, when α is large and there is broad support for fine-tuning models, the cost (initial cost) of sending untuned models may be low.
[0191] In some examples, the model prior and the global AE model can be trained jointly. In some cases, the model prior can be trained end-to-end with the instance RDM-AE model instead of manually setting the parameter ω.
[0192] As mentioned above, fine-tuning the neural compression system at the data points that are sent to and decoded by the receiver can provide advantages and benefits in terms of rate and / or distortion. Figure 10 Figure 1000 shows an example RD-AE model with fine-tuning on the data points (e.g., images or frames, video, audio data, etc.) to be sent to the receiver, and an example rate distortion of an RD-AE model without fine-tuning on the data points to be sent to the receiver.
[0193] In this example, Figure 1000 shows rate distortion 1002 of the RD-AE model that has not been fine-tuned as described herein, and rate distortion 1004 of the fine-tuned RD-AE model. As shown in rate distortion 1002 and rate distortion 1004, the fine-tuned RD-AE model has significantly higher compression (rate distortion) performance than the other RD-AE models.
[0194] The RD-AE model 1004, associated with rate distortion, can be fine-tuned using data points to be compressed and sent to the receiver. In some examples, the compressed AE can be fine-tuned at data point x as follows: In some cases, this can allow for high compression (rate-distortion or R / D) performance on data point x. Fine-tuned prior parameters (e.g., ) and the parameters of the fine-tuned decoder (e.g., This can be shared with the receiver for decoding the bitstream. The receiver can use finely tuned priors (e.g., ) and the parameters of the fine-tuned decoder (e.g., This is used to decode the bitstream generated by the finely tuned RD-AE system.
[0195] Table 1 below shows the example tuning details:
[0196] Table 1
[0197]
[0198] The tuning details in Table 1 are merely illustrative examples. Those skilled in the art will recognize that other examples may include more, fewer, the same, and / or different tuning details.
[0199] Figure 11 This is a flowchart illustrating an example process 1100 for adaptive compression using an instance of a neural network compression system (e.g., neural network compression system 500, neural network compression system 600, neural network compression system 700, neural network compression system 800, neural network compression system 900) adapted to the input data being compressed.
[0200] At box 1102, process 1100 may include: receiving input data via a neural network compression system for compression by the neural network compression system. The neural network compression system can be used on a training dataset (e.g., Figure 6 The neural network compression system is trained on the training dataset 602. In some examples, the neural network compression system may include global model parameters {φ} pre-trained on the training set D (e.g., the training dataset). D ,θ D Input data may include image data, video data, audio data, and / or any other data.
[0201] At box 1104, process 1100 may include: determining a set of updates for the neural network compression system. In some examples, the set of updates may include updated model parameters tuned using the input data (e.g., model parameter updates (θ) and / or quantized model parameter updates). ).
[0202] In some examples, determining a set of updates for a neural network compression system may include: processing (e.g., encoding / decoding) the input data at the neural network compression system; determining one or more losses (e.g., RDM loss) for the neural network compression system based on the processed input data; and tuning the model parameters of the neural network compression system based on the one or more losses. The tuned model parameters may include a set of updates for the neural network compression system.
[0203] In some examples, one or more losses may include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream; a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data; and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream. In some cases, one or more losses may be calculated using Equation 2 above.
[0204] At box 1106, process 1100 may include: generating a first bitstream (e.g., bitstream 510, bitstream 520, bitstream 710, bitstream 810, bitstream 916) comprising a compressed version of the input data using a neural network compression system with latent priors (e.g., latent priors 506, 622, 708, 806, 910). In some examples, the first bitstream may be generated by encoding the input data into a latent space representation and entropy encoding that latent space representation.
[0205] At box 1108, process 1100 may include: using latent priors and model priors (e.g., model priors 714, 816, 912, 1013, 1014, 1015, 1016, 1017, 1018, 1019 ... Generates updated model parameters (e.g., quantized model parameter updates). The compressed version of the second bitstream (e.g., bitstream 510, bitstream 520, bitstream 712, bitstream 811, bitstream 918).
[0206] In some examples, model priors may include independent Gaussian network priors, independent Laplacian network priors, and / or independent Spike and Slab network priors.
[0207] In some examples, generating a second bitstream may include: entropy encoding a latent prior using a model prior via a neural network compression system; and entropy encoding updated model parameters using a model prior via a neural network compression system.
[0208] In some examples, updated model parameters may include neural network parameters such as weights, biases, etc. In some examples, updated model parameters may include one or more updated parameters of the decoder model. Input data can be used to tune one or more updated parameters (e.g., tune / adjust to reduce one or more losses in the input data).
[0209] In some examples, the updated model parameters may include one or more updated parameters of the encoder model. These updated parameters may be tuned using the input data. In some cases, the first bitstream may be generated by a neural network compression system using one or more updated parameters.
[0210] In some examples, generating the second bitstream may include: encoding the input data into a latent spatial representation of the input data using one or more updated parameters via a neural network compression system; and entropy encoding the latent spatial representation into a first bitstream using latent priors via a neural network compression system.
[0211] At box 1110, process 1100 may include outputting a first bitstream and a second bitstream for transmission to a receiver. In some examples, the receiver may include a decoder (e.g., decoder 514, decoder 610, decoder 718, decoder 814, decoder 922). In some examples, the second bitstream may also include a compressed version of the potential prior and a compressed version of the model prior.
[0212] In some cases, process 1100 may further include: generating a concatenated bitstream comprising a first bitstream and a second bitstream; and transmitting the concatenated bitstream to a receiver. For example, a neural network compression system may concatenate the first bitstream and the second bitstream and transmit the concatenated bitstream to a receiver. In other cases, process 1100 may include: transmitting the first bitstream and the second bitstream to a receiver separately.
[0213] In some cases, process 1100 may include: generating model parameters for a neural network compression system based on a training dataset; tuning the model parameters for the neural network compression system using input data; and determining a set of updates based on the differences between the model parameters and the tuned model parameters. In some examples, the model parameters may be tuned based on the input data, the bit size of a compressed version of the input data, the bit size of a set of updates, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0214] In some cases, model parameters can be tuned based on the input data and the ratio (e.g., rate / distortion ratio) of the cost of sending a set of updates to the distortion between the input data and the reconstructed data generated from the compressed version of the input data. In some examples, this cost can be based on the bit size of a set of updates. In some examples, tuning model parameters may include determining to add / include one or more parameters in the tuned model parameters based on the following: including one or more parameters in the tuned model parameters is accompanied by a reduction in the bit size of the compressed version of the input data and / or a reduction in distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0215] In some examples, the receiver may include a decoder (e.g., decoder 514, decoder 610, decoder 718, decoder 814, decoder 922). In some cases, process 1100 may include: receiving data comprising a first bitstream and a second bitstream via an encoder. In some cases, process 1100 may also include: decoding a compressed version of the input data based on a set of updated parameters of the second bitstream via a decoder; and generating a reconstructed version of the input data based on the compressed version of the input data in the first bitstream using a set of updated parameters via a decoder.
[0216] In some examples, process 1100 may include training (e.g., based on a training dataset) a neural network compression system by reducing rate distortion and model rate loss. In some examples, the model rate reflects the length of the bitstream used to send model updates.
[0217] In some examples, process 1100 can implement Algorithm 1 and / or Algorithm 2 as described above.
[0218] Figure 12 This is a flowchart illustrating an example of a process 1200 for compressing one or more images. At block 1202, process 1200 may include: receiving image content for compression. At block 1204, process 1200 may include: updating a neural network compression system by training the neural network compression system to improve the compression performance for compressing images. In some examples, training the neural network compression system may include: training an autoencoder system associated with the neural network compression system.
[0219] At box 1206, process 1200 may include: encoding the received image content into a latent spatial representation using an updated neural network compression system.
[0220] At box 1208, process 1200 may include: generating a compressed version of the encoded image content using an updated neural network compression system based on a probabilistic model and a first code subset. The first code subset may include a portion of the latent space representation. For example, the latent space representation may be divided into a first code subset and one or more additional code subsets.
[0221] At box 1210, the generation process 1200 may include: generating a compressed version of the updated neural network compression system using a probabilistic model and the updated neural network compression system. In some cases, the compressed version of the updated neural network compression system may include quantized parameter updates of the neural network compression system, and may exclude outdated parameters of the neural network compression system.
[0222] At box 1212, process 1200 may include: outputting a compressed version of the updated neural network compression system and a compressed version of the encoded image content for transmission. The compressed version of the updated neural network compression system and the compressed version of the encoded image content may be sent to a receiver (e.g., a decoder) for decoding. The compressed version of the neural network compression system may include the complete model of the neural network compression system or updated model parameters of the neural network compression system.
[0223] In some examples, process 1200 can implement the previously described algorithm 1.
[0224] Figure 13 This is a flowchart illustrating an example of a process 1300 for decompressing one or more images using the techniques described herein. At block 1302, process 1300 includes: receiving a compressed version of an updated neural network compression system (and / or one or more parameters of the neural network compression system) and a compressed version encoding image content. The compressed version of the updated neural network compression system and the compressed version encoding image content can be as described above... Figure 11 Or as described in 12, it is generated and transmitted. In some cases, the encoded image content may include a first subset of codes from which the latent spatial representation of the image content from which the encoded image content is generated.
[0225] In some cases, the compressed version of an updated neural network compression system may include model parameters. In some examples, the model parameters may include updated model parameters and exclude one or more other model parameters.
[0226] At box 1304, process 1300 may include: decompressing the compressed version of the updated neural network compression system into the updated neural network compression system model using a shared probability model.
[0227] At box 1306, the decompression process 1300 may include: decompressing a compressed version of the encoded image content into a latent spatial representation using an updated probabilistic model and an updated neural network compression system model.
[0228] At box 1308, the process 1300 may include: generating reconstructed image content using an updated neural network compression system model and a latent spatial representation.
[0229] At box 1310, process 1300 may include: outputting the reconstructed image content.
[0230] In some examples, process 1300 can implement the above algorithm 2.
[0231] In some examples, the processes described herein (e.g., process 1100, process 1200, process 1300, and / or any other process described herein) may be executed by a computing device or apparatus. In one example, process 1100 and / or 1200 may be executed by... Figure 4 The transmitting device 410 of the system 400 shown is executed. In another example, processes 1100, 1200, and / or 1300 can be performed by, according to... Figure 4 The system 400 shown or Figure 14 The computing device of the computing system 1400 shown is executed.
[0232] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, wearable devices (e.g., VR headsets, AR headsets), AR glasses, connected watches or smartwatches or other wearable devices), server computers, autonomous vehicles or computing devices of autonomous vehicles, robotic devices, televisions, and / or any other computing device with the resources to perform the processes described herein, including processes 1100, 1200, and 1300. In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP) based data or other types of data.
[0233] Components of a computing device can be implemented in a circuit. For example, a component may include electronic circuitry or other electronic hardware and / or may be implemented using such electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof and / or may be implemented using such software to perform the various operations described herein.
[0234] Processes 1100, 1200, and 1300 are shown as logic flowcharts, whose operations represent a series of operations that can be implemented using hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents a computer-executable instruction stored on one or more computer-readable storage media, which, when executed by one or more processors, performs the described operation. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which operations are described should not be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.
[0235] Furthermore, the processes 1100, 1200, and 1300 and / or any other processes described herein can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors via hardware or a combination thereof. As described above, the code can be stored, for example, on a computer-readable or machine-readable storage medium in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0236] Figure 14 This is a diagram illustrating an example of a system used to implement certain aspects of the techniques described herein. Specifically, Figure 14 An example of a computing system 1400 is shown, which can be, for example, any computing device constituting an internal computing system, a remote computing system, a camera, or any component thereof, in which the components of the system communicate with each other using connection 1405. Connection 1405 can be a physical connection using a bus, or a direct connection to processor 1410, such as in a chipset architecture. Connection 1405 can also be a virtual connection, a network connection, or a logical connection.
[0237] In some embodiments, the computing system 1400 is a distributed system, wherein the functions described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functions described for that component. In some embodiments, the components may be physical or virtual devices.
[0238] Example system 1400 includes at least one processing unit (CPU or processor) 1410 and connections 1405 that couple various system components, including system memory 1415 (e.g., read-only memory (ROM) 1420 and random access memory (RAM) 1425), to processor 1410. Computing system 1400 may include a cache 1412 of high-speed memory that is directly connected to, adjacent to, or integrated into processor 1410.
[0239] Processor 1410 may include any general-purpose processor and hardware or software services, such as services 1432, 1434, and 1436 stored in storage device 1430, which are configured to control processor 1410 and dedicated processors in which software instructions are incorporated into the actual processor design. Processor 1410 may essentially be a completely independent computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.
[0240] To enable user interaction, the computing system 1400 includes an input device 1445, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. The computing system 1400 may also include an output device 1435, which can be one or more of multiple output mechanisms. In some cases, a multimodal system allows the user to provide multiple types of input / output to communicate with the computing system 1400. The computing system 1400 may include a communication interface 1440, which typically controls and manages user input and system output.
[0241] The communication interface may use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communications, including transceivers using the following: audio jack / plug, microphone jack / plug, Universal Serial Bus (USB) port / plug, Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs Wireless signal transmission Low-power (BLE) wireless signal transmission Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short-range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), global microwave access interoperability (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.
[0242] The communication interface 1440 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 1400 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the U.S. Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no limitations on operation on any particular hardware arrangement, therefore the basic features described herein can be readily replaced with improved hardware or firmware arrangements (when they are under development).
[0243] Storage device 1630 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing computer-accessible data, such as magnetic tape cassettes, flash memory cards, solid-state storage devices, digital multifunction disks, magnetic tape cassettes, floppy disks, hard disks, magnetic tape, magnetic stripes / tapes, any other magnetic storage media, flash memory, memory storage, any other solid-state storage, CD-ROM, rewritable CD, DVD, Blu-ray Disc, holographic disc, another optical medium, Secure Digital (SD) card, microSD card, memory Cards, smart card chips, EMV chips, SIM cards, micro / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), other memory chips or cassette memories, and / or combinations thereof.
[0244] Storage device 1430 may include software services, servers, services, etc., which cause the system to perform a certain function when processor 1410 executes code defining such software. In some embodiments, hardware services performing a particular function may include software components stored in a computer-readable medium and connected to necessary hardware components (e.g., processor 1410, connection 1405, output device 1435, etc.) to perform that function. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored, and does not include carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as optical discs (CDs) or digital versatile optical discs (DVDs), flash memory, memory, or storage devices. Code and / or machine-executable instructions may be stored on a computer-readable medium, which may represent processes, functions, subroutines, programs, routines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be connected to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameter data, etc., can be passed, forwarded, or sent via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.
[0245] In some embodiments, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0246] Specific details have been provided in the foregoing description to provide a thorough understanding of the embodiments and examples presented herein. However, it will be understood by those skilled in the art that these embodiments can be practiced without using these specific details. For clarity of explanation, in some instances, the techniques described herein may be presented as comprising individual functional blocks, including devices, device components, steps or routines in a software-embodied method, or a combination of hardware and software. Other components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary details. For example, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details to avoid obscuring the embodiments.
[0247] Various embodiments can be described as processes or methods, depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. While a flowchart can describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of these operations can be rearranged. A process terminates upon completion of its operations, but a process may have additional steps not shown in the diagram. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0248] The processes and methods according to the examples above can be implemented using computer-executable instructions stored in or available from a computer-readable medium. For example, such instructions may include instructions and data that cause a general-purpose computer, special-purpose computer, or processing device, or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. Some portions of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, used information, and / or information created during the methods according to the examples include hard disks or optical disks, flash memory, USB devices equipped with non-volatile memory, network storage devices, etc.
[0249] Devices implementing the processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take on any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, and so on. The functionality described herein may also be embodied in peripheral devices or add-in cards. As a further example, such functionality may also be implemented on a circuit board between different chips or different processes executed in a single device.
[0250] Instructions, the medium for conveying these instructions, the computing resources for executing them, and other structures for supporting such computing resources are example units for providing the functionality described in this disclosure.
[0251] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that these inventive concepts may be embodied and employed differently in other ways, and the appended claims are intended to be construed as including such variations beyond those limited by the prior art. Various features and aspects of the above-described applications may be used individually or in combination. Furthermore, embodiments may be used in any number of settings and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, this description and the accompanying drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods have been described in a particular order. It should be recognized that, in alternative embodiments, these methods may be performed in a different order than that described.
[0252] Those skilled in the art will understand that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols without departing from the scope of this specification.
[0253] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform those operations, by programming programmable electronic circuits (such as microprocessors or other suitable electronic circuits) to perform those operations, or any combination thereof.
[0254] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0255] The use of "at least one" and / or "one or more" in the language of claims, or other languages, indicates that one or more members of the set (in any combination) satisfy the claim. For example, the statement language referring to "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the statement language referring to "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The use of "at least one" and / or "one or more" in the language of claims does not limit the set to items listed in that set. For example, the statement language referring to "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0256] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the examples disclosed herein can all be implemented as electronic hardware, computer software, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the illustrative components, blocks, modules, circuits, and steps above have been generally described in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in alternative ways for each specific application, but such implementation decisions should not be construed as a departure from the scope of this application.
[0257] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handheld devices, or integrated circuit devices with a wide range of uses, including applications in wireless communication handheld devices and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as a discrete but interoperable logic device. If implemented in software, these techniques can be implemented at least in part by a computationally readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (e.g., synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, etc. Additionally or alternatively, these technologies may be implemented at least in part by a computationally readable communication medium that carries or transmits program code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer, for example, to propagate signals or waveforms.
[0258] The program code can be executed by a processor, which may include one or more processors such as a digital signal processor (DSP), a general-purpose microprocessor, an application-specific integrated circuit (ASIC), a field-programmable array (FPGA), or other equivalent integrated or discrete logic circuitry. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, or the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, a combination of one or more microprocessors with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.
[0259] Illustrative examples of this disclosure include:
[0260] Aspect 1: An apparatus comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: receive input data via a neural network compression system for compression via the neural network compression system; determine a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generate a first bitstream comprising a compressed version of the input data via the neural network compression system using latent priors; generate a second bitstream comprising a compressed version of the updated model parameters via the neural network compression system using the latent priors and model priors; and output the first bitstream and the second bitstream for transmission to a receiver.
[0261] Aspect 2: The apparatus according to aspect 1, wherein the second bitstream further includes a compressed version of the potential prior and a compressed version of the model prior.
[0262] Aspect 3: The apparatus according to any one of aspects 1 to 2, wherein the one or more processors are configured to: generate a concatenated bit stream comprising the first bit stream and the second bit stream; and transmit the concatenated bit stream to the receiver.
[0263] Aspect 4: The apparatus according to any one of aspects 1 to 3, wherein, in order to generate the second bitstream, the one or more processors are configured to: entropy encode the latent prior using the model prior through the neural network compression system; and entropy encode the updated model parameters using the model prior through the neural network compression system.
[0264] Aspect 5: The apparatus according to any of aspects 1 to 4, wherein the updated model parameters include one or more updated parameters of the decoder model, the one or more updated parameters being tuned using the input data.
[0265] Aspect 6: The apparatus according to any of aspects 1 to 5, wherein the updated model parameters include one or more updated parameters of the encoder model, the one or more updated parameters being tuned using the input data, wherein the first bitstream is generated by the neural network compression system using the one or more updated parameters.
[0266] Aspect 7: The apparatus according to aspect 6, wherein, in order to generate the second bitstream, the one or more processors are configured to: encode the input data into a latent spatial representation of the input data using the one or more updated parameters via the neural network compression system; and entropy encode the latent spatial representation into the first bitstream using the latent prior via the neural network compression system.
[0267] Aspect 8: An apparatus according to any one of aspects 1 to 7, wherein the one or more processors are configured to: generate model parameters of the neural network compression system based on a training dataset used to train the neural network compression system; tune the model parameters of the neural network compression system using the input data; and determine the set of updates based on the difference between the model parameters and the tuned model parameters.
[0268] Aspect 9: The apparatus according to aspect 8, wherein the model parameters are tuned based on the input data, the bit size of the compressed version of the input data, the set of updated bit sizes, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0269] Aspect 10: The apparatus according to aspect 8, wherein the model parameters are tuned based on the input data and the ratio of the cost of sending the set of updates to the distortion between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
[0270] Aspect 11: The apparatus according to aspect 8, wherein, in order to tune the model parameters, the one or more processors are configured to: determine to include one or more parameters in the tuned model parameters based on the following: including the one or more parameters in the tuned model parameters is accompanied by a reduction of at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0271] Aspect 12: An apparatus according to any of aspects 1 to 11, wherein, in order to determine the set of updates for the neural network compression system, the one or more processors are configured to: process the input data at the neural network compression system; determine one or more losses of the neural network compression system based on the processed input data; and tune model parameters of the neural network compression system based on the one or more losses, the tuned model parameters including the set of updates for the neural network compression system.
[0272] Aspect 13: The apparatus according to aspect 12, wherein the one or more losses include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream, a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
[0273] Aspect 14: An apparatus according to any one of aspects 1 to 13, wherein the receiver includes an encoder, and wherein the one or more processors are configured to: receive data including a first bitstream and a second bitstream via the encoder; decode the compressed version of the updated model parameters based on the second bitstream via the decoder; and generate a reconstructed version of the input data based on the compressed version of the input data in the first bitstream via the decoder using the set of updated parameters.
[0274] Aspect 15: An apparatus according to any of aspects 1 to 14, wherein the one or more processors are configured to train the neural network compression system by reducing rate distortion and model rate loss, wherein the model rate reflects the length of the bit stream used to send model updates.
[0275] Aspect 16: The apparatus according to any of aspects 1 to 15, wherein the model prior includes at least one of the following: independent Gaussian network prior, independent Laplace network prior, and independent Spike and Slab network prior.
[0276] Aspect 17: The apparatus according to any one of aspects 1 to 16, wherein the apparatus includes a mobile device.
[0277] Aspect 18: The apparatus according to any one of aspects 1 to 17 further includes a camera configured to capture the input data.
[0278] Aspect 19: A method comprising: receiving input data via a neural network compression system for compression via the neural network compression system; determining a set of updates for the neural network compression system, the set of updates including updated model parameters tuned using the input data; generating a first bitstream comprising a compressed version of the input data via the neural network compression system using a latent prior; generating a second bitstream comprising a compressed version of the updated model parameters via the neural network compression system using the latent prior and a model prior; and outputting the first bitstream and the second bitstream for transmission to a receiver.
[0279] Aspect 20: The method according to aspect 19, wherein the second bitstream further includes a compressed version of the potential prior and a compressed version of the model prior.
[0280] Aspect 21: The method according to any of aspects 19 to 20, wherein the one or more processors are configured to: generate a concatenated bit stream comprising the first bit stream and the second bit stream; and transmit the concatenated bit stream to the receiver.
[0281] Aspect 22: The method according to any of aspects 19 to 21, wherein generating the second bitstream comprises: entropy encoding the latent prior using the model prior through the neural network compression system; and entropy encoding the updated model parameters using the model prior through the neural network compression system.
[0282] Aspect 23: The method according to any of aspects 19 to 22, wherein the updated model parameters include one or more updated parameters of the decoder model, the one or more updated parameters being tuned using the input data.
[0283] Aspect 24: The method according to any of aspects 19 to 23, wherein the updated model parameters include one or more updated parameters of the encoder model, the one or more updated parameters being tuned using the input data, wherein the first bitstream is generated by the neural network compression system using the one or more updated parameters.
[0284] Aspect 25: According to the method of aspect 24, generating the second bitstream includes: encoding the input data into a latent spatial representation of the input data using the one or more updated parameters via the neural network compression system; and entropy encoding the latent spatial representation into the first bitstream using the latent prior via the neural network compression system.
[0285] Aspect 26: The method according to any of aspects 19 to 25, wherein the one or more processors are configured to: generate model parameters of the neural network compression system based on a training dataset used to train the neural network compression system; tune the model parameters of the neural network compression system using the input data; and determine the set of updates based on the difference between the model parameters and the tuned model parameters.
[0286] Aspect 27: According to the method of aspect 26, wherein the model parameters are tuned based on the input data, the bit size of the compressed version of the input data, the set of updated bit sizes, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0287] Aspect 28: According to the method of aspect 26, wherein the model parameters are tuned based on the input data and the ratio of the cost of sending the set of updates to the distortion between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
[0288] Aspect 29: According to the method of aspect 26, wherein tuning the model parameters includes: determining to include one or more parameters in the tuned model parameters based on the following: including the one or more parameters in the tuned model parameters is accompanied by a reduction of at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
[0289] Aspect 30: The method according to any of aspects 19 to 29, wherein determining the set of updates for the neural network compression system comprises: processing the input data at the neural network compression system; determining one or more losses of the neural network compression system based on the processed input data; and tuning model parameters of the neural network compression system based on the one or more losses, the tuned model parameters including the set of updates for the neural network compression system.
[0290] Aspect 31: The method according to aspect 30, wherein the one or more losses include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream, a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
[0291] Aspect 32: The method according to any of aspects 19 to 31, wherein the receiver includes an encoder, and wherein the one or more processors are configured to: receive data including a first bitstream and a second bitstream via the encoder; decode the compressed version of the updated model parameters based on the second bitstream via the decoder; and generate a reconstructed version of the input data based on the compressed version of the input data in the first bitstream via the decoder using the set of updated parameters.
[0292] Aspect 33: The method according to any of aspects 19 to 32, wherein the one or more processors are configured to: train the neural network compression system by reducing rate distortion and model rate loss, wherein the model rate reflects the length of the bit stream used to send model updates.
[0293] Aspect 34: The method according to any of aspects 19 to 33, wherein the model prior includes at least one of the following: independent Gaussian network prior, independent Laplace network prior, and independent Spike and Slab network prior.
[0294] Aspect 35: A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of aspects 19 to 33.
[0295] Aspect 36: An apparatus comprising a unit for performing the method according to any one of aspects 19 to 33.
Claims
1. An apparatus for data processing, comprising: Memory; as well as One or more processors coupled to the memory, the one or more processors being configured to: Input data is received through a neural network compression system for compression by the neural network compression system. Determine a set of updates for the neural network compression system, the set of updates including one or more updated parameters of the encoder model, the one or more updated parameters being tuned using the input data; The neural network compression system uses latent priors and one or more updated parameters to generate a first bitstream of a compressed version of the input data. The neural network compression system uses the latent priors and model priors to generate a second bitstream that includes one or more updated parameters; as well as The first bitstream and the second bitstream are output for transmission to the receiver. In order to generate the second bitstream, the one or more processors are configured to: The input data is encoded into a latent spatial representation of the input data using the one or more updated parameters via the neural network compression system; and The neural network compression system uses the latent prior to encode the latent spatial representation entropy into the first bit stream.
2. The apparatus according to claim 1, wherein, The second bitstream also includes a compressed version of the potential prior and a compressed version of the model prior.
3. The apparatus according to claim 1, wherein, The one or more processors are configured to: Generate a concatenated bitstream comprising the first bitstream and the second bitstream; and The concatenated bit stream is sent to the receiver.
4. The apparatus according to claim 1, wherein, In order to generate the second bitstream, the one or more processors are configured to: The neural network compression system uses the model prior to entropy encode the latent prior; and The neural network compression system uses the model prior to entropy encode one or more updated parameters.
5. The apparatus according to claim 1, wherein, The one or more processors are configured to: The model parameters of the neural network compression system are generated based on the training dataset used to train the neural network compression system. The input data is used to tune the model parameters of the neural network compression system; as well as The set of updates is determined based on the difference between the model parameters and the tuned model parameters.
6. The apparatus according to claim 5, wherein, The model parameters are tuned based on the input data, the bit size of the compressed version of the input data, the set of updated bit sizes, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
7. The apparatus according to claim 5, wherein, The model parameters are tuned based on the input data and the cost of sending the set of updates relative to the distortion ratio between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
8. The apparatus according to claim 5, wherein, In order to tune the model parameters, the one or more processors are configured to: The inclusion of one or more parameters in the tuned model parameters is determined based on the following: including the one or more parameters in the tuned model parameters is accompanied by a reduction in at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
9. The apparatus according to claim 1, wherein, In order to determine the set of updates for the neural network compression system, the one or more processors are configured to: Process the input data at the neural network compression system; One or more losses of the neural network compression system are determined based on the processed input data; as well as The model parameters of the neural network compression system are tuned based on the one or more losses, and the tuned model parameters include the set of updates for the neural network compression system.
10. The apparatus according to claim 9, wherein, The one or more losses include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream, a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
11. The apparatus according to claim 1, wherein, The one or more processors are configured to: The neural network compression system is trained by reducing rate distortion and model rate loss, where the model rate reflects the length of the bit stream used to send model updates.
12. The apparatus according to claim 1, wherein, The model priors include at least one of the following: independent Gaussian network priors, independent Laplace network priors, and independent Spike and Slab network priors.
13. The apparatus according to claim 1, wherein, The device includes a mobile device.
14. The apparatus of claim 1, further comprising a camera configured to capture the input data.
15. A method for data processing, comprising: Input data is received through a neural network compression system for compression by the neural network compression system. Determine a set of updates for the neural network compression system, the set of updates including one or more updated parameters of the encoder model, the one or more updated parameters being tuned using the input data; The neural network compression system uses latent priors and one or more updated parameters to generate a first bitstream that includes a compressed version of the input data. The neural network compression system uses the latent priors and model priors to generate a second bitstream that includes one or more updated parameters; as well as The first bitstream and the second bitstream are output for transmission to the receiver. Generating the second bit stream includes: The input data is encoded into a latent spatial representation of the input data using the one or more updated parameters via the neural network compression system; and The neural network compression system uses the latent prior to encode the latent spatial representation entropy into the first bit stream.
16. The method according to claim 15, wherein, The second bitstream also includes a compressed version of the potential prior and a compressed version of the model prior.
17. The method according to claim 15, wherein, The one or more processors are configured to: Generate a concatenated bitstream comprising the first bitstream and the second bitstream; and The concatenated bit stream is sent to the receiver.
18. The method according to claim 15, wherein, Generating the second bitstream includes: The neural network compression system uses the model prior to entropy encode the latent prior; and The neural network compression system uses the model prior to entropy encode one or more updated parameters.
19. The method according to claim 15, wherein, The one or more processors are configured to: The model parameters of the neural network compression system are generated based on the training dataset used to train the neural network compression system. The input data is used to tune the model parameters of the neural network compression system; as well as The set of updates is determined based on the difference between the model parameters and the tuned model parameters.
20. The method according to claim 19, wherein, The model parameters are tuned based on the input data and at least one of the following: the bit size of the compressed version of the input data, the bit size of the set of updates, the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and the ratio of the cost of sending the set of updates to the distortion between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
21. The method according to claim 19, wherein, Tuning the model parameters includes determining to include one or more parameters in the tuned model parameters based on the following: including the one or more parameters in the tuned model parameters is accompanied by a reduction in at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
22. The method according to claim 15, wherein, Determining the set of updates for the neural network compression system includes: Process the input data at the neural network compression system; One or more losses of the neural network compression system are determined based on the processed input data; and The model parameters of the neural network compression system are tuned based on the one or more losses, and the tuned model parameters include the set of updates used for the neural network compression system. The one or more losses include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream; a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data; and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
23. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to: Input data is received through a neural network compression system for compression by the neural network compression system. Determine a set of updates for the neural network compression system, the set of updates including one or more updated parameters of the encoder model, the one or more updated parameters being tuned using the input data; The neural network compression system uses latent priors and one or more updated parameters to generate a first bitstream that includes a compressed version of the input data. The neural network compression system uses the latent priors and model priors to generate a second bitstream that includes one or more updated parameters; as well as The first bitstream and the second bitstream are output for transmission to the receiver. In order to generate the second bit stream, the instruction further causes the one or more processors to: The input data is encoded into a latent spatial representation of the input data using the one or more updated parameters via the neural network compression system; and The neural network compression system uses the latent prior to encode the latent spatial representation entropy into the first bit stream.
24. The non-transitory computer-readable medium according to claim 23, wherein, The second bitstream also includes a compressed version of the potential prior and a compressed version of the model prior.
25. The non-transitory computer-readable medium according to claim 23, wherein, The instructions also cause the one or more processors to: Generate a concatenated bitstream comprising the first bitstream and the second bitstream; and The concatenated bit stream is sent to the receiver.
26. The non-transitory computer-readable medium according to claim 23, wherein, In order to generate the second bitstream, the instructions also cause the one or more processors to: The neural network compression system uses the model prior to entropy encode the latent prior; and The neural network compression system uses the model prior to entropy encode one or more updated parameters.
27. The non-transitory computer-readable medium according to claim 23, wherein, The instructions also cause the one or more processors to: The model parameters of the neural network compression system are generated based on the training dataset used to train the neural network compression system. The input data is used to tune the model parameters of the neural network compression system; as well as The set of updates is determined based on the difference between the model parameters and the tuned model parameters.
28. The non-transitory computer-readable medium according to claim 27, wherein, The model parameters are tuned based on the input data, the bit size of the compressed version of the input data, the set of updated bit sizes, and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
29. The non-transitory computer-readable medium according to claim 27, wherein, The model parameters are tuned based on the input data and the cost of sending the set of updates relative to the distortion ratio between the input data and the reconstructed data generated from the compressed version of the input data, the cost being based on the bit size of the set of updates.
30. The non-transitory computer-readable medium according to claim 27, wherein, In order to tune the model parameters, the instructions also cause the one or more processors to determine whether to include one or more parameters in the tuned model parameters based on the following: Including the one or more parameters in the tuned model parameters is accompanied by a reduction in at least one of the following: the bit size of the compressed version of the input data and the distortion between the input data and the reconstructed data generated from the compressed version of the input data.
31. The non-transitory computer-readable medium according to claim 23, wherein, In order to determine the set of updates for the neural network compression system, the instructions also cause the one or more processors to: Process the input data at the neural network compression system; One or more losses of the neural network compression system are determined based on the processed input data; as well as The model parameters of the neural network compression system are tuned based on the one or more losses, and the tuned model parameters include the set of updates for the neural network compression system.
32. The non-transitory computer-readable medium according to claim 31, wherein, The one or more losses include: a rate loss associated with the rate at which the compressed version of the input data is transmitted based on the size of the first bitstream, a distortion loss associated with the distortion between the input data and the reconstructed data generated from the compressed version of the input data, and a model rate loss associated with the rate at which the compressed version of the updated model parameters is transmitted based on the size of the second bitstream.
33. The non-transitory computer-readable medium according to claim 23, wherein, The instructions also cause the one or more processors to: The neural network compression system is trained by reducing rate distortion and model rate loss, where the model rate reflects the length of the bit stream used to send model updates.