Method and apparatus for trimming selected parameter set in depth coding system
By selecting and optimizing input specific weight subsets, the problem of fine-tuning of a single image deep neural network model in image compression is solved, and efficient image decoding and compression effects are achieved.
Patent Information
- Application Number
- CN202380074127.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-21
- Filing Date
- 2023-10-06
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to effectively fine-tune the deep neural network model of a single image in image compression, resulting in the possible suboptimal reasoning results, and traditional methods have shortcomings in peak signal-to-noise ratios in single image compression.
By selecting and optimizing input-specific weight subsets, the process that can be reproduced by the inference engine is used to implicitly select the weight subset optimized for a specific input, avoiding the transmission of the weighted identifiers, thereby achieving efficient decoding of the image.
This method reduces the bitstream size in image compression, improves the efficiency and quality of image decoding, and is suitable for compression of a single image without increasing the bit length.
Smart Images

Figure CN120112915A_ABST
Abstract
Description
[0001] This application claims priority to European application number 22306599.6 filed on October 21, 2022, which is incorporated herein by reference in its entirety. Technical Field
[0002] At least one of the present embodiments relates generally to neural networks, and more particularly to fine-tuning a selected set of parameters of a deep neural network. Background Art
[0003] Deep neural networks consist of multiple neural layers such as convolutional layers. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called a bias, and then applies a nonlinear function to the resulting value. The shape of the tensor (and other properties) and the type of nonlinear function are called the "architecture" of the network. The values of the tensors and biases are referred to as "weights" below. The weights and, if applicable, the parameters of the nonlinear function are called "parameters". The architecture and parameters define a "model".
[0004] The model M can be trained on a database of images D to learn its weights. In supervised learning, this database consists of input / output pairs (i, o), and the model M is a function that tries to predict the output given the input: The weights are optimized to minimize the training loss, such as Where d measures the difference between the actual output and the predicted output. As an example, d can be the squared error or the Euclidean distance. The loss function can also contain additional terms, such as regularization terms. The value of the parameter is denoted by θ below. Using the trained model is called inference.
[0005] Training is successful when the resulting value of the loss is small. A trained model performs well on average for all inputs, but may be suboptimal for any single input. In some applications such as compression, inference is part of a two-step system where the input is first prepared or viewed by an optimizer (the encoder in compression), and then in a second step, typically in another device, the input is processed by an inference engine (within the decoder in compression). In such a system, the inference results can be improved by fine-tuning (in other words, by retraining) the weights of the model for each input in the optimizer individually. By retraining M specifically for that input, transmitting the weight update δ to the inference engine in addition to the input, and adding δ to θ before inference, the reconstructed output to better match the desired result. The retraining loss for fine-tuning can be:
[0006]
[0007] Image and video compression is a fundamental task in image processing, which has become crucial in an era of massive popularity and increase in video streaming. Thanks to decades of great efforts by the community, traditional methods have reached state-of-the-art rate / distortion performance and dominate current industrial codec solutions. End-to-end trainable deep models have recently emerged as an alternative and achieved promising results. Now, they beat the best traditional compression methods (VVC, Versatile Video Coding) even in terms of peak signal-to-noise ratio for single image compression. Summary of the invention
[0008] In at least one embodiment, a deep neural network based encoding system for an image determines update parameters for a deep neural network model used to decode the image. These parameters are determined by an encoder and provided to a decoder to update the decoder's model before decoding the image. This provides structural sparsity by fine-tuning only some parameters of the neural decoder. The update is performed on a set of parameters selected based on an embedding representing the encoded image, so that information related to the selection of parameters to be updated does not need to be transmitted. A more general optimizer / inference engine that implements data transformations and application to sound upsampling is also described.
[0009] According to a first aspect, a method includes: obtaining input data; selecting a subset of parameters for fine-tuning a model of a neural network; determining a parameter update for the selected subset of parameters based on a loss function; and packaging the input data and the parameter update.
[0010] According to a second aspect, a method includes: obtaining input data and parameter updates for a selected parameter subset; selecting the parameter subset for fine-tuning a model of a neural network based on the parameter optimization and the input data; updating the model of the neural network based on the parameter updates for the selected parameter subset; and determining output data by processing the input data using the updated neural network.
[0011] According to a third aspect, an apparatus includes a processor configured to: obtain input data; select a subset of parameters for fine-tuning a model of a neural network; determine a parameter update for the selected subset of parameters based on a loss function; and package the input data and the parameter update.
[0012] According to a fourth aspect, an apparatus includes a processor configured to: obtain input data and parameter updates for a selected parameter subset; select the parameter subset for fine-tuning a model of a neural network based on the parameter optimization and the input data; update the model of the neural network based on the parameter updates for the selected parameter subset; and determine output data by processing the input data using the updated neural network.
[0013] In a first variation of the first and third aspects applicable to encoding an image, the input data is an image, the first neural network is used for decoding and the second neural network is used for encoding, the second neural network is updated using parameter updates for the selected parameter subset, the method further comprising: determining an embedding by encoding the image using the second neural network; quantizing the embedding; and performing the selection of the parameter subset based on the quantized embedding.
[0014] In a second variation of the first and third aspects applicable to compressed sound, the input data is an audio signal, and the method further comprises: compressing the audio signal; decompressing the compressed audio signal; performing selection of a parameter subset based on the decompressed compressed audio signal; and packaging the compressed audio signal and the parameter update.
[0015] In a first variant of the second and fourth aspects suitable for decoding an image, the selection of the parameter subset is further based on the obtained quantized embedding.
[0016] In a second variant of the second and fourth aspects suitable for decoding sound, the selection of the parameter subset is further based on the obtained decompressed audio signal.
[0017] According to a fifth aspect of at least one embodiment, a computer program is presented, the computer program comprising program code instructions executable by a processor, the computer program implementing at least the steps of the method according to the first aspect or the second aspect when executed on the processor.
[0018] According to a sixth aspect of at least one embodiment, a non-transitory computer-readable medium is presented, which includes program code instructions executable by a processor, which when executed on the processor implements the steps of the method according to at least the first aspect or the second aspect.
[0019] In variations of the first, second, third and fourth embodiments, the parameters are selected from a set comprising biases, weights, parameters of a nonlinear function of the model, a subset of layers of the model, a specific layer of the model, biases of a specific layer of the model, and a subset of neurons of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 An example of an end-to-end neural network-based compression system for encoding images using a deep neural network is shown.
[0021] Figure 2 A process of an optimizer in accordance with at least one embodiment of a general data transformation system is shown.
[0022] Figure 3A process of an inference engine in accordance with at least one embodiment of a general data transformation system is shown.
[0023] Figure 4 An architectural diagram of an optimizer and inference engine according to at least one embodiment is shown.
[0024] Figure 5A A process of an encoder in the context of an end-to-end image compression system in accordance with at least one embodiment is shown.
[0025] Figure 5B A process of a decoder in the context of an end-to-end image compression system in accordance with at least one embodiment is shown.
[0026] Figure 6 The architecture of an encoder in the context of an end-to-end image compression system according to at least one embodiment is shown.
[0027] Figure 7 The architecture of a decoder in the context of an end-to-end image compression system in accordance with at least one embodiment is shown.
[0028] Figure 8 An example is shown of size information of a bitstream generated according to at least one embodiment compared to a bitstream of the same input generated without any of the presented embodiments.
[0029] Fig. 9 An example of applying an optimizer and an inference engine in the context of sound enhancement is shown in accordance with at least one embodiment.
[0030] Fig.10 A block diagram is shown of an example of a system in which various aspects and embodiments are implemented.
[0031] Fig.11 A process of an image encoder according to at least one embodiment is shown.
[0032] Fig.12 A process for an image decoder according to at least one embodiment is shown. DETAILED DESCRIPTION
[0033] Figure 1An example of an end-to-end neural network based compression system 100 for encoding an image using a deep neural network is shown. An input image x to be compressed is first processed by a device 110 comprising a deep neural network encoder (hereinafter referred to as a deep encoder or encoder). The output y of the encoder is called the embedding of the image. The embedding is converted into a bitstream 120 by undergoing a quantizer Q and then an entropy encoder AE. The resulting bitstream thus comprises an encoded quantized embedding of the input image. The bitstream is provided to a device 130 comprising a deep neural network decoder 130 (hereinafter referred to as a deep decoder or decoder). The bitstream is decoded by undergoing an arithmetic decoder AD to reconstruct the quantized embedding The reconstructed quantized embedding can be processed by the deep decoder to obtain the decompressed image
[0034] Deep encoders and decoders consist of multiple neural layers. Usually, encoders and decoders are fixed, based on a predetermined model that should be known at encoding and decoding time. The encoder and decoder models are trained, for example, simultaneously so that they are compatible. Together they are sometimes referred to as "autoencoders", i.e., models that encode an input and then reconstruct that input. The architecture of the decoder is usually almost the opposite of the encoder, but some layers or their order may be slightly different. The set of parameters of the decoder is denoted by Ω below.
[0035] Many end-to-end architectures have been proposed. They may be more Figure 1 The architectures shown in are more complex, but they all retain a deep encoder and decoder. State-of-the-art models can rival traditional codecs such as Versatile Video Coding (VVC) in terms of rate / distortion tradeoff.
[0036] The model M must be trained on a massive image database D to learn the weights of the encoder and decoder. Typically, the weights are optimized to minimize the rate / distortion training loss, for example, expressed as:
[0037]
[0038] where p M represents the probability of quantized embedding according to M (hence, this term is a theoretical lower bound on the bitstream size of the encoded quantized embedding), represents a measure of the distortion between the original image and the reconstructed image (e.g., mean squared error, multi-scale structural similarity index metric (MS-SSIM), information weighted structural similarity index metric (IWSSIM), video multi-method assessment fusion (VMAF), visual information fidelity (VIF), peak signal-to-noise ratio human visual system modified (PSNR-HVS-M), normalized Laplacian pyramid distance (NLPD), or feature similarity index metric (FSIM)), and λ represents a parameter that controls the trade-off between the rate (r) and distortion (d) terms.
[0039] Typically, the architecture is trained multiple times with different values of λ to produce a set of models {M i}. Typically, different architectures produce models with different r / d points. To compare these architectures, the r / d points of each architecture are interpolated to obtain a function d(r) for each architecture that provides a distortion estimate for any value of rate.
[0040] Figure 1 The deep decoder proposed in can decode any type of image. In other words, it performs well on average for all images, but may be suboptimal for any single image. The rate / distortion tradeoff for a single video can be improved by retraining the decoder specifically for that video and transmitting the decoder's weight updates δ in addition to the quantized embeddings of the video's intra frames. Before decoding the quantized embeddings, δ is added to θ. This technique is denoted fine-tuning. The weight updates δ are determined by a fine-tuning algorithm that minimizes a loss function, which can be, for example:
[0041]
[0042] where p Δ (.) represents the probability density of weight update, represents the image reconstructed by the decoder whose weights have been updated by δ, and β represents the trade-off between the two losses.
[0043] However, this approach does not achieve rate / distortion improvements for a single image due to the increased code size caused by including weight updates. In an example solution, an additional term can be added to the loss to enforce a global sparsity constraint on δ, so that many weight updates have the same value (0) to make the encoding more efficient.
[0044] Current methods of fine-tuning the decoder using a global sparsity constraint lead to improved performance in terms of rate / distortion for encoding video. However, due to the increased code size caused by including weight updates, this approach is not suitable for a single image even with a global sparsity constraint.
[0045] A second solution proposes to fine-tune the decoder for a single image by updating a fixed subset of weights for all images or a subset of weights specific to each image. In the latter case, the updated weights must be identified in the bitstream.
[0046] Previous approaches (e.g., overfitting specific weights) necessarily suffer from a suboptimality problem. When the same subset of weights is optimized for each input, the weight selection is not optimal for each input. When a subset of weights is selected specifically for each input, these weights must be identified in the bitstream, thus increasing the bit length. To limit this additional cost, weights are usually selected on a block-by-block basis (e.g., layer-by-layer basis), thus also limiting the reduction in distortion.
[0047] The embodiments described below are designed with the above in mind and are based on a new fine-tuning process that proposes implicitly selecting a subset of weights optimized for a specific input using a process that can be reproduced by an inference engine. In other words, the embodiments are based on selecting and optimizing a subset of weights specific to an input without transmitting the identifiers (or locations) of these weights. Therefore, this weight selection is more suitable for a specific input (such as an image or frame, GoP, tile, or other input), but does not increase the bit length because the location of the selected weights does not need to be transmitted. Updates still need to be transmitted. The cost is an increase in computational cost in the inference engine.
[0048] One embodiment is directed to a general data conversion system including fine-tuning capabilities. In an end-to-end compressed embodiment, the inference engine is an end-to-end decoder. In at least one embodiment, the principles are applied to a video compression system including a video encoder and a video decoder and allow the size of the encoded video bitstream generated by the encoder to be reduced because it does not include any information identifying weights to be updated.
[0049] According to at least one embodiment, a general data conversion system based on selecting and optimizing input specific weight subsets without transmitting identifiers of these weights can be implemented as follows: Figure 2 and Figure 3 The optimizer and inference engine shown are implemented.
[0050] Figure 2 200 of an optimizer according to at least one embodiment of a general data conversion system is shown. The optimizer contains a model M that is the same as the model of the inference engine. The model implements any function that maps an input domain to an output domain. The process 200 is, for example, Fig.10The device 1000 is implemented, and more specifically, by the processor 1010 of such a device. In step 210, the optimizer obtains data representing input i and optionally data representing target output o (i.e., the expected output of the inference engine). If there is a target output, the output is the target of the model M, and the optimizer will attempt to optimize the parameters of the model so that the output is closer to the target output. If there is no target output, the optimizer can attempt to optimize the metric on the output without considering the target. These will be referred to as "no reference" metrics. For example, if the output domain of the model is an image, a metric such as BRISQUE (blind / no reference image space quality assessor) can be used, if the domain is related to a probability distribution over labels, then the probability of the label can be maximized, or if the output domain is an audio signal, the optimization can attempt to limit the clipping or Gaussianity of the signal.
[0051] In step 220, the optimizer selects a weight subset of fixed size s, based on the input i
[0052] ω * =f(M,i).
[0053] This computation will be reproduced by the inference engine. Therefore, this step does not depend on the target output (as the inference engine cannot access it), but only on the quantized embedding.
[0054] The ideal choice might be a solution to the following optimization problem:
[0055]
[0056] where δ ω represents the update corresponding to the parameter ω. However, since it depends on the target output o, the above problem formulation cannot be used by the inference engine, so in at least one embodiment, another method is proposed to replace it. In the following, the selection of the subset ω is described. * Three different methods.
[0057] The idea of the first approach is to select the weights that have the greatest impact on the target output of the model M after modification. This property can be estimated by the gradient of the target output of the model. This can be easily calculated by using the back-propagation algorithm (commonly used for training models). Therefore, in this first approach, a subset of weights is calculated as follows:
[0058]
[0059] in represents the gradient with respect to the parameter Ω.
[0060] The second approach proposes to use a machine learning algorithm to directly infer the subset from the input. A possible option is to use a supervised learning algorithm. In this case, we use an optimal subset ω consisting of pairs of inputs i and that input * (i) is used to train the second machine learning model N. The second element ω * (i) is the output of the model N. This model N can then be used as a function f(M, i) in the inference engine to determine the subset of parameters to update based on the quantized embeddings. This function may be known to both the encoder and the decoder, so that the location / identifier of the updated weights does not need to be transmitted.
[0061] A third approach is to use a reinforcement learning algorithm. Such an algorithm can, for example, gradually build up ω by adding or removing elements of Ω, fine-tuning the updates of these weights, and using the resulting r / d tradeoff as the algorithm's reward. * .
[0062] Many variations of these methods are contemplated. In a first variation, in a subset Instead of optimizing on Ω. For example, the finite subset may consist of biases and / or weights and / or parameters of a nonlinear function and / or any subset of these elements. Such a subset may, for example, be defined as a subset of layers, such as the last k layers, or the biases of the last k layers, or a subset of neurons.
[0063] In a second variation, additional constraints are imposed on the allowed values of ω. For example, Ω (or Ω′) may be divided into non-overlapping subsets And the search may be restricted to a subset ω such that for any pair (ω i ,ω j ),ω i ,ω j Do not belong to the same subsetΩ l Alternatively, the constraint could be that at most m elements of ω belong to any subset Ω l One motivation for this approach is to spread the impact of the update across the entire model.
[0064] The weight selection can be performed using a no-reference loss. In this case, an optimization can be performed to maximize a function that quantifies the quality of the output of M(i). As an example, the BRISQUE metric can be used to evaluate the quality of an image. Other no-reference losses are also discussed above.
[0065] In an embodiment, the process is not limited to deep neural networks, but any machine learning model may be used.
[0066] In step 230, the fine-tuning algorithm calculates the parameter ω * Corresponding updates The fine-tuning loss is, for example:
[0067]
[0068] This loss can be the same as or different from the loss used in the second step. The loss can also contain additional terms, for example, terms that impose some constraints on the weights (such as sparsity constraints). In a variation of this embodiment, the input i can be Joint fine-tuning.
[0069] If no output is provided, the update may be computed by optimizing any loss that does not require a reference signal, such as the no-reference loss described above.
[0070] In step 240, the input and weight updates are prepared for the inference engine, for example, by packaging the input and weight updates into a data set that is stored together.
[0071] Figure 3 FIG. 300 shows a process of an inference engine according to at least one embodiment of a general data conversion system. Fig.10 The method is implemented by the apparatus 1000 of the embodiment of the present invention, and more specifically, by the processor 1010 or the decoder 1030 of such an apparatus. In step 310, the device obtains an input and an associated parameter update for a selected parameter subset. In step 320, the subset selection of parameters is recalculated based on the input and the model M using the same process as in the optimizer (or a process that gives the same result). In step 330, the model M is updated based on the recalculated parameter subset and parameter update. In step 340, the updated model M processes the input and determines the output.
[0072] In at least one embodiment, Figure 2 The process of 200 and Figure 3 An example of a parameter used in process 300 is a weight, such as Figure 4 shown.
[0073] Figure 4 An architecture diagram of an optimizer and an inference engine according to at least one embodiment is shown. The optimizer (400) obtains an input i (411) and optionally obtains a target output o (412). In step 420, the optimizer selects a weight subset of fixed size s, based on the input i. In step 430, the fine-tuning algorithm calculates the parameter ω * Corresponding updates Finally, the input and weight updates are stored together and / or prepared for the inference engine (410). The inference engine obtains data (440) including the input (441) and the associated weight updates (443). In step 450, the weight subset ω* is recalculated from the input and the model M using the same process as in the optimizer (or a process that gives the same results). In step 460, the model M is updated based on the recalculated weight subset and weight updates. In step 470, the updated model M processes the input to determine the output (480).
[0074] For readability, the description and drawings refer to updating weights. However, any other parameter of the neural network can be updated using the same techniques. In other words, the embodiments described below as applied to weights also apply more generally to any parameter of a neural network model, i.e., parameters selected from the set consisting of biases of nonlinear functions of the model, weights, parameters, a subset of layers of the model, a specific layer of the model, biases of a specific layer of the model, and a subset of neurons of the model.
[0075] According to at least one embodiment, a general data transformation system (e.g., a system for selecting and optimizing a specific subset of weights as an input without transmitting identifiers of these weights) is provided. Figure 2 , Figure 3 , Figure 4 The processes of these devices are respectively described in Figure 5A and Figure 5B The architectures of these devices are shown in Figure 6 and Figure 7 In this context of end-to-end compression, the inference engine is the decoder, the optimizer is the encoder, and the weight selection process is done in both the encoder and the decoder. Here, the output o and input i of the general method are the image / frame x to be encoded and the corresponding quantized embedding vector Image encoders and decoders can be used as basic components of video compression systems.
[0076] Figure 5A The process of an encoder in the context of an end-to-end image compression system according to at least one embodiment is shown. Process 500A, for example, consists of Fig.10 The invention is implemented by the device 1000, and more specifically, by the processor 1010 or the encoder 1030 of such a device. Figure 6 An example of the architecture of such an encoder 600 is shown. The encoder is based on Figure 4 The same principles of the optimizer 400 are applied, but in the context of end-to-end image compression.
[0077] In step 510, the encoder obtains an image x. In step 515, the encoder determines an embedding vector y by using a depth encoder (610). In step 520, the embedding vector y is quantized using a quantizer (611). In step 525, the encoder calculates the embedding vector y based on the quantized embedding vector To select a weight subset ω of fixed size s, * :
[0078]
[0079] This corresponds to the above Figure 2 The selection is done by a selection element 620 also present in the decoder to allow the decoder to perform the same operation. In the first method proposed above, i.e. when selecting parameters that have a greater impact on the output at the decoder, a subset of weights can be calculated, for example, as follows:
[0080]
[0081] in represents the gradient, and M dec Represents the decoder part of a deep neural network.
[0082] The second approach (using a machine learning model N) is similar to the above approach. However, in addition to improvements in model predictions, the loss used to train these models can also take into account the bit length of parameter updates. In other words, these models will be trained to produce a subset of weights that will achieve the best r / d tradeoff rather than just distortion.
[0083] In step 530 (associated with element 630 of the architecture), the fine-tuning algorithm calculates the parameter ω * Corresponding updates For end-to-end encoding, the fine-tuning loss can be:
[0084]
[0085] The loss may also contain additional terms, for example, terms that impose some additional constraints on the weights (such as sparsity constraints).
[0086] In a variation of this embodiment, the quantization embedding Can be used with Joint fine-tuning. In this case, another loss is used, for example:
[0087]
[0088] In step 535, these weight updates are typically quantized (631) and encoded (632). These quantized weight updates are represented by express.
[0089] In step 540, a bitstream is generated (640), for example, by aggregating the following data:
[0090] Quantized Embedding For example, the quantized embedding is encoded by an arithmetic encoder (612) or another encoder, thereby generating an encoded quantized embedding (641), and
[0091] (Quantized) weight update For example, the (quantization) weight update is encoded by an arithmetic encoder (632) or another encoder, thereby generating an encoded quantization weight update (643).
[0092] Optionally, the quantization and encoding of the weight updates may depend on some parameters. These parameters may be the same for all images, or some or all of them may be fine-tuned for each image. In the latter case, the bitstream also includes the values of these parameters (denoted by C) inserted as coding information (644).
[0093] Embedded quantization and coding can also depend on additional parameters. These elements can be arranged in any order or even interleaved in the bitstream.
[0094] Figure 5B FIG. 5 shows a process of a decoder in the context of an end-to-end image compression system according to at least one embodiment. Fig.10 The invention is implemented by the device 1000, and more specifically, by the processor 1010 or the decoder 1030 of such a device. Figure 7 The architecture of such a decoder 700 is shown. The decoder is based on Figure 4 The same principles of the inference engine 410 are applied, but in the context of end-to-end image compression.
[0095] In step 550, the quantization embedding and weight updates are extracted from the bitstream (640). Both the quantization embedding and the quantization weight updates are optionally decoded using parameters carried by the encoding information also extracted from the bitstream (711 and 713).
[0096] In step 560, a subset of weights ω is determined (620) based on the quantized embedding * This step must produce the same result as the corresponding step 540 of the encoder. This can be achieved by using the same process (620). An advantage of the embodiments described herein is that the subset ω * It does not need to be included in the bitstream, but at the expense of extra computations (620) in the decoder to perform the selection.
[0097] In step 570, the depth decoder is updated (720) based on the weight subset and the quantization weight update. In step 580, the image is then decoded (730) by the updated decoder according to the quantized embedding.
[0098] The above embodiments are based on a system in which the reversible operations associated with the quantization of the weight updates are also inverted in the AD block. The same system can be described using additional blocks that perform these operations, such as those called "dequantization" or "inverse quantization". An example of such a reversible operation is scaling the weight updates before quantization to change the quantization resolution.
[0099] Figure 8 An example of comparing size information of a bitstream generated according to at least one embodiment with a bitstream of the same input generated without any of the presented embodiments is shown. Figure 5A An example implementation of process 500A of , based on an embodiment related to end-to-end image compression, generates a bitstream 800. The bitstream includes an encoded quantized embedding 801, an encoded weight update 802, and optional encoding information 803. The sizes of these different elements are 28160 bits, 1440 bits, and 40 bits, respectively. These specific numbers are obtained when using the "cheng2020_anchor" end-to-end encoder of the compressAI library as model M. The bitstream 810 is generated based on the most advanced end-to-end neural network compression system supporting fine-tuning, based on the same input and using the same settings. In contrast to the proposed embodiment, such a system requires information 811 identifying the weights to be updated (i.e., for example, their positions based on indices) to be passed from the encoder to the decoder. Although the other elements of the bitstream 810 have the same size as in the bitstream 800, the additional data 811 increases the size of the encoded message. Another implementation will produce data of other sizes, but will still provide the same advantage: reducing the size of the generated bitstream and thus increasing the performance of the end-to-end image compression system.
[0100] According to at least one embodiment, a general data transformation system (e.g., a system for selecting and optimizing a specific subset of weights as an input without transmitting identifiers of these weights) is provided. Figure 2 , Figure 3 , Figure 4 Described) is applied to sound enhancement or compression systems.
[0101] Fig. 9An example of applying an optimizer (900) and an inference engine (901) in the context of sound enhancement according to at least one embodiment is shown. Deep sound upsampling is a sound improvement method in which an audio signal is converted by an upsampling neural network that increases the number of sampling points. One possible use of sound upsampling is to improve the quality of low-frequency or downsampled audio signals. In this context, the method can be applied as follows to improve the audio quality of downsampled audio signals sent to a device. The downsampled original audio signal (or more precisely, the audio that the inference engine will receive) is the input i (911). The original high-frequency audio signal is the target output (o912). The upsampled neural network is model M.
[0102] The raw audio signal may be pre-processed by an optimizer before being sent to an audio playback device (such as a mobile device or a computer device that reads an audio file streamed by a server or any other device suitable for playing audio. In this case, the optimizer would be implemented in the audio server and the inference engine would be implemented in the audio playback device.
[0103] The server first obtains the audio file and the high frequency target signal. The server then prepares the original audio file for transmission (if not already done), for example, by compressing the original audio file (915) to obtain a signal 918, and processes the original audio file, for example, decoding or decompressing (916) to recover a signal 919 to be used by the inference engine. As an example, this may mean encoding (i.e., compressing) and decoding (i.e., decompressing) the signal. The modified signal (919) is then used to select a subset of weights of the model M to be modified (920). The weights are then updated Optimized or fine-tuned (930), prepared to be sent to the inference engine and concatenated with the audio file ready for transmission (940). Both are sent to a device with an inference engine or stored for later use.
[0104] On the audio playback device, the audio signal (941) and weight updates (943) are received and recovered by the inference engine. This includes any processing (916) (e.g., decoding or decompression) performed in the inference engine. The recovered audio signal (944) is then used to determine a subset of weights (950), and the subset and the weight updates are used (960) to update the model M to an updated model M'. In other words, the selection of the parameter subset is independent of the information representing the parameter updates (940). Finally, the received audio signal (944) is used as input to the updated model M' (970), and a resulting upsampled audio signal (980) is generated. The audio signal can be played to a user or used for any other purpose.
[0105] The optimizer 900 and the inference engine 901 are composed of, for example, Fig.10The optimizer 900 and the inference engine 901 are respectively based on the device 1000 and the decoder 1030 of the device. Figure 4 The optimizer 400 and Figure 4 The same principles of the inference engine 410 are applied, but in the context of end-to-end image compression.
[0106] Fig.10 1 shows a block diagram of an example of a system in which various aspects and embodiments are implemented. The system 1000 may be embodied as an apparatus including various components described below and configured to perform one or more aspects described in the present application, such as Figure 4 Optimizer 400, or Figure 4 The reasoning engine 410, or Figure 6 Encoder 600, or Figure 7 Decoder 700, or Fig. 9 Optimizer 900, or Fig. 9 Such a system can achieve Figure 2 The optimizer process 200, or Figure 3 The reasoning process 300, or Figure 5A The encoding process 500A, or Figure 5B Decoding process 500B, or Fig.11 The encoding process 1101, or Fig.12 The decoding process 1201 of system 1000. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, encoders, code converters, and servers. The elements of system 1000 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed over multiple ICs and / or discrete components. In various embodiments, system 1000 is coupled to other similar systems or other electronic devices via, for example, a communication bus or by dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.
[0107] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement, for example, various aspects described in this document. The processor 1010 may include embedded memory, input and output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile storage, including but not limited to EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash memory, magnetic disk drive, and / or optical disk drive. As a non-limiting example, the storage device 1040 may include an internal storage device, an attached storage device, and / or a network accessible storage device.
[0108] The system 1000 includes an encoder / decoder module 1030, which is configured to process data, for example, to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that can be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of the encoding and decoding modules. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of the system 1000, or may be combined within the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0109] Program code to be loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from equations, formulas, operations, and operational logic processing.
[0110] In several embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store the operating system of the television. In at least one embodiment, a fast external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations such as for MPEG-2, HEVC, or VVC (Universal Video Coding).
[0111] Input to the elements of system 1000 may be provided through various input devices, as indicated in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives an RF signal transmitted over the air, such as by a broadcaster, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0112] In various embodiments, the input device of block 1130 has associated corresponding input processing elements known in the art. For example, the RF portion may be associated with elements necessary to: (i) select a desired frequency (also referred to as selecting a signal, or band limiting a signal to a frequency band), (ii) downconvert the selected signal, (iii) again bandlimit to a narrower frequency band to select a signal frequency band that may be referred to as a channel in some embodiments, (iv) demodulate the downconverted and bandlimited signal, (v) perform error correction, and (vi) demultiplex to select a desired packet stream. The RF portion of various embodiments includes one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a frequency band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF portion may include a tuner that performs various of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near baseband frequency) or baseband. In a set-top box embodiment, the RF part and its associated input processing element receive the RF signal transmitted by wired (for example, cable) medium, and filter to the target frequency band again by filtering, down-conversion and perform frequency selection.Various embodiments rearrange the order of above-mentioned (and other) elements, remove some and / or add other elements of similar or different functions in these elements.Adding element can include inserting element between existing element, for example, such as inserting amplifier and analog-to-digital converter.In various embodiments, the RF part comprises antenna.
[0113] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting the system 1000 to other electronic devices via USB and / or HDMI connections. It should be appreciated that aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, for example, in a separate input processing IC or in the processor 1010 as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, in a separate interface IC or in the processor 1010 as desired. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data streams as desired for presentation on an output device.
[0114] The various components of the system 1000 may be disposed in an integrated housing in which the various components may be interconnected and transmit data between them using suitable connection means (eg, an internal bus 1140 known in the art, including an I2C bus, wiring, and a printed circuit board).
[0115] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data through the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0116] In various embodiments, a Wi-Fi network such as IEEE 802.11 is used to stream data to the system 1000. The Wi-Fi signals of these embodiments are received through a communication channel 1060 and a communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to external networks including the Internet to allow streaming applications and other top communications. Other embodiments provide streaming data to the system 1000 using a set-top box that delivers data through an HDMI connection of an input block 1130. Still other embodiments provide streaming data to the system 1000 using an RF connection of an input block 1130.
[0117] The system 1000 can provide output signals to various output devices, including a display 1100, a speaker 1110, and other peripheral devices 1120. In various examples of embodiments, the other peripheral devices 1120 include one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide functions based on the output of the system 1000. In various embodiments, control signals are transmitted between the system 1000 and the display 1100, the speaker 1110, or other peripheral devices 1120 using signaling such as AVLink, CEC, or other communication protocols that implement device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 1000 via dedicated connections through the respective interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to the system 1000 using a communication channel 1060 via a communication interface 1050. The display 1100 and the speaker 1110 can be integrated into a single unit with other components of the system 1000 in an electronic device (e.g., such as a television). In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip, for example.
[0118] For example, if the RF input section 1130 is part of a separate set-top box, the display 1100 and the speaker 1110 can be separated from one or more of the other components alternatively. In various embodiments where the display 1100 and the speaker 1110 are external components, an output signal can be provided via a dedicated output connection (including, for example, an HDMI port, a USB port, or a COMP output). The implementation described herein can be implemented in, for example, a method or process, a device, a software program, a data stream, or a signal. Even if only discussed in the context of a single implementation form (for example, discussed only as a method), the implementation of the features discussed in other forms (for example, a device or a program) can also be implemented. The device can be implemented, for example, with appropriate hardware, software, and firmware. The method can be implemented in, for example, a device, for example, such as a processor that generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes a communication device, for example, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between terminal users.
[0119] Fig.11 FIG. 1101 shows a process of an image encoder according to at least one embodiment. Fig.10The method is implemented by the apparatus 1000 of the present invention, and more specifically, by the processor 1010 or the encoder 1030 of such an apparatus. In step 1111, the processor obtains input data. In step 1112, the processor selects a parameter subset based on the input data, the parameter subset being used to fine-tune the model of the first neural network. In step 1113, the processor determines a parameter update for the selected parameter subset based on the loss function. In step 1114, the processor packages the input data and information representing the parameter update.
[0120] Fig.12 FIG. 1201 shows a process of an image decoder according to at least one embodiment. Fig.10 The method is implemented by the apparatus 1000 of the present invention, and more specifically, by the processor 1010 or the decoder 1030 of such an apparatus. In step 1211, the processor obtains input data and information representing parameter updates for the selected parameter subset. In step 1212, the processor selects a parameter subset based on the input data, the parameter subset being used to fine-tune the model of the first neural network. In step 1213, the processor updates the model of the first neural network based on the parameter updates for the selected parameter subset. In step 1214, the processor determines output data by processing the input data using the updated first neural network.
[0121] References to "one embodiment" or "an embodiment" or "an implementation" or "implementations" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" or "in one implementation" or "in an implementation" and any other variations appearing in various places throughout this specification are not necessarily all referring to the same embodiment.
[0122] Additionally, this application or its claims may refer to "determining" various information. Determining information may include one or more of: for example, estimating information, calculating information, predicting information, or retrieving information from a memory.
[0123] Additionally, the application or its claims may refer to "accessing" various information. Accessing information may include one or more of: for example, receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, predicting information, or estimating information.
[0124] Furthermore, the application or its claims may refer to "receiving" various information. Like "accessing," receiving is intended to be a broad term. Receiving information may include one or more of: for example, accessing information or retrieving information (e.g., from a memory or optical media storage device). Furthermore, "receiving" is often involved in one way or another during an operation, such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0125] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", any of the following uses of " / ", "and / or", and "at least one of" are intended to include selecting only the first listed option (A), or only the second listed option (B), or selecting both options (A and B). As another example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include selecting only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or only the first and second listed options (A and B), or only the first and third listed options (A and C), or only the second and third listed options (B and C), or all three options (A and B and C). It will be readily appreciated by those of ordinary skill in this and related arts that this can be extended to as many listed items as possible.
[0126] It will be apparent to those skilled in the art that implementations may generate a variety of signals formatted to carry information that may be stored or transmitted, for example. Information may include, for example, instructions for executing a method, or data generated by one of the described implementations. For example, a signal may be formatted to carry a bitstream of the described embodiments. Such a signal may be formatted as, for example, an electromagnetic wave (e.g., using a radio frequency portion of a spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. As is known, the signal may be transmitted over a variety of different wired or wireless links. The signal may be stored on a processor readable medium.
Claims
1. A method, include: Get input data; selecting a subset of parameters for fine-tuning a model of a first neural network based on the input data; determining a parameter update for the selected subset of parameters based on the loss function; as well as The input data and information representing the parameter updates are packaged.
2. The method according to claim 1, in, Determining parameter updates is also based on the target output.
3. The method according to any one of claims 1 or 2, in, The parameter subset is selected based on the gradient of the output of the model to select the parameters with the greatest impact.
4. The method according to any one of claims 1 or 2, in, The selecting the subset of parameters is based on a loss function using a no-reference metric determined based on the input data.
5. The method according to any one of claims 1 to 4, in, The parameters are selected from a set comprising: biases of the model, weights of the model, parameters of a nonlinear function of the model, a subset of layers of the model, a specific layer of the model, the biases of a specific layer of the model, and a subset of neurons of the model.
6. The method according to any one of claims 1 to 5, wherein the method is suitable for encoding an image, wherein the first neural network is used for decoding and the second neural network is used for encoding, and wherein the method further comprises: include: - determining an embedding by encoding the image using the second neural network; - quantizing the embedding; as well as - performing a selection of a subset of parameters based on the quantized embedding, wherein the first neural network is updated using the parameter update for the selected subset of parameters.
7. The method according to claim 6, in, Determining parameter updates is also based on the image.
8. A method according to any one of claims 1 to 5, suitable for compressing sound, wherein the input data is an audio signal, the method further comprising: include: compressing the audio signal; Decompress the compressed audio signal; performing the selection of the parameter subset based on the decompressed compressed audio signal; as well as Package the compressed audio signal and parameter updates.
9. A method, include: obtaining input data and updated information representing a selected subset of parameters; selecting a subset of parameters for fine-tuning a model of a first neural network based on the input data; updating a model of the first neural network based on the update for the selected subset of parameters; as well as Output data is determined by processing the input data using the updated first neural network.
10. The method of claim 9, adapted for decoding an image, wherein the input data is a quantized embedding.
11. The method according to claim 9, adapted for decoding sound, in, The selection of the parameter subset is also based on the obtained decompressed audio signal.
12. The method according to any one of claims 1 to 11, in, The selection of the parameter subset is independent of the information representing the parameter update.
13. An apparatus comprising a processor, the processor being configured to: Get input data; selecting a subset of parameters for fine-tuning a model of a first neural network based on the input data; determining a parameter update for the selected subset of parameters based on the loss function; as well as The input data and information representing the parameter updates are packaged.
14. The device according to claim 13, in, Determining parameter updates is also based on the target output.
15. The device according to any one of claims 13 or 14, in, The parameter subset is selected based on the gradient of the output of the model to select the parameters with the greatest impact.
16. The device according to any one of claims 13 or 14, in, The selecting the subset of parameters is based on a loss function using a no-reference metric determined based on the input data.
17. The device according to any one of claims 13 to 16, in, The parameters are selected from a set comprising: biases of the model, weights of the model, parameters of a nonlinear function of the model, a subset of layers of the model, a specific layer of the model, the biases of a specific layer of the model, and a subset of neurons of the model.
18. The apparatus according to any one of claims 13 to 17, adapted to encode an image, wherein the first neural network is used for decoding and the second neural network is used for encoding, the processor being further configured to: determining an embedding by encoding the image using the second neural network; quantizing the embedding; and The selection of a subset of parameters is performed based on the quantized embedding, in, The first neural network is updated using the parameter updates for the selected subset of parameters.
19. The device according to claim 18, in, Determining parameter updates is also based on the image.
20. The apparatus according to any one of claims 13 to 17, adapted for compressing sound, wherein the input data is an audio signal, the processor being further configured to: compressing the audio signal; Decompress the compressed audio signal; The selecting of the subset of parameters is performed based on the decompressed compressed audio signal; and Package the compressed audio signal and parameter updates.
21. An apparatus comprising a processor, the processor being configured to: obtaining input data and information representing parameter updates for the selected subset of parameters; selecting a subset of parameters for fine-tuning a model of a first neural network based on the input data; updating a model of the first neural network based on the parameter updates for the selected subset of parameters; as well as Output data is determined by processing the input data using the updated first neural network.
22. The apparatus of claim 21, adapted to decode an image, wherein the input data is a quantized embedding.
23. The device according to claim 21, adapted to decode sound, in, The selection of the parameter subset is also based on the obtained decompressed audio signal.
24. The device according to any one of claims 13 to 23, in, The information indicative of the parameter update does not include information indicative of the selection of the parameter subset.
25. A computer program comprising program code instructions for implementing the method according to at least one of claims 1 to 12 when the program code instructions are executed by a processor.
26. A non-transitory computer-readable medium comprising program code instructions for implementing the method according to at least one of claims 1 to 12 when the program code instructions are executed by a processor.