Data Compression and Reconstruction Using a Sparse Meta-Learning Neural Network
By employing a data-reconstruction neural network to determine sparse updates for a subset of network parameters, the system addresses the scalability issues of existing INR methods, achieving improved compression rates and reconstruction quality.
Patent Information
- Application Number
- JP2024564946
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-03
- Filing Date
- 2023-05-03
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-05-03
AI Technical Summary
Existing implicit neural representation (INR) methods for data compression require transmitting all dense network parameters for each data signal, which does not scale well for real-world compression scenarios and large-scale data signals.
The system uses a data-reconstruction neural network to compress input signals by determining updates to a subset of network parameters, encouraging sparse updates, and transmitting only these sparse updates, thereby reducing the compression cost while maintaining high reconstruction quality.
This approach significantly improves the compression rate by reducing the data required for communication between the compression and restoration systems, while maintaining high reconstruction quality across various data modalities.
Smart Images

Figure 2025517129000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Application No. 63 / 338,018, filed on May 3, 2022. The disclosure of the prior application is considered a part of this application and is incorporated by reference into the disclosure of this application.
[0002] This specification relates to compressing and reconstructing input signals using a machine - learning model.
Background Art
[0003] As an example, a neural network is a machine - learning model that uses one or more layers of non - linear units to predict the output of received inputs. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to another layer in the network, such as the next hidden layer or the output layer. Each layer of the network generates an output from the received inputs according to the current values of each set of weights.
Summary of the Invention
Means for Solving the Problems
[0004] This specification describes a system implemented as a computer program on one or more computers that compresses an input signal using a data - reconstruction neural network having network parameters.
[0005] In particular, the system uses the input signal to determine an update to a shared value of a subset of the network parameters.
[0006] The system then generates a compressed representation of the input signal that identifies each update to the subset of the network parameters.
[0007] A system or another system can use a data reconstruction neural network to restore an input signal by determining an updated value of a network parameter using a shared value and a compressed representation and reconstructing the input signal using the updated value.
[0008] In one aspect, the method comprises the step of maintaining data specifying a shared value of network parameters of a data reconstruction neural network, wherein the data reconstruction neural network is configured to receive an input specifying coordinates from a coordinate space of an input data signal, process the input according to the network parameters, and generate, as output, one or more predicted values of the input data signal at the specified coordinates; the step of receiving a new input signal including one or more respective new values at each of a plurality of new coordinates; the step of determining respective updates for each of a subset of the network parameters, the step of determining, in each of one or more internal iterations, one or more sets of current values of the network parameters of the data reconstruction neural network, and for each set of current values, determining respective gate values for each network parameter within the subset that specify whether respective updates of the subset are set to zero according to a set of distribution parameters; the step of determining a set of current values by setting current values for any network parameters not within the subset based on the shared value of the network parameters, setting current values for any network parameters within the subset for which the respective gate values specify that respective updates of the subset are set to zero based on the shared value of the network parameters, and setting current values for any network parameters within the subset for which the respective gate values specify that respective updates of the subset are not set to zero based on the shared value of the network parameters and the respective updates of the network parameters; the step of, for each of one or more sets of current values of the network parameters, processing an input specifying the new coordinates according to the set of current values of the network parameters using the data reconstruction neural network to generate one or more current predicted values for each of the new coordinates; and the respective updates of the network parameters within the subset, and for each set of current values,(i) For each new coordinate, a reconstruction quality term that measures the error between one or more current predicted values of the new coordinate and one or more new values of the new coordinate in the new input signal, and (ii) for each of the distribution parameters of an internal loss function that includes a differentiable sparsity term that penalizes non-zero updates to a subset of the network parameters, determining a respective gradient; for each of the subset of network parameters and the distribution parameters, updating each respective update using each respective gradient; and generating a compressed representation of the new input signal that identifies each respective update for the subset of network parameters.
[0009] In some implementations, the method includes storing the compressed representation in association with data that identifies the new input signal.
[0010] In some implementations, the method includes transmitting the compressed representation over a data communication network.
[0011] In some implementations, for each of the subset of network parameters, the step of determining each respective update includes, after one or more internal iterations, determining a respective final gate value for each network parameter in the subset according to the distribution parameters after one or more internal iterations; setting the final update of the network parameter to zero for any network parameter in the subset for which each respective update of the subset is specified by the respective final gate value to be set to zero; and setting the respective final update of the network parameter based on the respective update of the network parameter after one or more internal iterations for any network parameter in the subset for which each respective update of the subset is specified by the respective final gate value not to be set to zero.
[0012] In some implementations, a subset of the network parameters is a suitable subset of the network parameters.
[0013] In some implementations, a first neural network layer within a neural network has network parameters that include (i) a weight tensor and (ii) a modulation tensor, where the modulation tensor is within the subset and the weight tensor is not within the subset.
[0014] In some implementations, the first neural network layer is configured to perform operations that include calculating an affine transformation between the weight tensor and an input to the layer, and applying the modulation tensor to the output of the affine transformation.
[0015] In some implementations, the network parameters of the first neural network layer further include (iii) a bias tensor, where the bias tensor is not within the subset, and applying the modulation tensor to the output of the affine transformation includes applying the modulation tensor and the bias tensor to the output of the affine transformation.
[0016] In some implementations, the subset includes all of the network parameters of the neural network.
[0017] In some implementations, the method includes maintaining data that specifies shared distribution parameters, and setting the distribution parameters to be equal to the shared distribution parameters prior to a first iteration of one or more internal iterations.
[0018] In some implementations, the method includes training a neural network with a plurality of training signals to determine a shared value of the network parameters.
[0019] In some implementations, the step of training a neural network with a plurality of training signals includes training the neural network to minimize an internal loss function that is evaluated after performing a fixed number of internal training steps starting from a given set of shared values of network parameters, for a given set of shared values.
[0020] In some implementations, the step of training a neural network with a plurality of training signals includes training the neural network with a plurality of training signals to determine a shared value of network parameters and a shared distribution parameter.
[0021] In some implementations, the step of training a neural network with a plurality of training signals includes training the neural network to minimize an internal loss function that is evaluated after performing a fixed number of internal training steps starting from a given set of shared values of network parameters and a given set of shared distribution parameters, for a given set of shared values of network parameters and a given set of shared distribution parameters.
[0022] In some implementations, the step of determining respective gate values for each network parameter within a subset, which specifies whether each update of the subset is set to zero according to a set of distribution parameters, includes the step of sampling noise from a noise distribution and the step of mapping the set of distribution parameters and the sampled noise to the respective gate values of the network parameters within the subset.
[0023] In some implementations, the step of mapping the set of distribution parameters and the sampled noise to the respective gate values of the network parameters within the subset includes the step of applying a hard rectification to the value determined from the distribution parameters and the sampled noise.
[0024] In some implementations, the compressed representation of the new input signal identifies only non-zero updates of network parameters within a subset of the network parameters.
[0025] In some implementations, each network parameter within the subset has a respective gate value that is different from each other network parameter within the subset.
[0026] In some implementations, two or more network parameters within the subset share the same respective gate value.
[0027] In some implementations, the differentiable sparsity term measures the sum of the respective probabilities for each gate value, where each probability for each gate value is defined by distribution parameters and specifies the likelihood that the respective update of one or more network parameters corresponding to the gate value is set to a non-zero value.
[0028] In some implementations, the new input signal is an image, each coordinate corresponds to a respective pixel of the image in a two-dimensional coordinate space, and one or more respective values include one or more intensity values of the pixel.
[0029] In some implementations, the new input signal is a three-dimensional image, each coordinate corresponds to a respective voxel of the image in a three-dimensional coordinate space, and one or more respective values include one or more intensity values of the voxel.
[0030] In some implementations, the new input signal is a point cloud, each coordinate corresponds to a respective point in a three-dimensional coordinate space, and one or more respective values include the respective intensity of each point.
[0031] In some implementations, the new input signal is video, each coordinate is a three - dimensional coordinate identifying the spatial position within the video frame of a pixel from the video, and one or more respective values include one or more intensity values of the pixel.
[0032] In some implementations, the new input signal is an audio signal, each coordinate is a respective point in time within the audio signal, and one or more respective values include one or more values defining the amplitude of the audio signal at each point in time.
[0033] In some implementations, the new input signal represents a signed distance function, and one or more respective values include the signed distance from the boundary of the object at the corresponding coordinate.
[0034] In some implementations, the new input signal represents a rendered scene.
[0035] In another aspect, the method includes receiving a request to reconstruct an input data signal, training a data reconstruction neural network to reconstruct the input signal while applying a differentiable sparsity term that penalizes the update of a subset of non - zero network parameters, thereby obtaining (i) data specifying shared values of the network parameters of the data reconstruction neural network, and (ii) data specifying respective updates of a subset of the network parameters determined for the input data signal, and generating a reconstructed input signal, including processing an input specifying a coordinate using the data reconstruction neural network according to the values of the network parameters defined by the shared values and the respective updates to generate one or more values of the reconstructed input signal at the coordinate for each of a plurality of coordinates from the coordinate space of the input data signal.
[0036] In some implementations, the reconstructed input signal has respective values for more coordinates than the input signal.
[0037] In another aspect, the present specification describes one or more computer-readable storage media storing a compressed representation of a new data signal, the compressed representation of the data signal being to maintain data specifying a shared value of network parameters of a data reconstruction neural network, the data reconstruction neural network being configured to receive an input specifying coordinates from a coordinate space of an input data signal, process the input according to the network parameters, and generate, as an output, one or more predicted values of the input data signal at the specified coordinates, maintaining, receiving a new input signal including one or more respective new values at each of a plurality of new coordinates, and determining a respective update for each of a subset of the network parameters, determining the respective updates being determining, for each of one or more internal iterations, one or more sets of current values of the network parameters of the data reconstruction neural network, and for each set of current values, determining a respective gate value for each network parameter within the subset that specifies whether the respective update of the subset is set to zero according to a set of distribution parameters, setting a current value for any network parameter not within the subset based on the shared value of the network parameters, setting a current value for any network parameter within the subset for which the respective update of the subset is set to zero based on the shared value of the network parameters, and setting a current value for any network parameter within the subset for which the respective update of the subset is not set to zero based on the shared value of the network parameters and the respective update of the network parameters, thereby determining one or more sets of current values of the network parameters, and for each of one or more sets of current values of the network parameters, using the data reconstruction neural network to generate one or more current predicted values for each of the new coordinates according to the set of current values of the network parameters,Determining each update, including processing an input specifying new coordinates; for each update of the network parameters within the subset and each set of current values, (i) for each new coordinate, a reconstruction quality term measuring the error between one or more current predicted values of the new coordinate and one or more new values of the new coordinate in the new input signal, and (ii) for non-zero updates to a subset of the network parameters, determining a respective gradient for each of the distribution parameters of an internal loss function including a differentiable sparsity term imposing a penalty; updating each respective update using the respective gradient for each of the subset of network parameters and the distribution parameters; and generating a compressed representation of the new input signal identifying each respective update to the subset of network parameters.
[0038] In another aspect, the present specification describes a compressed representation of a data signal, such as a bitstream, where the compressed representation of the data signal is to maintain data that specifies a shared value of network parameters of a data reconstruction neural network. The data reconstruction neural network is configured to receive an input that specifies coordinates from the coordinate space of the input data signal, process the input according to the network parameters, and generate, as outputs, one or more predicted values of the input data signal at the specified coordinates. This includes maintaining, receiving a new input signal that includes one or more respective new values at each of a plurality of new coordinates, and determining a respective update for each of a subset of the network parameters. Determining each respective update is to determine one or more sets of current values of the network parameters of the data reconstruction neural network in each of one or more internal iterations. For each set of current values, determining a respective gate value for each network parameter within the subset to specify whether each respective update of the subset is set to zero according to a set of distribution parameters, setting the current value for any network parameter not within the subset based on the shared value of the network parameters, setting the current value for any network parameter within the subset where each respective gate value specifies that each respective update of the subset is set to zero based on the shared value of the network parameters, and setting the current value for any network parameter within the subset where each respective gate value specifies that each respective update of the subset is not set to zero based on the shared value of the network parameters and each respective update of the network parameters, thereby determining one or more sets of current values of the network parameters. For each of one or more sets of current values of the network parameters, using the data reconstruction neural network to generate one or more current predicted values for each of the new coordinates according to the set of current values of the network parameters for each of the new coordinates.Determining each update, including processing an input specifying new coordinates; for each update of the network parameters within the subset and each set of current values, determining a respective gradient with respect to each of the distribution parameters of an internal loss function including a reconstruction quality term that measures an error between one or more current predicted values of the new coordinates and one or more new values of the new coordinates in the new input signal for each new coordinate, and a differentiable sparsity term that penalizes non-zero updates for the subset of network parameters; updating each respective update using each respective gradient for each of the subset of network parameters and the distribution parameters; and generating a compressed representation of the new input signal that identifies each respective update for the subset of network parameters.
[0039] Certain embodiments of the subject matter described herein can be implemented to realize one or more of the following advantages.
[0040] In some data compression techniques, neural networks that map from a coordinate space to an underlying continuous signal are used to compress data. This type of approach is referred to as an implicit neural representation (INR), and through careful architecture search, the INR has been shown to exhibit superior performance compared to other established compression methods for lower dimensional data or when a low compression rate is required.
[0041] However, with these approaches, for each data signal to be compressed, a set of all dense network parameters of a neural network (a "data reconstruction neural network") that maps from coordinates to underlying signal values needs to be transmitted from the compression system to the restoration system. Thus, the INR has not been shown to scale to real-world compression scenarios and large-scale real-world data signals.
[0042] This specification describes techniques for significantly improving the compression rate of INR (thus significantly reducing the compression cost) while maintaining a high reconstruction quality (e.g., low reconstruction error). In particular, this specification describes techniques for significantly reducing the amount of data that needs to be communicated between the compression system and the restoration system for each signal within the INR framework while maintaining a high reconstruction quality. More specifically, this improvement in the compression rate is achieved by learning the update of the parameter values for each signal for a subset of the network parameters in a way that encourages the updates to be sparse, such that only these sparse updates need to be transmitted to the restoration system, thus maintaining a high compression rate. Further, the described techniques can be applied to achieve a high reconstruction quality in various diverse data modalities such as images, manifolds, signed distance functions, 3D shapes, and scenes.
[0043] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
Brief Description of the Drawings
[0044]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0045] Like reference numerals and designations in the various drawings indicate like elements.
[0046] FIG. 1 shows an exemplary compression system 100 and an exemplary restoration system 150.
[0047] Systems 100 and 150 are examples of systems implemented as computer programs on one or more computers in one or more locations that can implement the systems, components, and techniques described below.
[0048] Compression system 100 is a system that uses data reconstruction neural network 110 to compress input signal 102 to generate a compressed representation 112 of input signal 102, e.g., a bitstream encoding signal-specific information necessary to reconstruct input signal 102.
[0049] Restoration system 150 is a system that uses data reconstruction neural network 110 to restore input signal 102 from compressed representation 112, e.g., to generate a reconstruction 152 of input signal 102.
[0050] In general, the compression system 100 and the restoration system 150 may be located in the same place or in separate places. That is, the compression system 100 can be implemented on the same set of one or more computers as the restoration system 150, or on a different set of one or more computers in a different location from the restoration system 150.
[0051] The compressed representation generated by the compression system 100 can be provided to the restoration system 150 in any of various ways.
[0052] For example, the compressed data can be stored in association with data identifying the corresponding input signal (e.g., in a physical data storage device or a logical data storage area), and then retrieved from storage and provided to the restoration system 150.
[0053] As another example, the compressed data can be transmitted to a destination via a communication network (e.g., the Internet), where it is then retrieved and provided to the restoration system 150.
[0054] The input signal 102 is a data signal having one or more respective values for each of a plurality of coordinates in a coordinate space.
[0055] For example, the input signal can be an image having respective intensity values for each of one or more channels for each of a plurality of pixels in a two-dimensional coordinate space, e.g., a two-dimensional grid. Thus, each coordinate corresponds to a respective pixel of the image in the two-dimensional coordinate space, and the one or more respective values are one or more intensity values of the pixel. For example, in the case of an RGB image, the one or more respective values can include the R, G, and B values of the pixel.
[0056] As another example, the input signal can be a 3D image, where each coordinate corresponds to a respective voxel of the image in a 3D coordinate space, such as a 3D grid, and one or more respective values are one or more intensity values of the voxel. Examples of such images include computer tomography (CT) images, magnetic resonance imaging (MRI) images, ultrasonic images, X-ray images, mammogram images, fluoroscopic images, or positron emission tomography (PET) images.
[0057] As another example, the input signal can be a point cloud, where each coordinate can correspond to a respective point in a 3D coordinate space, and one or more respective values can include the respective intensity of each point. This value can optionally include other components of the points in the point cloud, such as return, second return, extension, etc. For example, the input signal can be a point cloud captured by a laser sensor, such as a LiDAR sensor.
[0058] As another example, the input signal can be a video, where each coordinate is a 3D coordinate that identifies the spatial position of a pixel within a video frame of the video, such as the x, y, t coordinates of the pixel, and one or more respective values are one or more intensity values of the pixel.
[0059] As another example, the input signal can be an audio signal, where each coordinate is a respective point in time within the audio signal, and one or more respective values include one or more values that define the amplitude of the audio signal at each point in time, such as raw amplitude values, compressed amplitude values, companded amplitude values, or compressed and companded amplitude values.
[0060] As can be seen from the above examples, the input signal can be any suitable signal sensed by one or more sensors, such as one or more sensors configured to sense a real-world environment.
[0061] As another example, the input signal can represent, for example, a signed distance function of an object or a set of one or more objects, and each of the one or more values includes the signed distance from the boundary of the corresponding object (or set of objects) of the coordinates.
[0062] As another example, the input signal can represent a rendered scene, for example, a 3D rendered scene, the coordinates represent points in the coordinate space of the scene, such as a three-dimensional coordinate system, and each value can include the density value of the scene at the point and one or more color values. The 3D rendered scene may be a 3D scene rendered based on one or more 2D images, such as one or more 2D images captured by a camera. That is, the input signal can represent the information necessary to generate a 2D image of the 3D rendered scene from any new perspective, which is different from the 2D image captured by the camera. In other words, the input signal can represent the scene in the Neural Radiance Field (NeRF) framework.
[0063] The data reconstruction neural network 110 is a neural network configured to receive an input specifying coordinates from the coordinate space of the input data signal, process the input according to network parameters, and generate, as an output, one or more predicted values of the input data signal at the specified coordinates, that is, the predicted values of one or more values of the input data signal at the specified coordinates. Therefore, the neural network 110 can be used to generate or reconstruct the data signal by providing the coordinates as an input to the neural network 110 for each coordinate of the data signal and obtaining one or more predicted values of the data signal at the coordinates as an output.
[0064] The data reconstruction neural network 110 can generally have any suitable architecture that enables the neural network 110 to map an input specifying coordinates to one or more predicted values of those coordinates.
[0065] As an example, the neural network 110 can be a multi-layer perceptron (MLP). Optionally, the MLP can be extended with positional encoding, a sine wave activation function, or both.
[0066] To compress and reconstruct data signals using the neural network 110, the compression system 100 and the reconstruction system 150 each maintain data specifying a shared value 120 of the network parameters of the data reconstruction neural network 110.
[0067] These values are determined before receiving any given data item to be compressed and are used as part of the compression and reconstruction processes of each subsequent data signal, so they are called "shared" values.
[0068] For example, the shared value 120 can be determined by training the data reconstruction neural network 110 on a set of training data that includes a plurality of different data signals, for example, by meta-learning. An example of training the data reconstruction neural network 110 will be described in more detail below with reference to FIG. 5.
[0069] When the system 100 receives a new input signal 102 that includes one or more respective values ("new values") at each of a plurality of coordinates ("new coordinates") within the coordinate space, the system 100 determines a respective update 122 for each of a subset of the network parameters.
[0070] In some implementations, the subset includes all of the network parameters.
[0071] In some other implementations, the subset is a suitable subset and includes fewer parameters than all network parameters. These implementations will be described in more detail below with reference to FIG. 3.
[0072] At a high level, system 100 generates update 122 in such a way that neural network 110 generates a more accurate reconstruction of new input signal 102 while encouraging update 122 to be sparse, i.e., such that updates 122 for most of the network parameters in the subset become zero.
[0073] The determination of update 122 for the subset of network parameters will be described in more detail below with reference to FIGS. 2 and 3.
[0074] After generating update 122, system 100 generates a compressed representation 112 of new input signal 102 that identifies each update 122 for the subset of network parameters. Since system 100 determines each update 122 in such a way that many of the updates 122 are zero, compressed representation 112 only needs to identify which parameters have non-zero updates and each update 122 for the identified parameters. This significantly reduces the compression cost while maintaining high-quality compression.
[0075] For example, system 100 can apply compression techniques to the data that identifies which parameters have non-zero updates and each update 122 for the identified parameters in order to generate compressed representation 112. The compression technique can be any suitable technique, such as entropy coding techniques like arithmetic coding or Huffman coding, or learned deep entropy coding techniques, or other suitable techniques.
[0076] As a specific example, the system 100 can apply a quantization method, such as a uniform quantization method, to non-zero updates to generate quantized updates, and then entropy code the quantized updates to generate a compressed representation 112.
[0077] Since the compressed representation only needs to identify non-zero parameter updates 122, it will be understood that the amount of data required to encode the updates 122 is already significantly reduced (i.e., without further compression) compared to the data required to represent all parameter updates. Thus, in an alternative example, the compressed representation 112 can represent the updates 122 without applying further compression techniques. That is, the updates 122 and the compressed representation 112 may be the same.
[0078] Once generated, the system 100 can store the compressed representation 112 in association with data identifying the new signal 102 for later use, or transmit the compressed representation 112 over a data communication network.
[0079] Upon receiving a request to reconstruct the input signal 102, the restoration system 150 obtains data specifying each update 122 of a subset of the network parameters determined for the input signal 102.
[0080] In particular, the system 150 can restore the compressed representation 112 using a restoration technique corresponding to the compression technique used to generate the compressed representation 112 and determine each update 122.
[0081] The system 150 then generates a reconstructed input signal 152 using the shared value 120 and each update 122. If no further compression is applied to the updates 122, the system 150 can simply receive the updates 122.
[0082] In particular, the system 150 can determine a new value 160 of a network parameter defined by a shared value 120 and respective updated values 122, for example, for each network parameter having a non-zero update 122, by combining the update 122 and the shared value 120 of the network parameter to determine the new value 160 of the network parameter. The system 150 can determine the new value 160 of a given network parameter, for example, by adding the update 122 and the shared value 120, or by subtracting the update 122 from the shared value 120.
[0083] To generate the reconstructed input signal 152, for each of a plurality of coordinates from the coordinate space of the input data signal 102, the system 150 processes an input that specifies the coordinates using the data reconstruction neural network 110 according to the new value 160 to generate a value of the reconstructed input signal 152 at the coordinates.
[0084] In some implementations, the restoration system 150 generates a "faithful" reconstruction that includes the reconstructed values of each coordinate of the input signal 102.
[0085] In some other implementations, the restoration system 150 can generate a higher-resolution reconstruction having values for each of more coordinates than the input signal 102. That is, the system 150 can perform super-resolution or in-filling to increase the resolution of the input signal 102 as part of generating the reconstruction 152.
[0086] FIG. 2 shows an example of compression and restoration of a new data signal using the data reconstruction neural network 110.
[0087] As shown in FIG. 2, the compression of the new data signal begins with a dense initialization value 210 of the network parameters of the data reconstruction neural network 110, which is also referred to herein as the "shared value" of the parameters.
[0088] These dense initialization values are called secret because there are no constraints imposed on the values that encourage a significant proportion of the values to be zero. As a specific example, these values may be determined by training the data reconstruction neural network 110 through meta-learning. This example will be described in more detail below with reference to FIG. 5.
[0089] Next, the compression system 100 performs a sparse adaptation 220 to determine a sparse update of a subset of the network parameters. The adaptation 220 generates an update of the dense initialization values 210 of the network parameters. The adaptation 220 is called "sparse" because during the adaptation, the system 100 is penalized for parameters within the subset having non-zero updates. As a result, the adaptation 220 generally ends with an update in which a large number of parameters within the subset are set to zero.
[0090] The execution of this sparse adaptation 220 to determine each update will be described in more detail below with reference to FIG. 3.
[0091] Next, the compression system 100 performs an encoding 230 to generate a compressed representation of the new data signal, such as a bitstream 232. The compressed representation contains sufficient information for the restoration system 150 to accurately reconstruct the new data signal.
[0092] In particular, since the dense initialization values 210 are shared by all data signals and the restoration system 150 also maintains the values 210 and an instance of the reconstruction neural network 110, the only information required by the restoration system 150 to reconstruct the data signal is the update determined by the sparse adaptation 220 by performing the adaptation 220.
[0093] Accordingly, compression system 100 performs encoding by applying compression techniques to the updates, generating a bitstream 232 that can be decoded by restoration system 150. Since adaptation 220 is "sparse", only a very small portion of the updates is non-zero and needs to be encoded in bitstream 232, further improving the compression ratio of the compression performed by system 100.
[0094] To reconstruct the data signal, restoration system 150 performs decoding 240 by restoring bitstream 232 to recover the updates and then combining the updates with the dense initialization value 210 (shared value) to generate a new value for the network parameters.
[0095] System 150 then performs dense reconstruction 250 according to the new value. The reconstruction is dense, thereby enabling high reconstruction quality, and its density is achieved by the dense initialization value 210 (shared value) that does not affect the compression cost per signal. Therefore, the signal-specific information is only the signal-specific updates, which are sparse, enabling a high compression ratio.
[0096] FIG. 3 is a flowchart of an exemplary process 300 for compressing a new data signal. For convenience, process 300 is described as being executed by a system of one or more computers located at one or more locations. For example, a compression system, such as compression system 100 of FIG. 1, appropriately programmed, can execute process 300.
[0097] The system maintains data specifying a shared value of the network parameters of the data reconstruction neural network (step 302).
[0098] The system then receives a new input signal including one or more respective values ("new values") at each of a plurality of coordinates ("new coordinates") in the coordinate space (step 304).
[0099] As described above, for each subset of network parameters, the system determines respective updates.
[0100] In some implementations, the subset includes all network parameters.
[0101] In some other implementations, the subset is an appropriate subset and includes fewer parameters than all network parameters.
[0102] As this specific example, one or more layers within a neural network can have network parameters that include (i) a weight tensor and (ii) a modulation tensor, where the modulation tensor is within the subset and the weight tensor is not within the subset.
[0103] For example, when processing an input, a layer having this architecture can be configured to perform operations that include computing an affine transformation between the weight tensor and the layer input to the layer, and applying the modulation tensor to the output of the affine transformation. For example, the affine transformation can be a multiplication or convolution between matrices or between a matrix and a vector. The layer can apply modulation coefficients by adding the modulation tensor to the output of the affine transformation.
[0104] Optionally, a layer having this architecture can also have a bias tensor that is not within the subset, and applying the modulation tensor to the output of the affine transformation can include applying both the modulation tensor and the bias tensor to the output of the affine transformation. For example, the layer can add both the modulation tensor and the bias tensor to the output of the affine transformation.
[0105] Thus, a layer having this architecture can adapt to each new input signal by changing only a small subset of the parameters, and the parameters not in the subset cannot be updated and thus do not need to be represented in a compressed form, so the compression cost of compression can be reduced.
[0106] To determine an update of the network parameters, the system performs one or more internal optimization iterations ("internal iterations"). In each internal iteration, the system optimizes an internal loss function such that the update promotes improving the reconstruction quality and at the same time promotes sparsity among the network parameters within the subset.
[0107] In particular, in each internal iteration, the system determines one or more sets of current values of the network parameters (step 306).
[0108] To generate a given set of current values in a given internal iteration, the system determines respective gate values for each network parameter within the subset that specify whether each update of the subset is set to zero according to a set of distribution parameters.
[0109] For example, the distribution parameters can be parameters of a continuous distribution that assigns respective probabilities to each gate. For example, the continuous distribution can be a hard concrete distribution or any other continuous distribution that allows the reparameterization trick. The system can use an appropriate estimator to sample from the continuous distribution at test time.
[0110] In the first internal iteration, the parameters can be shared distribution parameters common to all new data signals. That is, the system can maintain data specifying the shared distribution parameters and set the distribution parameters to be equal to the shared distribution parameters before the first of one or more internal iterations.
[0111] As an example, these shared distribution parameters can be determined and fixed in advance for all data signals.
[0112] As another example, the shared distribution parameters can be learned together with the shared values of the network parameters, for example, during the training of a data reconstruction neural network by meta-learning. That is, as part of meta-learning, the system effectively learns an initial sparse update configuration that has been determined to function well for the training data signals used for meta-learning.
[0113] In any subsequent internal iteration, the distribution parameters can be the distribution parameters after being updated in the previous internal iteration.
[0114] In this example, the system can determine each gate value by using the reparameterization trick, that is, sampling noise from a noise distribution and then performing a deterministic mapping from the distribution parameters and the noise to the respective probabilities of each gate, and then performing a hard rectification that maps each gate to zero (indicating that the update is set to zero) or 1 (indicating that the update is not set to zero). Performing the hard rectification means mapping a value s sampled from the distribution (or determined using a deterministic mapping) to 0 or 1 by applying a function g, such as g(s)=min(1,max(0,s)).
[0115] The system then determines a set of current values.
[0116] In particular, for any network parameter not in the subset, the system sets the current value based on the shared value of the network parameter, for example, sets it to be equal to the shared value of the network parameter.
[0117] For any network parameter within a subset where each gate value specifies that each update of the subset is set to zero, the system sets the current value based on the shared value of the network parameter, for example, sets it to be equal to the shared value.
[0118] For any network parameter within a subset where each gate value specifies that each update of the subset is not set to zero, the system sets the current value based on the shared value of the network parameter and each update of the network parameter, for example, sets it to be equal to the sum or difference of the shared value of the network parameter at the internal iteration time and each update, or to be equal to the sum or difference of the moving average of the shared value of the network parameter at the internal iteration time and each update at any previous internal iteration time.
[0119] In some implementations, each network parameter within a subset has a respective gate value that is different from each other network parameter within the subset, that is, the distribution parameter assigns probabilities to separate gate values for each of the network parameters.
[0120] In some other implementations, two or more network parameters within a subset share the same respective gate value. That is, the system imposes a sparsity pattern on the updates by requiring that a particular network parameter within the subset be jointly set to zero or jointly updated.
[0121] When multiple sets of current values are generated, the system can independently sample noise for each set of current values, that is, different sets of current values can assign updates of different network parameters within the subset to zero.
[0122] For each one or more sets of current values of network parameters, the system generates a current reconstruction of the new input signal using the set of current values of the network parameters (step 308). That is, for each set of current values and each of the new coordinates, the system processes an input specifying the new coordinates according to the set of current values of the network parameters using a data compression neural network to generate a current predicted value for the new coordinates.
[0123] In each inner iteration, the system determines, for each update of the network parameters within the subset and each of the distribution parameters of the internal loss function, the respective gradients (step 310).
[0124] In particular, the internal loss function, for each set of current values, includes (i) a reconstruction quality term that measures, for each new coordinate, the error between one or more current predicted values of the new coordinate and one or more new values of the new coordinate in the new input signal, and (ii) a differentiable sparsity term that penalizes non-zero updates for a subset of the network parameters. For example, the internal loss function can be the sum or weighted sum of the reconstruction quality term and the differentiable sparsity term. Thus, as described above, the internal loss function promotes improvement of the reconstruction quality and promotes making the updates sparse.
[0125] As an example, the reconstruction term can be a squared error term that measures, for each new coordinate, the sum or average of the squared errors, e.g., the L2 distance, between a vector of one or more current predicted values of the new coordinate and a vector of one or more new values for the new coordinate.
[0126] This system can use any of a variety of differentiable sparsity terms that penalize non-zero updates, where the larger the number of non-zero updates expected given the current distribution parameters, the larger the penalty.
[0127] As an example, a differentiable sparsity term can measure the sum of the respective probabilities for each gate value.
[0128] As described above, the respective probabilities for each gate value are defined by distribution parameters and specify the likelihood that each update of one or more network parameters corresponding to the gate value is set to a non-zero value.
[0129] For example, the probability of a given gate value can be expressed using the cumulative density function (CDF) of a continuous probability distribution. For example, given the current distribution parameters, it can be expressed as 1 minus the probability that the value s is less than or equal to zero according to the CDF.
[0130] The system then updates each respective update for a subset of network parameters and each of the distribution parameters using each gradient, for example, by applying an optimizer to the gradient (step 312).
[0131] After the last inner iteration, the system determines each (final) update of the subset of network parameters (step 314).
[0132] In particular, after the last inner iteration, the system determines the respective final gate values of each network parameter within the subset as described above according to the distribution parameters after the last inner iteration.
[0133] For any network parameter within the subset for which each respective final gate value specifies that each update of the subset is set to zero, the system sets the final update of the network parameter to zero.
[0134] For any network parameter within a subset where each final gate value specifies that the respective update of the subset is not set to zero, the system sets the respective final update of the network parameter based on the respective update of the network parameter after the last internal iteration. For example, the system can set the final update of the network parameter to be equal to the respective update of the network parameter after the last internal iteration, or set the final update of the network parameter to be equal to the moving average of the respective updates after each of one or more internal iterations.
[0135] The system then generates a compressed representation of the new input signal that identifies the respective updates for a subset of network parameters (step 316).
[0136] Due to the differentiable sparsity terms, many of the updates are zero, and in the compressed representation, it is only necessary to identify which parameters have non-zero updates and the respective updates of the identified parameters. This significantly reduces the compression cost while maintaining high-quality compression.
[0137] For example, as described above, the system can apply compression techniques to the data that identifies which parameters have non-zero updates and the respective updates of the identified parameters in order to generate the compressed representation.
[0138] FIG. 4 is a flowchart of an exemplary process 400 for generating a reconstruction of a new data signal. For convenience, process 400 is described as being executed by a system of one or more computers located in one or more locations. For example, a restoration system, such as restoration system 150 of FIG. 1, appropriately programmed, can execute process 400.
[0139] The system receives a request to reconstruct the input data signal (step 402).
[0140] The system obtains (step 404) (i) data specifying shared values of the network parameters of the data reconstruction neural network, and (ii) data specifying respective updates of a subset of the network parameters determined for the input data signal.
[0141] For example, the data specifying the shared values of the network parameters is maintained by the system and can be used when reconstructing all data signals.
[0142] The data specifying the respective updates of a subset of the network parameters may be determined by the compression system by applying a differentiable sparsity term that penalizes the updates of the non-zero subset of the network parameters while training the data reconstruction neural network to reconstruct the input signal, i.e., by training the data reconstruction neural network with the internal loss function described above.
[0143] The system generates a reconstructed input signal (step 406).
[0144] In particular, for each of a plurality of coordinates from the coordinate space of the input data signal, the system processes the input specifying the coordinate according to the values of the network parameters defined by the shared values and the respective updates to generate one or more values of the reconstructed input signal at the coordinate, using the data reconstruction neural network.
[0145] Thus, the reconstructed input signal includes, for each of the plurality of coordinates, one or more respective values generated by the coordinate data reconstruction neural network.
[0146] FIG. 5 is a flowchart of an exemplary process 500 for training a data reconstruction neural network through meta - learning to obtain a shared value of network parameters. For convenience, process 400 is described as being executed by a system of one or more computers located in one or more locations. For example, a compression system such as compression system 100 of FIG. 1, a restoration system such as restoration system 150 of FIG. 1, or a different training system appropriately programmed can execute process 500.
[0147] The system obtains a plurality of training signals (step 502).
[0148] The system trains a data reconstruction neural network with the plurality of training signals to determine a shared value of network parameters (step 504). Optionally, the system also jointly learns a shared value of distribution parameters, that is, the system trains a neural network with the plurality of training signals to determine a shared value of network parameters and shared distribution parameters.
[0149] In particular, the system trains the neural network to minimize an external loss function.
[0150] The external loss function measures the value of an internal loss function that is evaluated after performing a fixed number of internal training steps for a given set of training signals starting from a given set of shared values for a given set of shared values and given training signals.
[0151] When the system also determines the shared distribution parameters, the external loss function measures the value of the internal loss function that is evaluated after performing a fixed number of internal training steps for a given set of shared values, a given set of shared distribution parameters, and a given training signal, starting from the given set of shared values of the network parameters and the given set of shared distribution parameters for the training signal.
[0152] The system can use any of a variety of meta - learning techniques to minimize the external loss function. For example, the system can use model - agnostic meta - learning (MAML) techniques, such as directly optimizing a secondary objective by differentiating through the learning process, or by using a first - order approximation of this optimization, to update the shared values of the network parameters and, optionally, the shared distribution parameters.
[0153] Since sparsity is promoted by a differentiable sparsity term in the internal loss function, the computational cost of minimizing the external loss function is reduced as a result of computational savings achieved, for example, when the gradient of the internal loss function is zero with respect to many network parameters due to the imposed sparsity, compared to meta - learning approaches that do not employ a sparsity term in the internal loss.
[0154] Figure 6 shows an example 600 of the compression results achieved using the described technique (MSCN) compared to other approaches for three different image datasets: CelebA, SDF, and ImageNette. The results are shown in terms of the peak signal - to - noise ratio (PSNR), a metric commonly used to quantify the reconstruction quality. As can be seen from Figure 6, at a given compression rate (the proportion of remaining parameters, i.e., the parameters with non - zero updates), the described technique generally yields a higher - quality reconstruction than other approaches.
[0155] Figure 7 shows an example 700 of the compression results achieved using the described technique (MSCN) compared to another approach (Functa) for various types of data signals such as CelebA (images), ERA5 (manifolds), ShapeNet (voxel grids), and SRN Cars (rendered 3D scenes). In the example of Figure 7, the subset of network parameters is an appropriate subset and does not include all of the parameters of the neural network. The different columns of Figure 7 correspond to different sizes of the appropriate subset. As can be seen from Figure 7, at a given size of the subset, the described technique generally results in higher-quality reconstructions across various types of data signals than other approaches.
[0156] Figure 8 shows an exemplary reconstruction 800 of an exemplary image 802 generated by the described technique. In particular, Figure 8 shows exemplary reconstructions 804 generated using the described technique compared to reconstructions 806 generated using the existing technique Meta-Sparse-INR at the same sparsity rates, i.e., with different proportions of the network parameters within the subset set to zero, at various sparsity rates. In particular, Figure 8 shows reconstructions 804 and 806 at (from left to right) 90% sparsity, 95% sparsity, 97% sparsity, 98% sparsity, and 99% sparsity. As can be seen from Figure 8, the described technique consistently provides an improvement in reconstruction quality across all sparsity rates and can provide a consistent reconstruction of the original image 802 even at 99% sparsity.
[0157] Figure 9 shows an example 900 of the compression results achieved using the described technique when the values of the network parameters are sparse. That is, as described above, the system imposes no sparse constraints on the shared values of the network parameters. Thus, the current values of the network parameters used in any given internal iteration, and the new values of the network parameters used by the restoration system to reconstruct the data signal, are likely to be dense, i.e., likely to have no significant number of zero values. However, a sparse final network, i.e., a neural network with a significant number of parameters having zero values, may have desirable applications (e.g., for fast forward passes during inference when performing internal iterations or reconstructing data signals).
[0158] In such cases, the system can apply a gate value to the sum of the shared value and the update during internal iterations, i.e., make not only the update sparse but also the current values of the network parameters sparse. That is, for the parameters within a subset, if the gate value indicates setting the value of the network parameter to zero, the system sets the new value of the parameter (not just the update) to zero. Thus, the system realizes sparse current values of the network parameters within the subset instead of sparse updates.
[0159] In the example of Figure 9, the PSNRs of three different approaches in the SDF dataset are shown, such as the existing Meta-Sparse INR technique, the above-described approach with sparse updates (MSCN-δθ sparsity), and the above-described approach with sparse current values (MSCN-(θ 0 +δθ)). As can be seen from the example of Figure 9, the approach using sparse updates is improved over both alternatives at each percentage of remaining parameters, while the approach using sparse current values is consistently improved over the existing technique nonetheless. (MSCN-(θ 0In the case of the (+δθ) technique, it should be noted that the residual parameter means not only that the update of the parameter is sparse but also that the parameter has a non-zero value. Therefore, (MSCN-(θ 0 The (+δθ) technique can reduce the inference time and inference cost because there are many zero-valued parameters in the neural network by setting the current values of a large number of network parameters in the subset to zero.
[0160] This specification uses the term "configured" with respect to systems and computer program components. That one or more computer systems are configured to perform a particular operation or action means that the system has installed, during operation, software, firmware, hardware, or a combination thereof that causes the system to perform the operation or action. That one or more computer programs are configured to perform a particular operation or action means that the one or more programs include instructions that, when executed by a data processing apparatus, cause the apparatus to perform the operation or action.
[0161] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations of them. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus.
[0162] The term “data processing apparatus” refers to data processing hardware and includes any kind of apparatus, device, and machine for processing data, e.g., a programmable processor, a computer, or multiple processors or computers. The apparatus can be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations of them.
[0163] A computer program, also referred to or described as a program, software, software application, app, module, software module, script, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative languages, or procedural languages. It can be deployed in any form, such as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not necessarily correspond to a file in the file system. The program can be stored in a single file dedicated to the program or in multiple coordinated files, such as a file storing one or more modules, subprograms, or parts of the code, or as part of a file holding other programs or data, such as one or more scripts stored in a markup language document. A computer program can be deployed to execute on one computer or located at one site, or distributed across multiple sites and executed on multiple computers interconnected by a data communication network.
[0164] As used herein, the term "database" is used broadly to refer to any collection of data, and the data need not be structured in any particular way or structured at all and can be stored on a storage device at one or more locations. Thus, for example, an index database can contain multiple collections of data, each of which may be differently organized and accessed.
[0165] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more particular functions. In general, an engine is implemented as one or more software modules or components installed on one or more computers located in one or more locations. In some cases, one or more computers are dedicated to a particular engine, and in other cases, multiple engines can be installed on and executed on the same one or more computers.
[0166] The processes and logical flows described herein can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by, for example, a special purpose logic circuit such as an FPGA or ASIC, or by a combination of special purpose logic circuits and one or more programmed computers.
[0167] A computer suitable for the execution of a computer program can be based on a general-purpose microprocessor or a special-purpose microprocessor, or both, or other types of central processing units. Generally, the central processing unit receives instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a central processing unit for carrying out or executing instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special-purpose logic circuitry. Generally, a computer also includes one or more mass storage devices for storing data, such as, for example, magnetic, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices. However, a computer need not have such devices. Further, a computer can be incorporated in another device, such as, by way of several examples only, a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive.
[0168] Examples of computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as, by way of example only, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.
[0169] To provide interaction with a user, embodiments of the subject matter described herein can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can be used to provide interaction with the user. For example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be in any form, including acoustic, voice, or tactile input. Further, the computer can interact with the user by transmitting and receiving documents with the device used by the user, such as by sending a web page to a web browser on the user's device in response to a request received from the web browser. Also, the computer can interact with the user by sending a text message or other form of message to a personal device, such as a smartphone running a messaging application, and receiving a response message from the user in turn.
[0170] A data processing apparatus for implementing a machine learning model can also include, for example, a dedicated hardware accelerator unit for processing common and computationally intensive parts of machine learning training or production, i.e., inference, workloads.
[0171] The machine learning model can be implemented and deployed using a machine learning framework, such as the TensorFlow framework, or the Jax framework.
[0172] Embodiments of the subject matter described herein can be implemented in a computing system that includes, for example, backend components as a data server, or includes middleware components such as an application server, or includes frontend components such as a graphical user interface, a web browser, or an app on a client computer that a user can interact with the implementations of the subject matter described herein, or includes any combination of one or more such backend, middleware, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium, such as, for example, a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0173] The computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The relationship between a client and a server arises from computer programs that run on respective computers and have a client-server relationship with each other. In some embodiments, the server sends data, such as an HTML page, to a user device to display data to a user interacting with a device operating as a client and to receive user input from the user. For example, data generated at a user device, such as as a result of user interaction, can be received at the server from the device.
[0174] This specification includes many specific implementation details, but these are not limitations on the scope of any invention or what may be claimed, but rather are to be construed as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, features are described above as acting in certain combinations and were initially claimed as such, but in some cases, one or more features from the claimed combination can be deleted, and the claimed combination can be directed to a sub-combination or a variation of a sub-combination.
[0175] Similarly, operations are shown in the drawings and described in the claims in a particular order, but this is not to be understood as requiring that such operations be performed in the particular order or sequence shown, or that all of the operations shown be performed to achieve a desirable result. In some situations, multitasking and parallel processing may be advantageous. Further, the separation of various system modules and components in the embodiments described above is not to be understood as required in all embodiments, and the described program components and systems can generally be incorporated together into a single software product or packaged into multiple software products.
[0176] Particular embodiments of the subject matter are described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve a desirable result. As one example, the processes shown in the accompanying drawings do not necessarily require the particular or sequential order shown to achieve a desirable result. In some cases, multitasking and parallel processing may be advantageous.
Description of Symbols
[0177] 100 Compression System 102 Input Signal 110 Data Reconstruction Neural Network 112 Compressed Representation 120 Shared Value 122 Update 150 Restoration System 152 Reconstructed Input Signal 160 New Value 210 Dense Initialization Value 220 Sparse Adaptation 230 Encoding 232 Bitstream 240 Decoding 250 Dense Reconstruction 300 Process 400 Process 500 Process 800 Reconstruction 802 Image 804 Reconstruction 806 Reconstruction
Claims
1. A method performed by one or more computers, comprising: maintaining data specifying a shared value of network parameters of a data reconstruction neural network, the data reconstruction neural network being configured to receive an input specifying coordinates from a coordinate space of an input data signal, process the input according to the network parameters, and generate, as an output, one or more predicted values of the input data signal at the specified coordinates; receiving a new input signal including one or more respective new values at each of a plurality of new coordinates; determining a respective update for each of a subset of the network parameters, wherein, for each of one or more internal iterations, determining one or more sets of current values of the network parameters of the data reconstruction neural network, and for each set of current values, determining respective gate values for each network parameter within the subset that specify whether the respective update of the subset is set to zero according to a set of distribution parameters; determining the set of current values, for any network parameter not within the subset, setting the current value based on the shared value of the network parameter; for any network parameter within the subset for which the respective gate value specifies that the respective update of the subset is set to zero, setting the current value based on the shared value of the network parameter; for any network parameter within the subset for which the respective gate value specifies that the respective update of the subset is not set to zero, setting the current value based on the shared value of the network parameter and the respective update of the network parameter; thereby determining the set of current values; including; for each of the one or more sets of current values of the network parameters, For each of the new coordinates, to generate one or more current predictions for the new coordinates, using the data reconstruction neural network, according to the current set of the network parameters, a step of processing an input specifying the new coordinates including the step of for each respective update of the network parameters within the subset and for each set of current values, (i) for each new coordinate, a reconstruction quality term measuring an error between the one or more current predictions of the new coordinate and the one or more new values of the new coordinate in the new input signal, and (ii) for each of the distribution parameters of an internal loss function including a differentiable sparsity term imposing a penalty on non-zero updates for the subset of network parameters, a step of determining respective gradients for each of the subset of network parameters and the distribution parameters, a step of updating the respective updates using the respective gradients a step of generating a compressed representation of the new input signal identifying the respective updates for the subset of network parameters including the method
2. a step of storing the compressed representation in association with data identifying the new input signal The method according to claim 1, further including
3. a step of transmitting the compressed representation via a data communication network The method according to any one of claims 1 or 2, further including
4. for each of the subset of network parameters, the step of determining each respective update is after the one or more internal iterations after the one or more internal iterations, according to the distribution parameters, a step of determining respective final gate values for each network parameter within the subset for any network parameter within the subset where the respective final gate value specifies that the respective update of the subset is set to zero, a step of setting the final update of the network parameter to zero For any network parameter in the subset for which each update of the subset is not set to zero, as specified by each final gate value, setting each final update of the network parameter based on each update of the network parameter after the one or more internal iterations The method according to any one of claims 1 to 3, further comprising **Claim 5** The method according to any one of claims 1 to 4, wherein the subset of network parameters is a suitable subset of the network parameters **Claim 6** The method according to claim 5, wherein a first neural network layer in the neural network has network parameters including (i) a weight tensor and (ii) a modulation tensor, the modulation tensor being within the subset and the weight tensor not being within the subset **Claim 7** The method according to claim 6, wherein the first neural network layer is configured to perform operations including calculating an affine transformation between the weight tensor and an input to the layer, and applying the modulation tensor to an output of the affine transformation **Claim 8** The method according to claim 7, wherein the network parameters of the first neural network layer further include (iii) a bias tensor, the bias tensor not being within the subset, and applying the modulation tensor to the output of the affine transformation includes applying the modulation tensor and the bias tensor to the output of the affine transformation **Claim 9** The method according to any one of claims 1 to 4, wherein the subset includes all of the network parameters of the neural network **Claim 10** Maintaining data specifying a shared distribution parameter Before a first iteration of the one or more internal iterations, setting the distribution parameter to be equal to the shared distribution parameter The method according to any one of claims 1 to 9, further comprising **Claim 11** Training the neural network with a plurality of training signals to determine the shared value of the network parameter The method according to any one of claims 1 to 10, further comprising **Claim 12** The step of training the neural network with the plurality of training signals comprises training the neural network to minimize the internal loss function, which is evaluated after executing a fixed number of internal training steps starting from the given set of shared values of the network parameters, for a given set of shared values The method according to claim 11, comprising the step of **Claim 13** The step of training the neural network with the plurality of training signals to determine the shared values of the network parameters comprises training the neural network with the plurality of training signals to determine the shared values of the network parameters and the shared distribution parameters The method according to claim 12 when dependent on claim 10, comprising the step of **Claim 14** The step of training the neural network with the plurality of training signals comprises training the neural network to minimize the internal loss function, which is evaluated after executing a fixed number of internal training steps starting from the given set of shared values of the network parameters and the given set of shared distribution parameters, for a given set of shared values and a given set of shared distribution parameters The method according to claim 13, comprising the step of **Claim 15** Determining respective gate values for each network parameter in the subset that specify whether each respective update of the subset is set to zero according to a set of distribution parameters, comprising sampling noise from a noise distribution, and mapping the set of distribution parameters and the sampled noise to the respective gate values of the network parameters in the subset The method according to any one of claims 1 to 14, comprising the step of **Claim 16** The step of mapping the set of distribution parameters and the sampled noise to the respective gate values of the network parameters in the subset comprises applying a hard rectification to a value determined from the distribution parameters and the sampled noise The method according to claim 15, comprising the step of **Claim 17** The method according to any one of claims 1 to 16, wherein the compressed representation of the new input signal identifies only non-zero updates of network parameters within the subset of network parameters.
18. The method according to any one of claims 1 to 17, wherein each network parameter within the subset has a respective gate value that is different from each other network parameter within the subset.
19. The method according to any one of claims 1 to 17, wherein two or more network parameters within the subset share the same respective gate value.
20. The method according to any one of claims 1 to 19, wherein the differentiable sparsity term measures the sum of the respective probabilities for each gate value, and the respective probabilities for each gate value are defined by the distribution parameter and specify the likelihood that the respective update of one or more network parameters corresponding to the gate value is set to a non-zero value.
21. The method according to any one of claims 1 to 20, wherein the new input signal is an image, each coordinate corresponds to a respective pixel of the image in a two-dimensional coordinate space, and the one or more respective values include one or more intensity values of the pixel.
22. The method according to any one of claims 1 to 20, wherein the new input signal is a three-dimensional image, each coordinate corresponds to a respective voxel of the image in a three-dimensional coordinate space, and the one or more respective values include one or more intensity values of the voxel.
23. The method according to any one of claims 1 to 20, wherein the new input signal is a point cloud, each coordinate corresponds to a respective point in a three-dimensional coordinate space, and the one or more respective values include the respective intensity of the respective point.
24. The method according to any one of claims 1 to 20, wherein the new input signal is a video, each coordinate is a three-dimensional coordinate that identifies the spatial position within a video frame of a pixel from the video, and the one or more respective values include one or more intensity values of the pixel.
25. The method according to any one of claims 1 to 20, wherein the new input signal is an audio signal, each coordinate is a respective point in time within the audio signal, and each of the one or more respective values includes one or more values that define the amplitude of the audio signal at the respective point in time.
26. The method according to any one of claims 1 to 20, wherein the new input signal represents a signed distance function, and each of the one or more respective values includes the signed distance from the boundary of the object at the corresponding coordinate.
27. The method according to any one of claims 1 to 20, wherein the new input signal represents a rendered scene.
28. A method performed by one or more computers, comprising: receiving a request to reconstruct an input data signal; obtaining (i) data specifying a shared value of network parameters of the data reconstruction neural network and (ii) data specifying respective updates of the subset of network parameters determined for the input data signal, by training the data reconstruction neural network to reconstruct the input signal while applying a differentiable sparsity term that penalizes the update of a subset of non-zero network parameters; generating a reconstructed input signal, for each of a plurality of coordinates from the coordinate space of the input data signal, processing an input specifying the coordinate using the data reconstruction neural network according to values of the network parameters defined by the shared value and the respective updates to generate one or more values of the reconstructed input signal at the coordinate. including the steps including the method.
29. The method according to claim 28, wherein the reconstructed input signal has respective values for more coordinates than the input signal.
30. One or more computers, and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform respective operations of the method according to any one of claims 1 to 29. A system comprising.
31. One or more computer-readable storage media that, when executed by one or more computers, store instructions that cause the one or more computers to perform each operation of the method according to any one of claims 1 to 29. **Claim 32** One or more computer-readable storage media storing a compressed representation of a data signal, wherein the compressed representation of the data signal is generated by performing each operation according to any one of claims 1 to 29. **Claim 33** A compressed representation of a data signal, wherein the compressed representation of the data signal is generated by performing each operation according to any one of claims 1 to 29.
Citation Information
Patent Citations
Method and apparatus for content-adaptive online training in neural image compression
WO2022232842A1