Efficient implementation of dictionary-based implicit neural representation

By employing group-wise convolution and 1x1 convolution in CNNs, the computational inefficiencies of INR-based neural compression are addressed, resulting in faster and more efficient encoding and decoding processes.

WO2025153372A1PCT designated stage expired Publication Date: 2025-07-24INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/050370
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-19
Filing Date
2025-01-08
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing neural compression techniques, particularly Implicit Neural Representation (INR)-based methods, face high computational complexity due to sequential processing of multiple dictionary atoms, leading to synchronization issues and inefficient use of hardware accelerators.

Method used

Implementing a collection of multi-layer perceptron (MLP) functions using group-wise convolution and 1x1 convolution within a Convolutional Neural Network (CNN) to compute the output of multiple dictionary atoms in a single call, reducing the number of framework calls and optimizing computations for hardware accelerators.

Benefits of technology

This approach significantly reduces computational time and memory requirements by up to 13x, enhancing the practical deployment of INR-based methods in encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025050370_24072025_PF_FP_ABST
    Figure EP2025050370_24072025_PF_FP_ABST
Patent Text Reader

Abstract

A signal or a part of a signal can be encoded by approximating some layers of the INR network using a dictionary approximation. In one implementation, we propose a very efficient way to call the deep learning framework to perform the computations of the operations of the dictionary. In particular, we propose to implement a collection of multi-layer perceptron (MLP) functions using the convolutional neural network, based on group-wise convolution (also referred to as "group convolution") and 1x1 convolution. The proposed implementation is beneficial during both at encoding and decoding.
Need to check novelty before this filing date? Find Prior Art

Description

EFFICIENT IMPLEMENTATION OF DICTIONARY-BASED IMPLICIT NEURAL REPRESENTATIONTECHNICAL FIELD[1] The present embodiments generally relate to a method and an apparatus for neural compression.BACKGROUND[2] Neural compression or learning-based compression is the application of neural networks and other machine learning methods to data compression. Those techniques are currently being investigated by MPEG, and there is a new ad-hoc group which focuses on the Implicit Neural Representation-based (INR-based) compression within Working Group 4. Typically, INR-based compression techniques have a far lower computational complexity than end-to-end neural compression approaches.SUMMARY[3] According to an embodiment, a method of decoding video data representative of an image or a 3D scene is presented, comprising: obtaining a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; applying said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weighting outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combining said weighted output to form a weighted output; and decoding said region based on said weighted output.[4] According to another embodiment, a method of encoding video data representative of an image or a 3D scene is presented, comprising: obtaining a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; applying said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weighting outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combining said weighted output to form a weighted output; and reconstructing said region based on said weighted output.[5] According to another embodiment, an apparatus for decoding video data representativeof an image or a 3D scene is presented, comprising at least one memory and one or more processors, wherein said one or more processors are configured to: obtain a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; apply said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weight outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combine said weighted output to form a weighted output; and decode said region based on said weighted output.[6] According to another embodiment, an apparatus for encoding video data representative of an image or a 3D scene is presented, comprising at least one memory and one or more processors, wherein said one or more processors are configured to: obtain a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; apply said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weight outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combine said weighted output to form a weighted output; and reconstruct said region based on said weighted output.[7] One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform the encoding method or decoding method according to any of the embodiments described herein. One or more of the present embodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding video data according to the methods described herein.[8] One or more embodiments also provide a computer readable storage medium having stored thereon video data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving the video data generated according to the methods described herein.BRIEF DESCRIPTION OF THE DRAWINGS[9] FIG. 1 illustrates a block diagram of a system within which aspects of the present embodiments may be implemented.

[0010] FIG. 2 illustrates an example of a neural network used for Implicit Neural Representation (INR).

[0011] FIG. 3 illustrates a typical process to encode a signal using INR.

[0012] FIG. 4 illustrates a typical process to decode a signal using INR.

[0013] FIG. 5 illustrates a process of computing the output of the dictionary of INR functions.

[0014] FIG. 6A and FIG. 6B illustrate a process of computing the output of the dictionary ofINR functions with group convolution with the input repeated and not, respectively, according to an embodiment.

[0015] FIG. 7 illustrates calls to the deep learning framework and computations on the accelerator and CPU.

[0016] FIG. 8 illustrates calls to the deep learning framework and computations on the accelerator and CPU when group-wise convolution is used.DETAILED DESCRIPTION

[0017] FIG. 1 illustrates a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 100 may be embodied as a device including the various components described below and is configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 100 are distributed across multiple ICs and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports. In various embodiments, the system 100 is configured to implement one or more of the aspects described in this application.

[0018] The system 100 includes at least one processor 110 configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 110 may include embedded memory, input output interface, and various other circuitries as known in the art. The system 100 includes at least one memory 120 (e.g., a volatile memory device, and / or a non-volatile memory device). System 100 includes a storage device 140, which may include non-volatile memory and / or volatile memory, including, butnot limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 140 may include an internal storage device, an attached storage device, and / or a network accessible storage device, as non-limiting examples.

[0019] System 100 includes an encoder / decoder module 130 configured, for example, to process data to provide an encoded video or decoded video, and the encoder / decoder module 130 may include its own processor and memory. The encoder / decoder module 130 represents module(s) that may be included in a device to perform the encoding and / or decoding functions. As is known, a device may include one or both of the encoding and decoding modules. Additionally, encoder / decoder module 130 may be implemented as a separate element of system 100 or may be incorporated within processor 110 as a combination of hardware and software as known to those skilled in the art.

[0020] Program code to be loaded onto processor 110 or encoder / decoder 130 to perform the various aspects described in this application may be stored in storage device 140 and subsequently loaded onto memory 120 for execution by processor 110. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 may store one or more of various items during the performance of the processes described in this application. Such stored items may include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.

[0021] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing that is needed during encoding or decoding. In other embodiments, however, a memory external to the processing device (for example, the processing device may be either the processor 110 or the encoder / decoder module 130) is used for one or more of these functions. The external memory may be the memory 120 and / or the storage device 140, for example, a dynamic volatile memory and / or a non-volatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations, such as for MPEG-2, JPEG Pleno, MPEG-I, HEVC, VVC or MPEG VCM.

[0022] The input to the elements of system 100 may be provided through various input devicesas indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0023] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion may be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select (for example) a signal frequency band which may be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements to perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion may include a tuner that performs various of these functions, including, for example, down converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to baseband. In one set-top box embodiment, the RF portion and its associated input processing element receives an RF signal transmitted over a wired (for example, cable) medium, and performs frequency selection by filtering, down converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Adding elements may include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0024] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed- Solomon error correction, may be implemented, for example, within a separate input processing IC or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing may be implemented within separate interface ICs or within processor 110 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 110, and encoder / decoder 130 operating in combination with the memory and storage elements to process the datastream asnecessary for presentation on an output device.

[0025] Various elements of system 100 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and transmit data therebetween using suitable connection arrangement 115, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.

[0026] The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 may include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 may include, but is not limited to, a modem or network card and the communication channel 190 may be implemented, for example, within a wired and / or a wireless medium.

[0027] Data is streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802. 11. The Wi-Fi signal of these embodiments is received over the communications channel 190 and the communications interface 150 which are adapted for WiFi communications. The communications channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. Other embodiments provide streamed data to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105. Still other embodiments provide streamed data to the system 100 using the RF connection of the input block 105.

[0028] The system 100 may provide an output signal to various output devices, including a display 165, speakers 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display 165, speakers 175, or other peripheral devices 185 using signaling such as AV. Link, CEC, or other communications protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices may be connected to system 100 using the communications channel 190 via the communications interface 150. The display 165 and speakers 175 may be integrated in a single unit with the other components of system 100 in an electronic device, forexample, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.

[0029] The display 165 and speaker 175 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display 165 and speakers 175 are external components, the output signal may be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0030] FIG. 2 illustrates a simple neural network used for implicit neural representation (INR). Such a neural network used for INR can be referred to as an INR network. For clarity, we use for illustration a 2D signal such as an image, but INR can be used for signals of any dimension. INR parameterizes a signal as a function (200), which takes coordinates (210) as input and outputs approximated values (260) of a signal at these coordinates. INR has recently been applied to images, 2D videos or 3D objects among other applications. In the image case, the inputs (210) can be pixel coordinates (x, y) and the INR may output (260) the color values (r, g, b) or (y, u, v) of the input pixels. In the video case, the input can include the frame index t in addition to pixel coordinates.

[0031] The input coordinates may be modified by a transformation before being used as input for the neural network. This transformation can be a Fourier mapping, coordinate transformation, normalization etc. In this document, we will illustrate our methods using the Fourier mapping, where an input point with (%, y) pixel coordinates is mapped into a higher dimensional feature space before being passed through the network: y(v) = [cos(2nBv), sm(2'n:Bv)]T(1) where B is a random Gaussian matrix, whose each entry is drawn independently from a normal distribution N(0, o2).

[0032] The INR can be used to reconstruct a signal by computing the signal values for every necessary coordinate inputs. It can be used to upsample a signal by generating output for input coordinates corresponding to the upsampled pixels, for example the mean of the coordinates between two consecutive pixels for upsampling by a factor of 2.

[0033] An INR network (200) is typically a neural network composed of multiple neural layers, such as fully connected layers. In FIG. 2, the network has four neural layers (220, 230, 240, 250). Intermediate outputs (e.g., Nij, ithlayer and jthneuron) are represented by circles. Each neural layer can be described as a function that first multiplies the input by a tensor (e.g., SI-S4), adds a vector called the bias and then applies a nonlinear function on the resulting values (Nij). In this document, we may also refer to “neural layer” simply as “layer.” The shape (and other characteristics) of the tensor and the type of non-linear functions are called the architecture of the network. We will denote the values of the tensor and the bias by the term “weights.” The weights and, if applicable, the parameters of the non-linear functions, are called the parameters 6 of the network. The architecture and the parameters define a “model.” We will use fgto denote an INR function parameterized by 6.

[0034] FIG. 3 illustrates a typical process (300) to encode a signal (310) using INR. This is done by optimizing the parameters 6 (or a subset of them) of the INR network, by an INR- based encoder (320), to reconstruct the signal. The parameters 6 are quantized (330) with any general quantization scheme and entropy encoded (340) to create the output bitstream (350). For an image I of size (M x A), the parameters 0 or (the chosen subset) can for example be optimized by minimizing the following loss function:where D is a distortion which quantifies the difference between the image predicted (reconstructed) by fgand the original image / , R is the bitrate of the encoded parameters and A a trade-off parameter between D and R, D could be any differentiable distortion measure, such as mean squared error as shown in Eq. (3), and M and N are the width and height of an image, respectively. Other metrics such as LPIPS (Learned Perceptual Image Patch Similarity) can also be used in this case. The optimization of the parameters 6 is typically performed by a machine learning approach such as a batch / stochastic gradient descent method.

[0035] FIG. 4 illustrates a typical process (400) to decode a bitstream (410) compressed using INR. First the parameters are entropy decoded (420) from the bitstream. The parameters are then dequantized (430) and used by an INR-based decoder (440) to reconstruct the signal (450).

[0036] To decompress the signal, fgis evaluated at all relevant coordinates. These coordinates can be selected at decoding. A typical choice would be all pixel coordinates for an image or video. As an example, for a 256x256 pixel image, these coordinates could be all pairs (x, y) for all x e {0,1, ... ,255} and y G {0,1, ... ,255}. Other choices are possible, for example to upsample, downsample or extend the original image.

[0037] In a commonly owned EP Patent Application No. 23306188.6, entitled “Approximating Implicit Neural Representation through Learnt Dictionary Atoms” (our Attorney Docket No.2023PF00497), we have described how a signal or a part of a signal can be better encoded by approximating some parameters of the INR network using a dictionary approximation. By doing so, the INR network can be encoded by the weights of the non-approximated parts and some additional information that describes the approximated parts. In the remaining of the disclosure, we will take as an example an INR divided into head h and tail t layers, as f0= t0t. h0h, and approximate 9h. This is an example of a decomposition into a composite representation, and other types of decompositions could be used without any loss of generality. For the approximation, we proposed to leam a dictionary D = [d1, d2, ... dK] with k INR atoms so that each set of head layers is approximated (represented) using a sparse linear combination of the atoms of the dictionary. That is,

[0038] The sparse coefficients y are for example approximated by optimizing the following loss function, Hylli, (5)where 0his the weights of the head layers, D is the leamt dictionary, and y = [yn... , yk] is the sparse coefficients to be optimized. To enforce the sparsity in the coefficient vector, the LI norm is used and a is the trade-off between the two terms in the equation.

[0039] In a commonly owned EP Patent Application No. 23306514.3, entitled “Dictionary- driven Implicit Neural Representation for Image and Video Compression” (our Attorney Docket No. 2023PF00641), we have described how a signal or a part of a signal can be encoded by approximating some layers of the INR network using a dictionary approximation. This also allows encoding the INR network by the weights of the non-approximated parts and some additional information that describes the approximated parts. In the remaining of the disclosure, we will take as an example an INR divided into head h and tail t layers, as f0= t0. h0h, and approximate h0h. This is an example of a decomposition into a composite representation, and other types of decompositions could be used without any loss of generality. For the approximation, we proposed to leam a dictionary D = [f(i^fd2' ■■■ fdK\ with k INR functions so that the set of head layers is approximated (represented) using a sparse linear combination of the atoms of the dictionary. That is,

[0040] Such an approximation can for example be achieved by the optimization of thefollowing cost:where B is the spatial support of the INR approximation (all image, block, superpixel, etc). This optimization problem may also be modified to optimize the dictionary D and / or the weights 6t. It may also include additional losses, such as a f or l2losses on some or all of the optimized parameters.

[0041] The dictionary approximation part in equations (6) and (7), which contains a dictionary D with K INR functions can be illustrated in FIG. 5. In particular, Al, ... , AK are k INR functions in the dictionary which take same input (511, 512, 513) and outputs the feature maps (521, 522, 523), and these feature maps are combined using sparse coefficients (yn... , yfe). The combined feature maps (530) are transformed into the pixel intensity values (540).

[0042] Its implementation is performed as follows, where x and y represent the input and output, respectively. y = 0, for i = 1 to K y = y + fdtM * Yt end for

[0043] Thus, for the dictionary size of K, we need to perform K forward passes or to loop for K times over the K INR functions (atoms).

[0044] Processing deep neural networks is typically done using a deep learning accelerator, a hardware that is specialized for deep learning operations. We will call it an “accelerator.” Examples include graphical processing units (GPUs) or tensor processing units (TPUs). To use such a piece of hardware, the operations of the networks are typically described using a high-level framework such as Py Torch, TensorFlow or JAX. These frameworks then interface with lower-level libraries such as cuDNN and launch these neural operations on the specialized hardware. When performing the forward pass of a model (inferring output from input), the description of the operations may be iteratively analyzed by the CPU of the computer, sent to the accelerator through this process, their output is received by the CPU which moves to the next operation. Using this process, the CPU waits repeatedly to receive the output of the accelerator and makes many calls to the framework.

[0045] Referring back to FIG. 2, the INR functions are typically implemented using fullyconnected layers or multi-layer perceptron (MLP). In general, the computation of K atoms of INR functions is processed sequentially one by one, and at the end the results are aggregated with mixture coefficients. This involves multiple calls to the deep learning frameworks and thus heavy computation time as the number of K atoms increases. In order to obtain higher reconstruction quality, the number of atoms in the dictionary should be large, this results in heavy computation time.

[0046] These sequential computations can be parallelized when several accelerators are available, in this case the master accelerator can perform its own operation and wait for the output of the remaining accelerators to combine the output of all the accelerators for subsequent computations, in this case there might be a synchronization issue between the master and other accelerators, and there could be packet loss when there is a communication between the master and remaining nodes. Further, if there is some operation which requires gradient based optimization before the INR dictionary functions, then the gradients of the remaining accelerators need to be transferred, and also the graph needs to be tracked properly for back- propagation. These scenarios are not easy to handle with several accelerators. Further, if the INR functions in the dictionary is small MLP, then there could be more latency with several accelerators (might take more time in communication) than performing operations on a single accelerator. Lastly, if the sequential computation is parallelized on a single accelerator, then CPU of the computer sends the call to the accelerator and waits to receive the outputs of the several calls, again there could be an issue of synchronization, CPU scheduling the calls, gradients transfer and properly tracking the graph for back-propagation. The complexity of these issues could also vary with the high-level deep learning library framework.

[0047] In this document, we propose a very efficient way to call the deep learning framework to perform the computations of the operations of a dictionary (equations (6) and (7)) to lower the memory consumption and to reduce the number of calls to the framework, thus increasing the computational speed, which is very important from the practical deployment perspective. Furthermore, since we use fewer operations, the framework and / or lower-level library may be able to better optimize these operations for the accelerator. For example, by leveraging vector or tensor operations, multiple arithmetic operations can be performed in the same clock cycle. As another example, by using matrix operations rather than multiple scalar or vector operations, it may be possible to use algorithms with a lower number of operations. The proposed implementation is beneficial during both at encoding and decoding. For example, at the decoder side, the signal is reconstructed with the K INR functions and sparse coefficients asI x) =ar|d it can be implemented with the proposed fast implementation.

[0048] In one embodiment, we propose to compute the output of one layer of the K atoms of INR functions in a single call without sequential processing by leveraging two properties in the neural network architecture. Our proposal has a similar number of trainable parameters with the existing solution, but an efficient implementation. Therefore, there will be no discrepancy in the reconstruction quality. In particular, we propose to implement a collection of multi-layer perceptron (MLP) functions using the convolutional neural network: group-wise convolution (also referred to as “group convolution”) and 1x1 convolution.

[0049] Let us assume that a collection of INR functions is known by the encoder and decoder. These functions may be chosen randomly, learned on the first frame of a video sequence, or learned using a large database of images. Let us refer to this basis of INR functions as a collection of k functions parametrized by the vectors 01(.... 9k. leading to the corresponding functions F = {fe, ... , fdk] . One may interpret these functions as basis functions that ideally decompose any input signal as orthogonal basis. In the following, the functions will be referred to as atoms, dictionary components, or basis functions. All these terminologies refer to the same functions.

[0050] Each function in F is a multi-layer perceptron, and the output of the dictionary of INR functions are computed sequentially as shown in FIG. 5. Let us consider the INR functions in F consists of two layers (more generally, can be any number of layers) with non-linear activation function g and with hidden dimension of m, output dimension of o (here let us assume o = m), and input dimension of p. For example, each of the INR functions (fei,fd2> ■ ■ fdk) inF,canbedefined as follows in the PyTorch deep learning library. Here, Linear() represents MLP with a single layer, where in_features indicates the input dimension, out features indicates the output dimension, and “bias” indicates whether a linear layer includes bias or not, and g() represents the non-linear activation function. Both linear layer and g() correspond to calls to the framework and to neural operations that will be processed on the accelerator. {nn. Linear(in_features = p, out_features = m, bias = False) g() nn. Linear(in_features = m, out_features = m, bias = False) g(){nn. Linear(in_features = p, out_features = m, bias = False) g() nn. Linear(in_features = m, out_features = m, bias = False) g(){nn. Linear(in_features = p, out_features = m, bias = False) g() nn. Linear(in_features = m, out_features = m, bias = False) g()

[0051] FIG. 7 illustrates the resulting calls to the deep learning framework and computations on the accelerator (700) and CPU (710) for k atoms of two layers. The computation is first initialized, and the first operation is called (720). This may involve steps such as reserving memory space for the tensors or initializing the parameters. Calls to the framework and accelerator are not displayed for this step. The linear layers of the first atom are successively computed on the accelerator (721, 725), interleaved with the non-linear functions (723, 727). After each computation, the accelerator may wait on the CPU, for example to return the results to it or to request the next operation (722, 724, 726, 728). Then the second atom is processed in the same way: linear layer functions (731, 735) interleaved with the non-linear functions (733, 737) and waits for the CPU (732, 734, 736, 738). Subsequent atoms are then processed (not shown on the picture) until the last one: linear layer functions (741, 745) interleaved with the non-linear functions (743, 747) and waits for the CPU (742, 744, 746, 748).

[0052] We propose an alternative and efficient way of implementing the above-mentioned description by having only one INR function, and this INR function contains all the atoms in the INR dictionary functions by leveraging two properties in the convolution neural network (CNN). Therefore, we need only one forward pass to compute the output of the K dictionary functions.

[0053] MLP as CNN: More specifically, the first property is that MLP can be implemented using 2D convolutional neural network (CNN) using a filter size (kernel size) of 1 (1x1 convolution). For example, in PyTorch, MLP with a single layer is represented as nn.Linear(in_features = 3, out_features = 4), and it can be implemented using CNN as follows: nn.conv2d(in_channels = 3,out_channels = 4, kemel_size = 1) where conv2d() represents 2D CNN, in_channels indicates the number of input channels, out channels indicates the number of output channels, kemel size indicates the height andwidth of the convolving kernel (h, w), and it can include bias also.

[0054] It is important to note that the reshaping of the input data is required here when implementing MLP using 2D CNN. The MLP takes input data in the form of N x D, and the CNN takes input data in the form of 1 x D x N x 1, where N is the number of samples, and D is the number of features (channels) dimension.

[0055] Group convolution in CNN: The second property we leverage is the group convolution in the convolutional neural network. The group convolution performs convolution operation only for a group (chunk) of input channels, and each output channel only depends on the input channels within the group of channels. By this, the convolution kernel in the group convolution looks only the certain part of the input channels and neglects the remaining part of the input channels. That is, if the input has S channels, and the group is 2, then in the first group, convolution filters will perform convolution operation only on the first S / 2 channels of the input and in the second group, convolution filters will perform convolution operation only on remaining S / 2 channels of the input. In the fully connected layers it is not possible to group the dictionary atoms.

[0056] By using these two features, we propose a more efficient implementation for performing the forward pass of a dictionary aware training for the INR using deep learning frameworks. The k atoms of INR functions {fe, ... , fdk] in the F is implemented in one single c {omposite function as follows: nn. conv2d(in_channels = p * k, out_channels = m * k, kernel_size = 1, groups = k) g() nn. conv2d(in_channels = m * k, out_channels = m * k, kernel_size = 1, groups = k) g()

[0057] That is, as shown in FIG. 6A where k = 3 in this example, the input (610) is the pixel co-ordinates or Fourier mapping of the pixel co-ordinates, the first layer (620) is implemented using a CNN with input dimension in_channels = p*k (* denotes multiplication here, the input channels are replicated k times, thus input dimension is p * k, 610 ), output dimension out_channels = m*k, where this CNN performs group-wise convolution in k groups (620) with a kernel size kemel size = 1, followed by a non-linear activation function g(). Here Gl / Al indicates, group 1 / atom function 1, and this depends only on the Gl / Al partition of the input dimension. That is, Gl / Al 1x1 filter only operates on the input dimension partition Gl / Al. Likewise for remaining G2 / A2 and G3 / A3. The second layer (640) is implemented using a CNN on the output of the first layer (630) with input and output dimensions in_channels =out channels = m*k (number of output channels can be different from input channels), where this CNN also performs group-wise convolution in k groups with a kernel size kemel_size = 1, followed by a non-linear activation function g(). In the output of the second layer (650), Gl / Al belongs to the output of the first atom in the dictionary, G2 / A2 belongs to the output of the second atom in the dictionary, and so on. The sparse combination of the output of each atom (basis function) are performed (660) with linear mixture coefficients (not explicitly shown in 660). In an alternative embodiment, if some coefficients of the mixture are zero (or negligible), these coefficients and the associated dictionary atoms can be discarded and not included in the computation. Then the combined output is converted to the RGB output (680) with a MLP linear layer with three neurons (670).

[0058] Note that in this architecture, the mathematical operations remain exactly the same as in FIG. 5. Therefore, our proposal will have similar reconstruction quality.

[0059] Thus, instead of k forward (sequential) passes, each forward pass calling the linear() and g() operations I times, where I is the number of layers, we group computations from all the k atoms of INR functions in one call to the deep learning framework. Hence, only 21 such calls are required rather than 2lk to generate the output (650), thus reducing the computational time by a significant factor (13x according to our experiments). The first m number of dimensions of the output (650) belongs to the first atoms, and second m number of dimensions belong to the second atoms, and so on. To combine the output of all the INR functions in the dictionary, we use (660) the sparse linear combination of outputs of the atoms.

[0060] FIG. 8 illustrates the resulting calls to the deep learning framework and computations on the accelerator (800) and CPU (810) for k atoms of two layers. The computation is first initialized, and the first operation is called (820). As before, this may involve steps such as reserving memory space for the tensors or initializing the parameters. It may also involve reshaping or grouping some parameters. Calls to the framework and accelerator are not displayed for this step. The first linear layer of all atoms is computed on the accelerator (821), followed by a non-linear function (823). The second layer of all atoms (825) is then computed, followed by a non-linear function (827). After each computation, the accelerator may wait on the CPU, for example to return the results to it or to request the next operation (822, 824, 826, 828). The number of calls is thus much lower and the benefits outlined above apply.

[0061] Alternatively, F can be implemented as shown in FIG. 6B with the first layer as the normal convolution (because the input is same to all the atoms in the dictionary) and remainingl {ayers as the group convolution. nn. conv2d(in_channels = p, out_channels = m * k, kernel_size = 1) g() nn. conv2d(in_channels = m * k, out_channels = m * k, kernel_size = 1, groups = k) g()

[0062] As shown in FIG. 6B, the first convolution layer performs the operations in the first layer of all the atoms in one function call with input dimension in channels = p (610b), and output dimension out_channels = m*k (here * indicates the multiplication, 620b) using CNN with a kernel size of 1, followed by the non-linear activation function g(). Note that here the input dimensions are not repeated since the input is same to all the atoms in the dictionary. Here the first layer filters (620b) Al, A2, ..., operate on the entire channels of the input as in the convolution. From the second layer onwards, the group convolution is performed, until the last layer. The second layer (640b) is implemented using a CNN on the input (output of the first layer CNN 630b) with input and output dimensions m*k, where this CNN performs group- wise convolution in k groups with a kernel size of 1, followed by a non-linear activation function g(). This performs all the operations of the second layer of all the atoms in the dictionary in a one call.

[0063] To evaluate the benefits of our proposed implementation, we performed experiments on the dictionary driven training of INR and the computational time is reported in Table 1, which shows that our proposed implementation is 13x faster. Further, we perform the GPU memory analysis when all the INR dictionary functions and its computation are transferred to GPU, as illustrated in Table 2. Our proposed implementation also resulted in lower memory requirements; this shows the implementation with looping is inefficient also in memory usage.Table 1Table 2

[0064] In the encoding / decoding device, if there is specialized hardware accelerator or software accelerator which is optimized for the CNN operations, then our proposed implementation method would be highly efficient.

[0065] The proposed implementation is not specific to dictionary based INR representation. Rather, our proposal can be used on any dictionary based neural network methods, and mixture of expert models (Ensemble of neural networks).

[0066] Various numeric values are used in the present application. The specific values are for example purposes and the aspects described are not limited to these specific values.

[0067] Various methods are described herein, and each of the methods comprises one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., such as, for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an ordering to the modified operations unless specifically required. So, in this example, the first decoding need not be performed before the second decoding, and may occur, for example, before, during, or in an overlapping time period with the second decoding.

[0068] The implementations and aspects described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (for example, discussed only as a method), the implementation of features discussed may also be implemented in other forms (for example, an apparatus or program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. The methods may be implemented in, for example, an apparatus, for example, a processor, which refers to processing devices in general, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, for example, computers, cell phones, portable / personal digital assistants (“PDAs”), and other devices that facilitate communication of information between end-users.

[0069] Reference to “one embodiment” or “an embodiment” or “one implementation” or “an implementation”, as well as other variations thereof, means that a particular feature, structure, characteristic, and so forth described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrase “in one embodiment” or “in an embodiment” or “in one implementation” or “in an implementation”, as well any other variations, appearing in various places throughout this application are not necessarily all referring to the same embodiment.

[0070] Additionally, this application may refer to “determining” various pieces of information. Determining the information may include one or more of, for example, estimating the information, calculating the information, predicting the information, or retrieving the information from memory.

[0071] Further, this application may refer to “accessing” various pieces of information. Accessing the information may include one or more of, for example, receiving the information, retrieving the information (for example, from memory), storing the information, moving the information, copying the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0072] Additionally, this application may refer to “receiving” various pieces of information. Receiving is, as with “accessing”, intended to be a broad term. Receiving the information may include one or more of, for example, accessing the information, or retrieving the information (for example, from memory). Further, “receiving” is typically involved, in one way or another, during operations, for example, storing the information, processing the information, transmitting the information, moving the information, copying the information, erasing the information, calculating the information, determining the information, predicting the information, or estimating the information.

[0073] It is to be appreciated that the use of any of the following“and / or”, and “at least one of’, for example, in the cases of “A / B”, “A and / or B” and “at least one of A and B”, is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of both options (A and B). As a further example, in the cases of “A, B, and / or C” and “at least one of A, B, and C”, such phrasing is intended to encompass the selection of the first listed option (A) only, or the selection of the second listed option (B) only, or the selection of the third listed option (C) only, or the selection of the first and the second listed options (A and B) only, or the selection of the first and third listed options (A and C) only, or the selection of the second and third listed options (B and C) only, or the selection of all three options (A and B and C). This may be extended, as is clear to one of ordinary skill in this and related arts, for as many items as are listed.

[0074] As will be evident to one of ordinary skill in the art, implementations may produce a variety of signals formatted to carry information that may be, for example, stored or transmitted. The information may include, for example, instructions for performing a method, or data produced by one of the described implementations. For example, a signal may be formatted tocarry the bitstream of a described embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (for example, using a radio frequency portion of spectrum) or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links, as is known. The signal may be stored on a processor-readable medium.

Claims

CLAIMS1. A method of decoding video data representative of an image or a 3D scene, comprising: obtaining a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; applying said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weighting outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combining said weighted output to form a weighted output; and decoding said region based on said weighted output.

2. A method of encoding video data representative of an image or a 3D scene, comprising: obtaining a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; applying said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weighting outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combining said weighted output to form a weighted output; and reconstructing said region based on said weighted output.

3. An apparatus for decoding video data representative of an image or a 3D scene, comprising at least one memory and one or more processors, wherein said one or more processors are configured to: obtain a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; apply said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weight outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combine said weighted output to form a weighted output; anddecode said region based on said weighted output.

4. An apparatus for encoding video data representative of an image or a 3D scene, comprising at least one memory and one or more processors, wherein said one or more processors are configured to: obtain a plurality of coefficients corresponding to a plurality of basis functions for a region of said image or 3D scene; apply said plurality of basis functions based on a CNN (Convolutional Neural Network), wherein a layer of said CNN is implemented by group-wise convolution; weight outputs corresponding to said plurality of basis functions by said plurality of coefficients, respectively; combine said weighted output to form a weighted output; and reconstruct said region based on said weighted output.

5. The method of claim 1 or 2, or the apparatus of claim 3 or 4, wherein a kernel size of said group-wise convolution is 1.

6. The method of any one of claims 1, 2 and 5, or the apparatus of any one of claims 3-5, wherein said basis functions correspond to INR (Implicit Neural Representation) basis functions.

7. The method of any one of claims 1, 2, 5 and 6, or the apparatus of any one of claims 3-6, wherein K groups are used in said group-wise convolution, and wherein input to said CNN is based on coordinates of said video data, repeated by K times.

8. The method of claim 7, or the apparatus of claim 7, wherein an input dimension and output dimension of a basis function are p and m, respectively, and an input dimension and output dimension of said CNN are p*K and m*K, respectively.

9. The method of any one of claims 1, 2, 5 and 6, or the apparatus of any one of claims 3-6, wherein group-wise convolution is applied starting from the second layer of said CNN.

10. The method of claim 9, or the apparatus of claim 9, wherein a convolution is applied to the first layer of said CNN, wherein 1x1 convolution operates on all channels of input to said CNN.

11. The method of any one of claims 1, 2 and 5-10, or the apparatus of any one of claims 3-10, wherein said plurality of basis function represent head layers of an INR network, wherein a neural network corresponding to tail layers of said INR network is applied to said INR network.

12. The method of any one of claims 1, 2 and 5-11, or the apparatus of any one of claims 3-11, wherein said plurality of coefficients are sparse.

13. The method of any one of claims 1, 2 and 5-12, or the apparatus of any one of claims 3-12, wherein a set of basis functions corresponding to non-zero coefficients is selected from said plurality of basis functions, wherein a basis function is only applied if a corresponding coefficient is non-zero or non-negligible.

14. A signal comprising video data representative of an image or a 3D scene, formed by performing the method of any one of claims 2 and 5-13.

15. A computer readable storage medium having stored thereon instructions for encoding or decoding video data according to the method of any one of claims 1, 2 and 5-13.

Citation Information

Patent Citations

  • Implicit image and video compression using machine learning systems

    US20220385907A1

  • EP23306514A