Video compression using an implicit neural representation network

EP4744287A1Pending Publication Date: 2026-05-20INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
INTERDIGITAL CE PATENT HOLDINGS SAS
Filing Date
2024-06-25
Publication Date
2026-05-20

AI Technical Summary

Technical Problem

The existing implicit neural representation (INR) networks for video compression face limitations when increasing the dimensionality of Fourier mapping, leading to a large number of network parameters that hinder efficient image encoding and decoding, particularly in terms of bitrate and reconstruction quality.

Method used

The use of a cosine mapping instead of the conventional sine-cosine mapping in INR networks for video compression, which allows for twice the number of sampled frequencies to generate Fourier bases without increasing the complexity of encoding and decoding processes, thereby improving image reconstruction quality and compression efficiency.

Benefits of technology

The cosine mapping approach results in better rate-distortion performance, achieving about 6 dB PSNR improvement for the same bitrate and providing significant bitrate savings without adding complexity to the encoding and decoding processes, as demonstrated by experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024067757_16012025_PF_FP_ABST
    Figure EP2024067757_16012025_PF_FP_ABST
Patent Text Reader

Abstract

Apparatuses and methods are disclosed including techniques for encoding and decoding image regions of a video using implicit neural representations (INRs). Techniques disclosed provide forthe coding of an image region, including training an INR network using a cosine mapping. The INR network is trained to predict a pixel value of the image region based on respective pixel coordinates, where the training determines parameters of the INR network. The determined parameters of the INR network are then coded into a bitstream. Techniques disclosed also provide for the decoding of the image region from the bitstream. The decoding of the image region includes decoding from the bitstream the parameters of the INR network and reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] VIDEO COMPRESSION USING AN IMPLICIT NEURAL REPRESENTATION NETWORK

[0002] CROSS REFERENCE TO RELATED APPLICATIONS

[0003]

[0001] This application claims the benefit of European Application No. 23306179.5, filed on July 10, 2023, which is incorporated herein by reference in its entirety.

[0004] BACKGROUND

[0005] [2] An implicit neural representation (INR) network is a neural network that is trained to represent a specific image; the INR network is trained to predict a pixel value of the image when presented with the respective pixel coordinates. To improve the INR network’s capability to learn high frequency patterns in the image, a Fourier mapping can be used to map the pixel coordinates, generating therefrom features that are in turn fed into the INR network. Conventionally, such Fourier mapping consists of Fourier bases that are generated based on randomly sampled frequencies. The more frequency samples used (more Fourier bases), the better the INR network’s capability to learn the high frequency patterns in the image. However, increasing the dimensionality of the Fourier mapping (that is, increasing the number of input signals (features) the INR network has to process) is a limiting approach when the INR network is applied to image compression. This is because when the INR network is applied to encode an image, a large INR network translates to a large number of network parameters that need to be encoded into the bitstream.

[0006] SUMMARY

[0007] [3] Aspects disclosed in the present disclosure describe methods for encoding video data. The methods include encoding, into a bitstream, an image region from a video frame of the video data. The coding of the image region comprises training an INR network using a cosine mapping. The INR network is trained to predict a pixel value of the image region based on respective pixel coordinates, where the training determines the parameters of the INR network. The determined parameters of the INR network are then coded into the bitstream. Aspects disclosed in the present disclosure also describe methods for decoding video data. The methods include decoding, from a bitstream, an image region from a video frame of the video data. The decoding of the image region comprises decoding, from the bitstream, parameters of an INR network. The INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, where the training of the INR network determines the parameters that are decoded from the bitstream. The methods further comprise reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.

[0008] [4] Aspects disclosed in the present disclosure describe an apparatus for encoding video data. The apparatus comprises at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, can cause the apparatus to encode, into a bitstream, an image region from a video frame of the video data. The coding of the image region comprises training an INR network using a cosine mapping. The INR network is trained to predict a pixel value of the image region based on respective pixel coordinates, where the training determines the parameters of the INR network. The determined parameters of the INR network are then coded into the bitstream. Aspects disclosed in the present disclosure also describe an apparatus for decoding video data. The apparatus comprises at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, can cause the apparatus to decode, from a bitstream, an image region from a video frame of the video data. The decoding of the image region comprises decoding, from the bitstream, parameters of an INR network. The INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, where the training of the INR network determines the parameters that are decoded from the bitstream. The decoding further comprises reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.

[0009] [5] Further aspects disclosed in the present disclosure describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for encoding video data. The methods include encoding, into a bitstream, an image region from a video frame of the video data. The coding of the image region comprises training an INR network using a cosine mapping. The INR network is trained to predict a pixel value of the image region based on respective pixel coordinates, where the training determines the parameters of the INR network. The determined parameters of the INR network are then coded into the bitstream. Aspects disclosed in the present disclosure also describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for decoding video data. The methods include decoding, from a bitstream, an image region from a video frame of the video data. The decoding of the image region comprises decoding, from the bitstream, parameters of an INR network. The INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, where the training of the INR network determines the parameters that are decoded from the bitstream. The methods further comprise reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping. [6] This Summary is provided to introduce a selection of concepts in a simplified form that is further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure.

[0010] BRIEF DESCRIPTION OF THE DRAWINGS

[0011] [7] FIG. 1 is a block diagram of an example system, according to which aspects of the present embodiments can be implemented.

[0012] [8] FIG. 2 is a diagram illustrating an example INR network, according to which aspects of the present embodiments can be implemented.

[0013] [9] FIG. 3 is a block diagram of an example video encoder, according to which aspects of the present embodiments can be implemented.

[0014]

[0010] FIG. 4 is a block diagram of an example video decoder, according to which aspects of the present embodiments can be implemented.

[0015]

[0011] FIG. 5 is a diagram illustrating an example of an augmented INR network, according to which aspects of the present embodiments can be implemented.

[0016]

[0012] FIG. 6 is a flowchart of an example method for encoding video data, according to which aspects of the present embodiments can be implemented.

[0017]

[0013] FIG. 7 is a flowchart of an example method for decoding video data, according to which aspects of the present embodiments can be implemented.

[0018] DETAILED DESCRIPTION

[0019]

[0014] Representing data via INR networks is a relatively new technology that has only recently been investigated by the computer vision and computer graphics communities. INR networks are studied for compression where an efficient representation of images (as well as videos, surfaces, and volumes) is required. Neural networks, namely, autoencoders, have already been used to encode images; however, these networks are generically trained and thus their complexity increases with an increase in the number and dimensions of the images used for their training. In contrast, an INR network is trained to represent a specific image. Hence, unlike autoencoders, an INR-based encoder is not generic but is adapted (overfitted) to the image to be encoded and is, therefore, more efficient. Moreover, since the INR network is trained to predict an image based on the image’s coordinates, the trained network can be applied to progressively reconstruct the image based on any set of coordinates (e.g., a subset of the coordinates used in the training, or any other set of extrapolated or interpolated coordinates therefrom). Aspects described herein improve the reconstruction quality and the compression efficiency of INR-based compression, implemented with a Fourier mapping, as further described herein in reference to FIGS. 2-7.

[0020]

[0015] FIG. 1 illustrates a block diagram of an example system 100. System 100 can be embodied as a device and can be configured to perform one or more of the aspects described in this application. Examples of such devices, include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected home appliances, and servers. Elements of system 100, singly or in combination, can be embodied in an integrated circuit, multiple integrated circuits, and / or discrete components. For example, in at least one embodiment, the processing 110 and encoder / decoder 130 elements of system 100 are distributed across multiple integrated circuits and / or discrete components. In various embodiments, the system 100 is communicatively coupled to other systems, or to other electronic devices, via, for example, a communications bus or through dedicated input and / or output ports.

[0021]

[0016] The system 100 includes at least one processor 110 that can be configured to execute instructions loaded therein for implementing, for example, the various aspects described in this application. Processor 1 10 can include embedded memory, input and output interfaces, and various other circuitries as known in the art. The system 100 includes at least one memory 120, such as a volatile memory device and / or a non-volatile memory device. System 100 includes a storage device 140, which can include non-volatile memory and / or volatile memory, such as EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 140 can be an internal storage device, an attached storage device, and / or a network accessible storage device, for example.

[0022]

[0017] System 100 includes an encoder / decoder module 130 configured to process data to provide encoded video data or decoded video data. The encoder / decoder module 130 can include its own processor and memory. The encoder / decoder module 130 can be implemented as a separate element of system 100 or can be incorporated within processor 1 10 as a combination of hardware and / or software as known to those skilled in the art. Additionally, the encoder / decoder module 130 represents module(s) that can be implemented in a separate device to perform encoding and / or decoding functions.

[0023]

[0018] Program code that is to be loaded into processor 110 or into encoder / decoder 130 to perform the various aspects described in this application can be stored in a storage device 140 and subsequently loaded into memory 120 for execution by processor 1 10. In accordance with various embodiments, one or more of processor 110, memory 120, storage device 140, and encoder / decoder module 130 can store one or more of various items during the performance of the processes described in this application. Such stored items can include, but are not limited to, the input video, the decoded video or portions of the decoded video, the bitstream, matrices, variables, operational logic, and intermediate or final results from the processing of equations, formulas, operations.

[0024]

[0019] In several embodiments, memory inside of the processor 110 and / or the encoder / decoder module 130 is used to store instructions and to provide working memory for processing functions that are needed during encoding or decoding. In other embodiments, however, memory external to the processing device (where, for example, the processing device can be either the processor 110 or the encoder / decoder module 130) can be used for one or more of these functions. The external memory can be the memory 120 and / or the storage device 140 that may comprise, for example, a dynamic volatile memory and / or a nonvolatile flash memory. In several embodiments, an external non-volatile flash memory is used to store the operating system of a television. In at least one embodiment, a fast external dynamic volatile memory such as a RAM is used as working memory for video coding and decoding operations.

[0025]

[0020] The input to the elements of system 100 can be provided through various input devices as indicated in block 105. Such input devices include, but are not limited to, (i) an RF portion that receives an RF signal transmitted, for example, over the air by a broadcaster, (ii) a Composite input terminal (COMP), (iii) a USB input terminal, and / or (iv) an HDMI input terminal.

[0026]

[0021] In various embodiments, the input devices of block 105 have associated respective input processing elements as known in the art. For example, the RF portion can be associated with elements suitable for (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band of frequencies), (ii) down-converting the selected signal, (iii) band-limiting again to a narrower band of frequencies to select, for example, a signal frequency band which can be referred to as a channel in certain embodiments, (iv) demodulating the down converted and band-limited signal, (v) performing error correction, and (vi) demultiplexing to select the desired stream of data packets. The RF portion of various embodiments includes one or more elements that perform these functions, for example, frequency selectors, signal selectors, band-limiters, channel selectors, filters, downconverters, demodulators, error correctors, and demultiplexers. The RF portion can include a tuner that performs some of these functions, including, for example, down-converting the received signal to a lower frequency (for example, an intermediate frequency or a near-baseband frequency) or to a baseband. In one set-top box embodiment, the RF portion and its associated input processing element receive an RF signal transmitted over a wired (for example, cable) medium, and perform frequency selection by filtering, down-converting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above-described (and other) elements, remove some of these elements, and / or add other elements performing similar or different functions. Added elements can include inserting elements in between existing elements, for example, inserting amplifiers and an analog-to-digital converter. In various embodiments, the RF portion includes an antenna.

[0027]

[0022] Additionally, the USB and / or HDMI terminals can include respective interface processors for connecting system 100 to other electronic devices across USB and / or HDMI connections. It is to be understood that various aspects of input processing, for example, Reed-Solomon error correction, can be implemented, for example, within a separate input processing integrated circuit or within processor 110 as necessary. Similarly, aspects of USB or HDMI interface processing can be implemented within separate interface integrated circuits or within processor 1 10 as necessary. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1 10, and encoder / decoder 130 operating in combination with the memory and storage elements to process the data stream as necessary for presentation on an output device.

[0028]

[0023] Various elements of system 100 can be provided within an integrated housing. Within the integrated housing, the various elements can be interconnected and transmit data therebetween using a suitable connection arrangement 1 15, for example, an internal bus as known in the art, including the I2C bus, wiring, and printed circuit boards.

[0029]

[0024] The system 100 includes communication interface 150 that enables communication with other devices via communication channel 190. The communication interface 150 can include, but is not limited to, a transceiver configured to transmit and to receive data over communication channel 190. The communication interface 150 can include, but is not limited to, a modem or network card. The communication channel 190 can be implemented, for example, within a wired and / or a wireless medium.

[0030]

[0025] Data can be streamed to the system 100, in various embodiments, using a Wi-Fi network such as IEEE 802.1 1. The Wi-Fi signal of these embodiments is received over the communication channel 190 and the communication interface 150 which can be adapted for Wi-Fi communications. The communication channel 190 of these embodiments is typically connected to an access point or router that provides access to outside networks including the Internet for allowing streaming applications and other over-the-top communications. In other embodiments, data can be streamed to the system 100 using a set-top box that delivers the data over the HDMI connection of the input block 105 or data can be streamed to the system 100 using the RF connection of the input block 105.

[0031]

[0026] The system 100 can provide an output signal to various output devices, including a display device 165, an audio device (e.g., speaker(s)) 175, and other peripheral devices 185. The other peripheral devices 185 include, in various examples of embodiments, one or more of a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices that provide a function based on the output of the system 100. In various embodiments, control signals are communicated between the system 100 and the display device 165, the audio device 175, or the other peripheral devices 185 using signaling such as AV. link, CEC, or other communication protocols that enable device-to-device control with or without user intervention. The output devices can be communicatively coupled to system 100 via dedicated connections through respective interfaces 160, 170, and 180. Alternatively, the output devices can be connected to system 100 using the communication channel 190 via the communication interface 150. The display device 165 and the audio device 175 can be integrated in a single unit with the other components of system 100 in an electronic device, for example, a television. In various embodiments, the display interface 160 includes a display driver, for example, a timing controller (T Con) chip.

[0032]

[0027] Alternatively, the display device 165 and the audio device 175 can be separate from one or more of the other components, for example, if the RF portion of input 105 is part of a separate set-top box. In various embodiments in which the display device 165 and the audio device 175 are external components, the output signal can be provided via dedicated output connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0033]

[0028] FIG. 2 is a diagram illustrating an example INR network 200. An INR network is a neural network composed of multiple layers, each layer includes multiple nodes (denoted by circles). Generally, the architecture of a neural network is characterized by the number of layers, the number of layers’ nodes, and by the way the layers’ nodes are connected. In the example of FIG. 2, the network 200 has four layers 220, 230, 240, 250 that are fully connected. For example, the first layer 220 includes four nodes N11 , N12, N13, and N14 that each receives the coordinate values (x,y) 210 and each outputs an output signal that, in turn, feeds the nodes of the next layer, N21 , N22, N23, and N24. Likewise, the second layer 230 includes four nodes N21 , N22, N23, and N24, that each receives the output signals of nodes from the previous layer, N1 1 , N12, N13, and N14, and each outputs an output signal that, in turn, feeds the nodes of the next layer, N31 , N32, N33, and N34. The fourth layer 250 includes three nodes N41 , N42, and N43 that each receives the output signals of the nodes from the previous layer, N31 , N32, N33, N34 and each outputs a component of a pixel value 260, respectively, r, g, and b (or components of any other color model, such as y, u, and v).

[0029] Each node in the network 200 represents an operator that generates an output signal based on the node’s inputs. For example, node N21 of the second layer 230 receives as an input the output signals of nodes N11 , N12, N13, and N14, respectively, s1;s2, s3, and s4. These inputs are translated into an output signal soutthat feeds the nodes of the third layer 240. A node’s operator can be expressed as follows: where, L denotes the number of input signals (i.e., the number of nodes from the previous layer that connect to the node), s = {s; : i = 1 to L denotes the node’s input signal vector, soutdenotes the node’s output signal, p = {pt ■■ i = 0 to L denotes the node’s parameter vector (or weight vector), and A denotes an activation function (e.g., ReLU, Sigmoid, or Tanh). The weight vectors (and parameters of the activation functions, if such parameters exist) of respective nodes are collectively referred to as the parameters 9 of the network 200. These parameters 9 are determined through a training process. The network operation, denoted by fe, is therefore defined by the network parameters 9.

[0034]

[0030] Hence, an INR network 200 is trained to predict a pixel value 260 of an image based on the coordinates 210 of the pixel - that is, fe(x,y) = (r, g, b) (or f9(x,y) = (y,u, v)). During a training stage of an INR network 200, the parameters 9 (or a subset of them) are determined. This is done via an optimization process through which the parameters 9 that minimize a cost function can be determined. For example, the following cost function can be used:

[0035] Cost = D(I(x, y),fe(x,yy) + R(0) (2) where, D is a distortion measure, measuring the fidelity of the estimated pixel values f0(x,y) relative to the ground truth, that is, the corresponding pixel values from the original image / (x,y). And, where R is the resulting bitrate of the encoded parameters (e.g., encoded by quantization and entropy coding as discussed with respect to FIG. 3). A trade-off parameter A can be set to determine the balance between D and R. Note that the distortion measure D can be any metric that measure the distance (or similarity) between the original image / (x,y) and its estimated version fe(x,y), such as a mean squared error metric or a learned perceptual image patch similarity (LPIPS) metric. For example, a mean squared error metric can be expressed as: where, M and N are the width and height of the image I that the INR network is trained to predict. The optimization of the network parameters 9, according to equation (2), is typically performed by a machine learning optimization technique, applying, for example, a batch gradient descent algorithm.

[0036]

[0031] Following the training of the INR network 200 and using the optimal parameters 0 (obtained via the optimization process), the INR network can be applied to predict a pixel value based on its corresponding coordinate values. Using INR networks for the application of encoding and decoding images is further described with respect to FIGS. 3 and 4.

[0037]

[0032] FIG. 3 is a block diagram of an example video encoder 300. The video encoder 300 can be employed by the system 100 described in reference to FIG. 1. In the example of FIG. 3, the video encoder 300 includes an INR-based encoder 320, a quantizer 330, and an entropy-based encoder 340. The INR-based encoder 320 receives an input image 310 to be encoded. The input image 310 can be a still image or an image of a video frame. In an aspect, the input image 310 can be an image region of a still image or an image partition of a video frame. To code the input image 310, the INR-based encoder 320 trains an INR network (e.g., the INR networks 200, 500 described herein) to generate network parameters 0 representative of the input image 310. Specifically, based on the input image 310, the INR- based encoder 320 optimizes a cost function associated with the function fe, representative of the INR network, to determine the optimal parameters 0. For example, for an input image 310 with dimensions M and N, M times N pairs of pixel coordinates (x,y) and corresponding pixel values Z(x,y) can be used to train the INR network according to equation (2). The optimal parameters, generated by the INR-based encoder 320, are then quantized by the quantizer 330. Following quantization, the quantized parameters are entropy-coded, by the entropybased encoder 340, into a bitstream 350. That bitstream 350 can be used by a decoder to reconstruct the input image 310, as described in reference to FIG. 4.

[0038]

[0033] FIG. 4 is a block diagram of an example video decoder 400. The video decoder 400 can be employed by the system 100 described in reference to FIG. 1. The decoder 400 generally reverses the operation of the encoder 300 of FIG. 3. In the example of FIG. 4, the video decoder 400 includes an entropy-based decoder 420, a dequantizer 430, and an INR- based decoder 440. As illustrated, the decoder 400 receives the bitstream 410, 350 (generated by the encoder 300) and entropy-decodes therefrom the quantized INR network parameters. The dequantizer 430 is then employed to dequantize these quantized INR network parameters, resulting in a restored version of the INR network parameters to be provided to the INR-based decoder 440.

[0039]

[0034] The INR-based decoder 440 applies the trained INR network, defined by the restored INR network parameters, to generate the reconstructed image 450. Accordingly, the INR- based decoder 440 uses the INR network to predict the value of (or to evaluate fgusing the coordinates of) any pixel of the input image 310. Thus, the decoder 400 can be applied to: 1 ) reconstruct the whole encoded image 310; 2) to reconstruct only a region of the encoded image; or 3) to progressively reconstruct the encoded image. For example, at the encoder 300, an INR network may be trained to predict pixel values of an image with dimensions M=256 by N=256 based on the corresponding coordinates. At the decoder 400, pixel values of the image can be predicted by evaluating the trained INR network using: 1 ) the full coordinate set used for training, including all pairs of x e 0,1, ...,255 and y e 0,1, ...,255; 2) a subset of the full set, including coordinates from a region of the encoded image; or 3) a first subset of the full set, including coordinates that subsample the image (forming a low-resolution version of the encoded image) and then a second subset including the remaining coordinates. Any set of coordinates can be used to predict the corresponding pixel values, for example, in order to interpolate or to extrapolate the encoded image 310.

[0040]

[0035] In recent studies, it was found that training an INR network 200 (shown in FIG. 2) based on Fourier features extracted from the pixel coordinates 210 (instead of training directly based on the pixel coordinates) leads to better encoding performance (see, M. Tancik, et al., “Fourier features let networks learn high frequency functions in low dimensional domains,” Advances in Neural Information Processing Systems, pp. 7537-7547, 2020, thereinafter “Tancik”). In Tancik, it was shown that training an INR network 200 based on Fourier features allows the INR network to learn high-frequency components of the image. An INR network that is augmented by a Fourier mapping is described in reference to FIG. 5.

[0041]

[0036] FIG. 5 is a diagram illustrating an example of an augmented INR network 500. The augmented INR network 500 can be applied by the INR-based encoder 320 and the INR- based decoder 440 as described above. The augmented network 500 includes a mapping unit 520 and an INR network 530. The INR network 530, in principle, operates as described with respect to the INR network 200 of FIG. 2; however, in this case, the INR network 530 is trained based on features FI-FK(generated by the mapping unit 520) and not directly based on the image’s pixel coordinates 510. Thus, the mapping unit 520 can be configured to map pixel coordinates 510, using a Fourier mapping, into K Fourier features, FI-FK. A Fourier mapping can be defined by randomly selected frequencies (e.g., sampled from a Gaussian distribution) that may be used to generate respective Fourier bases (e.g., the Fourier mapping of equations (4) and (5)). Using Fourier features to train an INR network 530 can prevent a spectral bias. A spectral bias impairs the INR network’s capability to learn high frequency patterns of the image, thereby considerably impairing the visual quality of reconstructed pixels.

[0042]

[0037] The ability to learn high frequency patterns of an image depends on the number of frequencies used to generate the Fourier bases of the Fourier mapping. Current methods implement the Fourier mapping by sine and cosine functions with a mapping size of K - referred to herein as sine-cosine mapping. The number of frequencies used to generate the Fourier bases in the sine-cosine mapping is half the mapping size, that is, K / 2. When applying an INR network to image encoding (as illustrated with respect to FIGS. 3 and 4), using a higher mapping size (a larger K) increases the number of parameters 9 the INR network 530 uses, thereby increasing the bitrate required to encode these parameters. Therefore, restricting the mapping size K (to maintain a low bitrate of the coded network parameters) while maximizing the number of frequencies used (to maintain an optimal spectrum expressivity) may be crucial.

[0043]

[0038] According to aspects, an alternative Fourier mapping is disclosed herein, referred to herein as cosine mapping. In this cosine mapping the number of frequencies used to generate the Fourier bases is the same as the mapping size K. As shown herein, using cosine mapping improves image reconstruction quality (compared to the sine-cosine mapping) without increasing the complexity of the encoding and / or decoding operations.

[0044]

[0039] The conventionally used sine-cosine mapping is formulated as follows: where v = (x, y) denotes a pixel coordinate 510, is a Fourier basis vector, and atis a scalar factor, typically, set to 1 . A Fourier basis wtcan be generated by a frequency randomly sampled from a Gaussian distribution having a scale parameter a. As discussed above, the number of sampled frequencies and the mapping size have a major impact on the approximation accuracy of the image represented by an INR network and on the bitrate of the coded network parameters, respectively. The sine-cosine mapping in equation (4) is generated by m sampled frequencies, while the mapping size K (i.e., the dimension of y(v)) is 2m. Thus, the sine-cosine mapping of equation (4) is defined by a number of sampled frequencies that is half of the mapping size. For coding applications, where the mapping size limits the obtainable compression ratio, increasing the mapping size so that a larger number of sampled frequencies can be used is not advantageous.

[0045]

[0040] According to aspects described herein, a cosine mapping can be applied by the mapping unit 520, where only cosine functions are used as follows: y(v) = V2[cos where v = (x,y) denotes a pixel coordinate 510, is a Fourier basis vector, and btis a bias value. A Fourier basis wtcan be generated by a frequency randomly sampled from a Gaussian distribution having a scale parameter o , and the bias value can be randomly sampled from a uniform distribution within a range of [0,2TT], In contrast to the sine-cosine mapping of equation (4), the cosine mapping of equation (5) is defined by 2m number of sampled frequencies, that is the same as the mapping size K (i.e., dimension of y(v)). Hence, when using the cosine mapping to map the pixel coordinates 510, twice the number of frequencies are sampled and used to generate twice the number of respective Fourier bases (compared to the sine-cosine mapping). Notably, applying the cosine mapping (of equation (5)) does not increase the encoding 300 and decoding 400 complexity.

[0046]

[0041] Since Fourier mapping is carried out based on randomly sampled frequencies, based on which respective Fourier bases are formed, there is no need to code the resulting Fourier bases into the bitstream. To generate the same Fourier bases at the decoder end, the random seed used to generate the Fourier bases at the encoder end can be made available to the decoder. That random seed can be coded into the bitstream 350 or can be producible by the decoder by other means. For example, a default value can be used for the random seed or the random seed can be available from another source (e.g., the universal random seed can be used by the encoder and decoder).

[0047]

[0042] In an aspect, two random seeds can be used, one to randomly sample frequencies of the Fourier bases and a second to randomly sample the bias values. These two random seeds can be coded into the bitstream 350 or can be producible by the decoder by other means. For example, a default value may be used for each one or both of the random seeds or each one or both of the random seeds may be made available to the decoder from another source (e.g., the universal random seed(s) can be used by the encoder and decoder). In an aspect, one random seed may be used to randomly sample the frequencies of the Fourier bases and the bias values.

[0048]

[0043] Hence, the INR-based encoder 320 and the INR-based decoder 440 can be configured to apply cosine mapping (according to equation (5)). A flag can be used to indicate that a cosine mapping is applied as the Fourier mapping. In another aspect, the INR-based encoder 320 and the INR-based decoder 440 can be configured to apply either one of the Fourier mappings, the sine-cosine mapping (according to equation (4)) or the cosine mapping (according to equation (5)). The INR-based encoder may select to apply either mapping ((4) or (5)) to represent (encode) each image (or each image region) of a video frame in a video sequence, and a flag indicating which mapping is applied should then be coded into the bitstream.

[0049]

[0044] The encoding and decoding processes, when applying the cosine mapping, are further described below in reference to FIGS. 3-5.

[0050]

[0045] According to aspects, the encoder 300 is configured to code an input image 310 into a bitstream 350, 410, using the augmented INR network 500 and the cosine mapping described herein. The following steps may be carried out by the encoder 300.

[0051] • Step C1 : The cosine mapping (according to equation (5)) is determined by 1 ) sampling frequencies (used to generate the Fourier bases) from a Gaussian distribution with a scale a and 2) sampling respective bias values from a uniform distribution with a range of [0,2TT], These random samplings are performed using respective random seeds (optionally, derived from the same seed).

[0052] • Step C2: The coordinates 510 of the input image 310 are mapped, by the mapping unit 520, into K features, FI-FK, using the determined cosine mapping.

[0053] • Step C3: Based on the features FI-FK, the INR network 530 is trained to optimize the network parameters 9 using a cost function (see equation (2)).

[0054] • Step C4: The random seed(s), used in step C1 , is / are coded into the bitstream 350 (unless the random seed(s) can be independently derived by the decoder).

[0055] • Step C5: A flag is coded, into the bitstream, indicating that cosine mapping is used to map the pixel coordinates 510 of the input image 310.

[0056] • Step C6: The optimized network parameters 9 may be quantized 330. Any quantization technique may be used (e.g., a fixed bit quantization technique).

[0057] • Step C7: The quantized network parameters 9 may be entropy-coded 340 into the bitstream 350. Any entropy-based encoding technique may be used.

[0058] Note that the quantization of the network parameters 9 is optional. In the case where the network parameters 9 are quantized, in an aspect, quantization aware training of the INR network may be applied to reduce the impact of the quantization error (as described in European Application No. 22306480.9, filed on October 4, 2022, which is incorporated herein by reference in its entirety).

[0059]

[0046] According to aspects, the decoder 400 is configured to decode the bitstream 350, 410 to generate a reconstructed image 450, using the augmented INR network 500 and the cosine mapping described herein. The following steps may be carried out by the decoder 400.

[0060] • Step D1 : A flag is entropy-decoded 420 from the bitstream 410, indicating that the cosine mapping should be used by the mapping unit 520.

[0061] • Step D2: Random seed(s) (coded in step C4) are entropy-decoded 420 from the bitstream 410.

[0062] • Step D3: The INR network 530 parameters 9 are entropy-decoded 420 from the bitstream 410.

[0063] • Step D4: If quantization was applied in the encoding process (see step C6), quantization parameter(s) are decoded from the bitstream and applied to the dequantization 430 of respective data. • Step D5: Cosine mapping (according to equation (5)) is determined by 1 ) sampling frequencies (used to generate the Fourier bases) from a Gaussian distribution with a scale a and by 2) sampling respective bias values from a uniform distribution with a range of [0,2TT], These random samplings are performed using the random seed(s) decoded from the bitstream (step D2) or random seed(s) that are otherwise derivable, or obtainable, from another source.

[0064] • Step D6: The image 450 is reconstructed using the trained INR network 500 (defined by the INR parameters 9 decoded in step D3) by inference - that is, by evaluating the network 500 for each pixel’s coordinates 510 for which the corresponding pixel value 540 is to be predicted. Thus, each such pixel’s coordinates 510 are first mapped 520 (by the cosine mapping determined in step D5) into respective Fourier features FI-FK, and these features are then fed into the INR network 530 through which the corresponding pixel value is evaluated based on the decoded INR parameters 9.

[0065]

[0047] In an aspect, the training of the INR network 530 (to obtain the optimized network parameters 0) may also include optimizing the used random seed(s). For example, various random seed values can be used to generate samples of the frequencies and the respective bias values (as in steps C1 or D5), thereby determining respective cosine mapping candidates. And the random seed that corresponds to the cosine mapping candidate that (when used to reconstruct the image, as in step D6) provides the best quality of image reconstruction is selected.

[0066]

[0048] FIG. 6 is a flowchart of an example method for encoding video data 600 according to aspects described herein. The method 600 can be applied (e.g., by the encoder 300 of FIG. 3) to encode, into a bitstream 350, an image region 310 from a video frame of the video data. The coding of the image region can be performed according to steps 610 and 620. In step 610, using a cosine mapping, the INR network 530 is trained to predict a pixel value of the image region based on respective pixel coordinates, where, as described herein, the training determines parameters of the INR network. In step 620, the determined parameters of the INR network are coded into the bitstream. A flag, indicating that cosine mapping is used, can also be coded into the bitstream. According to method 600, the training of the INR network includes mapping coordinates of pixels from the image region into respective feature vectors and training the INR network based on these feature vectors. As explained above, the cosine mapping comprises a number of Fourier bases that is equal to a dimension K of the cosine mapping.

[0067]

[0049] Aspects of method 600 further include randomly sampling, based on a first random seed, frequencies used to generate the Fourier bases, and randomly sampling, based on a second random seed, the bias values. At least one of the first random seed and the second random seed can be coded into the bitstream to be made available to the decoder. In an aspect, one of the first and of the second random seeds can be obtained from each other. In another aspect at least one of the first and of the second random seeds is obtained from a predetermined seed value. In yet another aspect, a seed value can be selected for at least one of the first and of the second random seeds, so that a cosine mapping candidate associated with the selected seed value provides the best quality of image region reconstruction, wherein reconstruction of an image region with respect to a cosine mapping candidate is performed by evaluating the INR network using the parameters of the INR network and the cosine mapping candidate.

[0068]

[0050] FIG. 7 is a flowchart of an example method for decoding video data 700 according to aspects described herein. The method 700 can be applied (e.g., by the decoder 400 of FIG. 4) to decode, from a bitstream 410, an image region 450 from a video frame of the video data. The decoding of the image region can be performed according to steps 710 and 720. In step 710 parameters of an INR network are decoded from the bitstream. The INR network 530 is typically trained by the encoder 320 using a cosine mapping as described herein. The INR network 530 is trained to predict a pixel value of the image region based on respective pixel coordinates, where the training of the INR network determines the parameters that are decoded from the bitstream. In step 720, the image region is reconstructed by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping. A flag indicating that cosine mapping is used can also be decoded from the bitstream.

[0069]

[0051] As explained above, the used cosine mapping can be determined by the decoder 400 in the same manner it is determined by the encoder 300, using the same respective random seeds to randomly sample the frequencies used to generate the Fourier bases of the cosine mapping and to randomly sample the bias values of the cosine mapping. One or both of these respective random seeds can be decoded from the bitstream, can be obtained from each other, or can be obtained from a predetermined seed value.

[0070]

[0052] Applying the cosine mapping, as described above in reference to FIGS. 3-5, leads to a compression with better rate-distortion results as shown below by experiments conducted by the inventors. Moreover, this improved performance is provided without adding complexity to the encoding and decoding processes; that is, without increasing the INR network’s 530 complexity - the same number of network inputs, FI-FK, is maintained when applying the sinecosine mapping (according to equation (4)) and when applying the cosine mapping (according to equation (5)) by the mapping unit 520.

[0071]

[0053] Several sets of experiments were conducted to evaluate the performance of aspects described herein. First, experiments were conducted using a low mapping size of K = 8 and using an INR network consisting of 2 layers, each layer containing 28 nodes with an ReLLI activation function. Table 1 compares the performance achieved by using sine-cosine mapping (equation (4)) and cosine mapping (equation (5)), in terms of bitrate measured by bits-per-pixel (BPP) and in terms of peak signal-to-noise ratio (PSNR), obtained for an image from the Kodak dataset.

[0072]

[0054] Table 1 : Comparison between sine-cosine and cosine mappings.

[0073]

[0055] The results shown in Table 1 demonstrate that the number of Fourier bases in the mapping has a significant impact on the quality of image reconstruction. With the same mapping size but twice the number of Fourier bases, the cosine mapping (relative to the sinecosine mapping) improves the reconstruction quality by about 6 dB PSNR for the same bitrate.

[0074]

[0056] Next, experiments were conducted in which the depth of the neural network was varied. The mapping size was set to K = 16 and the width of the network (number of nodes per layer) was set to 28, while the depth of the network (number of layers) was varied, resulting different bitrates. The performances achieved by using sine-cosine mapping (equation (4)) and cosine mapping (equation (5)) are compared in terms of bitrate (BPP) and distortion measure (PSNR). Table 2 shows results obtained by using 1 to 4 network layers. The gain in using the cosine mapping relative to the sine-cosine mapping is also shown using the Bjontegaard-Delta metric, represented by BD Rate gain and BD PSNR gain (see, G. Bjontegaard, "Calculation of average PSNR differences between RD-curves," tech, rep., VCEG, April 2002. Contribution VCEG-M33).

[0075]

[0057] Table 2: Comparison between sine-cosine and cosine mappings across different number of network layers.

[0076]

[0058] Table 2 demonstrates that using the cosine mapping leads to better quality of reconstruction (higher PSNR results) for the same bitrates (BPP) relative to the sine-cosine mapping. The Bjontegaard-Delta metric, represented by the BD Rate gain and the BD PSNR gain, is calculated to quantify net bitrate saving due to the use of the cosine mapping instead of sine-cosine mapping (last two rows of the Table 2). These gains show that using cosine mapping provides about 14% BD Rate gain and about 0.6 dB PSNR gain on average. According to compression standards, these results are considered to be significant. Moreover, these gains are provided without introducing any additional complexity to processes in the encoder and the decoder and without increasing the required bitrate to code the symbols compared to the sine-cosine mapping.

[0077]

[0059] Further experiments were conducted using similar neural network architecture used in the literature (see, e.g., E. Dupont, et al., “COIN: Compression with Implicit Neural Representations,” ICLR Workshop, 2021 , hereinafter “COIN”) to show that the gains demonstrated herein are not specific to a certain network architecture. In these experiments, the mapping size was set to K = 32 and the network architecture (denoted by Q1 -Q4 in Table 3) corresponds to different quality levels (bitrates). Table 3 shows rate and distortion results across different bitrates when sine-cosine mapping is used and when cosine mapping is used, and with respect to the common network architecture used in COIN.

[0078]

[0060] Table 3: Comparison between sine-cosine and cosine mappings across different bitrates.

[0079]

[0061] Table 3 demonstrates that, relative to the sine-cosine mapping, the cosine mapping provides better rate distortion performance across different bitrates. As shown, the cosine mapping provides a significant BD rate gain of about 27% and BD PSNR gain of about 0.3 dB over the sine-cosine mapping.

[0062] The illustrations of the aspects described herein are intended to provide a general understanding of the structure, function, and operation of the various aspects. The illustrations are not intended to serve as a complete description of all of the elements and features of apparatuses and systems that utilize the structures or methods described herein. Many other aspects may be apparent to those of skill in the art upon reviewing the disclosure. Other aspects may be utilized and derived from the disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.

[0063] The description of the aspects is provided to enable the making or use of the aspects. Various modifications to these aspects will be readily apparent, and the generic principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

CLAIMS1 . A method for encoding video data, comprising: encoding, into a bitstream, an image region from a video frame of the video data, the coding of the image region comprises: training an implicit neural representation (INR) network, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training determines parameters of the INR network, and coding, into the bitstream, the determined parameters of the INR network.

2. The method according to claim 1 , further comprising: coding, into the bitstream, a flag indicating that the cosine mapping is used.

3. The method according to claim 1 or 2, wherein the training of the INR network comprises: mapping, using the cosine mapping, coordinates of pixels from the image region into respective feature vectors; and training the INR network based on the feature vectors.

4. The method according to any one of claims 1 to 3, wherein the cosine mapping comprises a number of Fourier bases that is equal to a dimension of the cosine mapping.

5. The method according to any one of claims 1 to 4, wherein the cosine mapping is defined by Fourier bases and respective bias values, and further comprising: randomly sampling, based on a first random seed, frequencies used to generate the Fourier bases; and randomly sampling, based on a second random seed, the bias values.

6. The method according to claim 5, further comprising: coding, into the bitstream, the first random seed, the second random seed, or the first and the second random seeds.

7. The method according to claim 5 or 6, wherein one of the first and the second random seeds is obtained from the other random seed.

8. The method according to any one of claims 5 to 7, wherein at least one ofthe first and the second random seeds is obtained from a predetermined seed value.

9. The method according to any one of claims 5 to 8, further comprising: selecting a seed value for at least one of the first and the second random seeds, so that a cosine mapping candidate associated with the selected seed value provides the best quality of image region reconstruction, wherein reconstruction of an image region with respect to a cosine mapping candidate is performed by evaluating the INR network using the parameters of the INR network and the cosine mapping candidate.

10. A method for decoding video data, comprising: decoding, from a bitstream, an image region from a video frame of the video data, the decoding of the image region comprises: decoding, from the bitstream, parameters of an INR network, the INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training of the INR network determines the parameters decoded from the bitstream, and reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.1 1 . The method according to claim 10, further comprising: decoding, from the bitstream, a flag indicating that the cosine mapping is used.

12. The method according to claim 10 or 11 , wherein the training of the INR network comprises: mapping, using the cosine mapping, coordinates of pixels from the image region into respective feature vectors; and training the INR network based on the feature vectors.

13. The method according to any one of claims 10 to 12, wherein the cosine mapping comprises a number of Fourier bases that is equal to a dimension of the cosine mapping.

14. The method according to any one of claims 10 to 13, wherein the cosine mapping is defined by Fourier bases and respective bias values, and further comprising: randomly sampling, based on a first random seed, frequencies used to generate the Fourier bases; and randomly sampling, based on a second random seed, the bias values.

15. The method according to claim 14, further comprising: decoding, from the bitstream, the first random seed, the second random seed, or the first and the second random seeds.

16. The method according to claim 14 or 15, wherein one of the first and the second random seeds is obtained from the other random seed.

17. The method according to any one of claims 14 to 16, wherein at least one of the first and the second random seeds is obtained from a predetermined seed value.

18. An apparatus for encoding video data, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to encode, into a bitstream, an image region from a video frame of the video data, the coding of the image region comprises: training an INR network, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training determines parameters of the INR network, and coding, into the bitstream, the determined parameters of the INR network.

19. The apparatus according to claim 18, wherein the instructions further cause the apparatus to: code, into the bitstream, a flag indicating that the cosine mapping is used.

20. The apparatus according to claim 18 or 19, wherein the cosine mapping comprises a number of Fourier bases that is equal to a dimension of the cosine mapping.21 . The apparatus according to any one of claims 18 to 20, wherein the cosine mapping is defined by Fourier bases and respective bias values, and wherein the instructions further cause the apparatus to: randomly sample, based on a first random seed, frequencies used to generate the Fourier bases; and randomly sample, based on a second random seed, the bias values.

22. The apparatus according to claim 21 , wherein the instructions further cause the apparatus to:code, into the bitstream, the first random seed, the second random seed, or the first and the second random seeds.

23. An apparatus for decoding video data, comprising: at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to decode, from a bitstream, an image region from a video frame of the video data, the decoding of the image region comprises: decoding, from the bitstream, parameters of an INR network, the INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training of the INR network determines the parameters decoded from the bitstream, and reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.

24. The apparatus according to claim 23, wherein the instructions further cause the apparatus to: decode, from the bitstream, a flag indicating that the cosine mapping is used.

25. The apparatus according to claim 23 or 24, wherein the cosine mapping comprises a number of Fourier bases that is equal to a dimension of the cosine mapping.

26. The apparatus according to any one of claims 23 to 25, wherein the cosine mapping is defined by Fourier bases and respective bias values, and wherein the instructions further cause the apparatus to: randomly sample, based on a first random seed, frequencies used to generate the Fourier bases; and randomly sample, based on a second random seed, the bias values.

27. The apparatus according to claim 26, further comprising: decoding, from the bitstream, the first random seed, the second random seed, or the first and the second random seeds.

28. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for encoding video data, the method comprising: encoding, into a bitstream, an image region from a video frame of the video data,the coding of the image region comprises: training an INR network, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training determines parameters of the INR network, and coding, into the bitstream, the determined parameters of the INR network.

29. A non-transitory computer-readable medium comprising instructions executable by at least one processor to perform a method for decoding video data, the method comprising: decoding, from a bitstream, an image region from a video frame of the video data, the decoding of the image region comprises: decoding, from the bitstream, parameters of an INR network, the INR network is trained, using a cosine mapping, to predict a pixel value of the image region based on respective pixel coordinates, wherein the training of the INR network determines the parameters decoded from the bitstream, and reconstructing the image region by evaluating the INR network using the decoded parameters of the INR network and the cosine mapping.