Warped temporal context for feature-based temporal inr

The hybrid INR approach with warped temporal context addresses inefficiencies in existing image and video compression by enhancing entropy encoding with spatial and temporal latent values, improving efficiency and reducing complexity for dynamic 3D objects and scenes.

WO2026041590A1PCT designated stage Publication Date: 2026-02-26INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/073547
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-22
Filing Date
2025-08-18
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Existing image and video compression technologies face challenges in efficiently leveraging spatial and temporal redundancy due to high computational complexity and inefficiencies in entropy encoding of latent features, particularly in dynamic 3D objects and scenes.

Method used

The implementation of a hybrid Implicit Neural Representation (INR) approach that utilizes a warped temporal context for entropy encoding, combining spatial and temporal latent values through optical flow to enhance the predictive power of latent feature encoding.

Benefits of technology

This method reduces computational complexity and improves encoding efficiency by using a warped temporal context, resulting in a more effective bitrate-distortion trade-off for dynamic 3D objects and scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025073547_26022026_PF_FP_ABST
    Figure EP2025073547_26022026_PF_FP_ABST
Patent Text Reader

Abstract

A method for encoding an input signal is provided, wherein a current latent representative of features of a current frame of the input signal is obtained, at least one reconstructed part of a reference latent representative of features of a reference frame of the input signal is warped to the current frame, a temporal context is determined for at least of one value of the current latent from the warped at least one reconstructed part of the reference latent, and the at least one value of the current latent is entropy encoded based at least on the temporal context. The input signal can be time-varying signal, such as a video, dynamic 3D object or 3D scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Docket No.2024P00587WO WARPED TEMPORAL CONTEXT FOR FEATURE-BASED TEMPORAL INR ______________________________________________ CROSS REFERENCE TO RELATED APPLICATIONSThis application claims the priority to European Application No. 24306385.6, filed on22 August 2024, which is incorporated herein by reference in its entirety.BACKGROUND The present embodiments generally relate to image, video and / or 3D scene compression using Feature-based Implicit Neural Representation (INR). To achieve high compression efficiency, image and video coding schemes usually employ prediction and transform to leverage spatial and temporal redundancy in the video content. Emerging technology makes use of neural networks. Among them, Implicit Neural Representation (INR) aims at parameterizing a function which takes coordinates as inputs and outputs values of a signal at these coordinates. INR can be used for instance for compressingimage, videos, 3D objects or haptic texture. Furthermore, these approaches have a far lowercomputational complexity than end-to-end neural compression approaches. BRIEF SUMMARYBriefly stated, in one embodiment, a method for encoding an input signal is provided, whereina current latent representative of features of a current frame of the input signal is obtained, atleast one reconstructed part of a reference latent representative of features of a reference frameof the input signal is warped to the current frame, a temporal context is determined for at leastof one value of the current latent from the warped at least one reconstructed part of the reference latent, and the at least one value of the current latent is entropy encoded based at least on the temporal context.The input signal can be time-varying signal, such as a video, dynamic 3D object or 3D scene.In another embodiment, a method for reconstructing the signal is provided, wherein at leastone part of a warped reference latent is obtained from at least one part of at least one previouslyreconstructed latent representative of features of a previously reconstructed frame of the signalto reconstruct, by warping to a current frame of the signal, the at least one part of the at leastone previously reconstructed latent, a temporal context is determined for at least of one value Docket No.2024P00587WO of a current latent representative of features of the current frame from the at least one part of the warped reference latent, and the at least one value of the current latent is entropy decoded based at least on the temporal context.In another embodiment, an apparatus is provided that comprises one or more processorsoperable to perform any one of the methods mentioned above.One or more embodiments also provide a computer program comprising instructions which when executed by one or more processors cause the one or more processors to perform any ofthe methods mentioned above. One or more of the present embodiments also provide a non-transitory computer readable medium and / or a computer readable storage medium havingstored thereon instructions for performing any of the methods mentioned above.One or more embodiments also provide a computer readable storage medium having stored thereon a bitstream generated according to the methods described herein. One or more embodiments also provide a method and apparatus for transmitting or receiving the bitstream generated according to the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS The following detailed description will be better understood when read in conjunction with the appended drawings, in which there are shown examples of one or more of the multiple embodiments of the present disclosure. It should be understood, however, that theembodiments described herein are not limited to the precise arrangements and instrumentalitiesshown in the drawings. In the drawings:FIG. 1 is a block diagram illustrating an example system according to one or moreembodiments of the present disclosure.FIG. 2 illustrates an example of a neural network for Implicit Neural Representation (INR).FIG. 3 illustrates an example of a method for encoding a signal using an INR.FIG. 4 illustrates an example of a neural architecture for hybrid-INR.FIG. 5 illustrates an example of a method for entropy encoding latent values using spatio- temporal context. FIG. 6 illustrates an example of a method for encoding a video frame using a COOL CHIC architecture. Docket No.2024P00587WOFIG. 7 illustrates an example of a method for obtaining a warped temporal context accordingto an embodiment. FIG. 8 illustrates an example of a method for entropy encoding latent values using spatial context and warped temporal context, according to an embodiment. FIG. 9 illustrates an example of a method for obtaining a warped temporal context according to another embodiment.FIG. 10 illustrates an example of a method for encoding an input signal according to anembodiment.FIG. 11 illustrates an example of a method for decoding / reconstructing a signal from abitstream according to an embodiment.FIG. 12 shows two remote devices communicating over a communication network inaccordance with an example of the present principles.FIG. 13 shows the syntax of a signal in accordance with an example of the present principles.DETAILED DESCRIPTION In describing the various embodiments of the present disclosure, certain terminology is used herein for convenience only and should not be considered as limiting such embodiments. In the drawings, the same reference numerals are employed for designating the same elements throughout the several figures and the present description. Referring to the drawings, there is shown in FIG. 1 a block diagram illustrating an example system 100 in which embodiments of the present disclosure can be implemented. Thesystem 100 may be an electronic device including, for example, a personal computer, laptopcomputer, mobile phone, tablet computer, multimedia set-top box, digital television receiver,personal video recording system, connected home appliance, vehicle control and / orentertainment system, and server. One or more elements of the system 100, singly or in combination, may be implemented as an integrated circuit (IC), multiple ICs, and / or discretecomponents. For example, in one embodiment, the processing, encoding and / or decodingelements of system 100 are distributed across multiple ICs and / or discrete components. In some embodiments, the system 100 is communicatively coupled to and / or in communication with other systems or devices, via, for example, a communications bus or dedicated input / output ports. Docket No.2024P00587WO One or more of the elements of system 100 may be provided within an integrated housing, with such elements being interconnected and able to transmit data therebetween using any suitable connection arrangement 115 generally known in the art, including, for example, an internal bus (e.g., I2C bus), wiring, and printed circuit boards. The system 100 includes at least one processor 110 configured to execute instructions for implementing the embodiments described herein, including signal / data coding andprocessing. The processor 110 may be a general-purpose processor or microprocessor, digitalsignal processor (DSP), one or more microprocessors in association with a DSP core, a controller, a microcontroller, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), a state machine, and the like. The processor 110 may include at least one central processing unit (CPU), embedded memory, input and output interfaces, and other circuitries. The system 100 includes at least one memory 120, for example, a volatile memorydevice and / or a non-volatile memory device. The system 100 includes a storage device 140,that may be or include non-volatile memory and / or dynamic volatile memory, includingEEPROM, ROM, PROM, RAM, DRAM, SRAM, DDR, flash, magnetic disk drives, solid state drives (SSD) and / or optical disk drives. The storage device 140 may be or include, for example, an internal storage device, an attached storage device, and / or a network accessible storage device. Although shown separately, the memory 120 and the storage device 140 may be collocated, integrated together, or otherwise combined. The system 100 includes an encoder / decoder module 130 configured to process videodata and to provide encoded video data or decoded video data. The encoder / decoder module130 may include one or more processors and / or memory (not shown). Although FIG. 1 depictsthe encoder / decoder module 130 as a separate element of system 100, it will be understood thatthe processor 110 and the encoder / decoder module 130 may be collocated and / or integratedtogether as a combination of hardware and / or software, e.g., in an electronic package or chip.The encoder / decoder module 130 may be or include one or more modules that may be includedin one or more separate devices that perform encoding and / or decoding functions.Instructions for execution by the processor 110 and / or the encoder / decoder module 130 may be stored in the storage device 140 and subsequently loaded into memory 120 for execution by the processor 110. In some embodiments, one or more of processor 110, memory120, storage device 140, and encoder / decoder module 130 may store one or more items whenperforming the processes disclosed herein. Such items may include input video, decoded video Docket No.2024P00587WO or portions thereof, bitstreams, matrices, variables, operational logic, and intermediate and / or final results from processing of equations, formulas, or operations. In some embodiments, the memory of the processor 110 and / or the encoder / decodermodule 130 is used to store instructions and / or provide working memory for video encodingand decoding functions. In some embodiments, memory external to the processor 110 and / orthe encoder / decoder module 130 (e.g., the memory 120 and / or the storage device 140) is usedfor one or more of these functions and / or, for example, to store the operating system of atelevision. The system 100 may obtain or receive information via one or more input devices,interfaces, and / or ports as indicated in input block 105. Examples of the input devices includea radio frequency (RF) device for transmitting and / or receiving RF signals over various media,for example, RF signals received over the air from a broadcaster; component video (COMP)inputs; a Universal Serial Bus (USB) input; and / or a High-Definition Multimedia Interface(HDMI) input. Other examples include composite video input (not shown). In someembodiments, the input devices are associated with respective input processing elements, e.g.,those generally known in the art. For example, the RF device may be associated with elementssuitable for selecting a desired frequency (e.g., selecting or band-limiting a signal) orperforming error correction on the signal. The USB and / or HDMI inputs may includerespective interface processors and transceivers (or transmitters and receivers) for coupling the system 100 to other devices via USB and / or HDMI ports or connections. Various forms of input processing may be implemented, for example, by and / or within a separate input processing device or the processor 110. The system 100 includes a communication interface 150 that enables wired and / orwireless communication with other devices, e.g., via a communication channel 190. Thecommunication interface 150 may include one or more transceivers, modems, network cardsand the like. The communication channel 190 may be or include wired and / or wireless mediums. In some embodiments, data may be streamed to the system 100 via wired and / orwireless networks. Examples of such wireless networks include cellular, Bluetooth or Wi-Fi(e.g., IEEE 802.11) networks. The wired and / or wireless networks may include one or more base stations (e.g., cellular base stations, access points, etc.), and / or user equipment (e.g. cellular user equipment, stations, etc.), and / or other network elements that communicate withthe system 100 via the communication interface 150 and communication channel 190, wherebythe system 100 may obtain data streamed from streaming applications (e.g., OTT services) via Docket No.2024P00587WOvarious networks, including the Internet. In some embodiments, data is streamed to the system100 via the input block 105 (e.g., using a set-top box that delivers data via the HDMI connectionor the RF connection). In some embodiments, data is received by the system 100 in a non-streaming manner. The system 100 may provide one or more output signals to one or more output devices.The output devices may include a display device 165 (e.g., touchscreen display, monitor, etc.),an audio device 175 (e.g., speakers), and other peripheral devices 185, including, for example,a stand-alone DVR, a disk player, a stereo system, a lighting system, and other devices thatprovide a function based on the output of the system 100. The display device 165 can be for atelevision, tablet, laptop, mobile phone, head-mounted display, or other device. In someembodiments, control signals are communicated between the system 100 and the display device 165, the audio device 175, and / or the peripheral devices 185, enabling device-to-device controlwith or without user intervention. The output devices may couple to and / or communicate withthe system 100 via dedicated connections via respective display, audio, and peripheralinterfaces 160, 170, 180. Alternatively, the output devices may couple to and / or communicatewith the system 100 via the communication channel 190 and the communication interface 150. The display device 165 and the audio device 175 may be collocated, integrated, or otherwise combined with the other components of system 100 in a single unit (e.g., a television). Alternatively, the display device 165 and the audio device 175 may be separate from one or more of the other components of the system 100. In embodiments in which thedisplay device 165 and the audio device 175 are external components, the output signals maybe provided via dedicated outputs and / or connections, including, for example, HDMI ports,USB ports, or COMP outputs.FIG. 2 illustrates an example of a neural network that can be used for implicit neuralrepresentation (INR). Such a neural network used for INR can be referred to as an INR network.The INR network is also referred as Coordinates Neural Representation. The INR network allows to obtain a compact representation of an input signal, an image for example. The INRnetwork models the input signal by an overfitted Multi Layer Perceptron (MLP), performing,in the case of an image for example, the mapping from pixel coordinates to its RGB values.The INR parameterizes a signal as a function 200, which takes coordinates 210 as input andoutputs values 220 of a signal at these coordinates. INR has recently been applied to image,videos or 3D objects among other applications. In the image case, the inputs 210 can be pixel Docket No.2024P00587WOcoordinates (^^, ^^) and the INR may output 220 the color values (^, ^, ^) or (^, ^, ^) of theinput pixel. The input coordinates may be modified by a transformation before being used as input for the neural network. This transformation can be a Fourier mapping, coordinate transformation, normalization etc. The INR can be used to reconstruct a signal by computing the signal values for every necessary coordinate inputs. It can be used to upsample a signal by generating output for input coordinates corresponding to the upsampled pixels, for example the mean of the coordinates between two consecutive pixels for upsampling by a factor of 2. An INR network 200 is typically a neural network, composed of multiple neural layers, such as fully connected layers. For example, in FIG. 2, the network has four layers. Intermediate outputs are represented by circles. Each neural layer can be described as a function that first multiplies the input by a tensor, adds a vector called the bias and then applies a nonlinear function on the resulting values. The shape (and other characteristics) of the tensor and the typeof non-linear functions are called the architecture of the network. The values of the tensor andthe bias are denoted by the term “weights”. The weights and, if applicable, the parameters of the non-linear functions, are called the parameters θ of the network. The architecture and the parameters define a “model”. We will use ^^to denote an INR function parameterized by θ. FIG. 3 illustrates an example of a method 300 to encode a signal 310 using an INR. This isdone by optimizing 320 the parameters θ (or a subset of them) of the INR network toreconstruct the signal and optionally encoding 330 them to create the output bitstream 350. Foran image x of size (^ × ^), the parameters θ can for example be optimized by minimizing thefollowing loss function: Loss = D(^, ^^) + λ^(θ)   where D is a distortion which quantifies the difference between the reconstructed image by ^^to the original image x, R is the bitrate of the encoded parameters and λ a trade-off parameterbetween D and R. D could be any differentiable distortion measure, such as mean squared erroras in the second equation. M and N are the width and height of the original image. Other metricssuch as LPIPS (learned perceptual image patch similarity) can also be used in this case. Theoptimization of the parameters θ is typically performed by a machine learning approach suchas a batch gradient descent method. Docket No.2024P00587WO To decompress the signal, ^^is evaluated at all relevant coordinates. These coordinates can be selected at decoding. A typical choice would be all pixel coordinates for an image or video. Asan example, for a 256x256 pixel image, these coordinates could be all pairs (^^, ^^) for all ^^ ∈0,1, … ,255 and ^^ ∈ 0,1, … ,255. Other choices are possible, for example to upsample,downsample or extend the original image. The bitstream encoding a signal is thus created by encoding the parameters of the neural network. This can be done by a neural compression codec such as Neural Network Coding (NNC) / ISO / IEC 15938-17 or MPEG-7 part 17 or by quantizing the parameters and / or pruning some neurons from the network.FIG. 4 illustrates an example of a neural architecture that is an INR variant, called in thefollowing Hybrid INR. The input coordinates 410 are first mapped to a vector of latent features430. Such a mapping 420 can for example rely on a lookup table, a partition of the input signal,a hash function and / or a linear combination of features. Multiple mappings may also be done, the associated features are then concatenated. It may also involve interpolation between neighboring vectors of features, for input coordinates that are not directly associated to a vectorof features. These features may also be upsampled if needed. A vector of features is obtainedby these mappings for given coordinates 410 and used as input for an INR synthesis network 440 that outputs the value 450 of the signal for the input coordinates 410. Using such an architecture rather than a plain INR network helps to handle local features of the input signal. Indeed, the features for a given location can be completely or mostly independent from featuresat other coordinates and thus can be tailored to each location. Encoding a signal using a hybrid-INR is similar to traditional INR: both the network parameters (440) and the features (430) areoptimized / learned by minimizing a loss as described above.An example of a hybrid-INR architecture, called COOL-CHIC, for image coding can be foundin Ladune T, Philippe P, Henry F, et al. (2023) Cool-chic: Coordinate-based low complexity hierarchical image codec. Proceedings of the IEEE / CVF International Conference onComputer Vision 13515–13522 and an extension of the COOL-CHIC system for video codingis described in Leguay T, Ladune T, Philippe P, Déforges O (2024) Cool-chic video: Learned video coding with 800 parameters. In: 2024 Data Compression Conference (DCC). IEEE, pp 23–32.An example of a multiresolution hash encoding (called Instant-NGP) can be found in “T.Müller, A. Evans, C. Schied, A. Keller, Instant Neural Graphics Primitives with a Docket No.2024P00587WOMultiresolution Hash Encoding, ACM Trans. Graph. Vol 41, N°4, July 2022” whichcomplements a fully connected neural network with a latent representation of the image. InInstant NGP, the latent representation of the image describes different spatial locations with different latent parameters. The latent representation is arranged into multiresolution levels,each level being independent and storing features vectors at the vertices of a grid. Toreconstruct an image for example, the fully connected neural network takes as input a featuresvector obtained for a given input coordinate of the image, the features vector being aconcatenation of results obtained for each level using hash tables and linear interpolation andauxiliary inputs and proves RGB values for the input coordinates.In hybrid INR approaches, the latent features y are typically the largest contributor to thebitstream size, several orders of magnitude larger than the one associated to the MLP parameters, except for small bitstream lengths. One solution to reduce the transmission cost ofthe latent features relies on quantization and entropy coding.The attention given to the entropy coding of the latent features y appears on the new formulation of the loss function: ^ m^,^i,n^^ ^^, ^^^upsample where ^^ is the quantized latent features and ^^(^^) is the discrete distribution over the quantizedlatent features. The distribution of the latent features y can for example be a known distributionor estimated using an auto regressive probability model. According to Equation (1), minimizing the rate associated to the transmission of the compressed version of the frame relies on minimizing the rate associated to the latent features. This can be achieved either by: ^Reducing the amount of information contained in the latent features risking a poorreconstruction of the sent frame and an increase of distortion in ^=^^^upsample(^^)^. (less information in ^^ -> more distortion in ^).^ Or we can try to estimate the distribution of the sent latent as close as possible to thereal (unknown) one using a well-chosen probability model. If that probability model is trained, it must be included in the bitstream.Due to its high dimension the modeling of the joint distribution of ^^ is usually untrackable. Docket No.2024P00587WOInstead it is typical to factorize ^^(^^) and use a set of C context latents c^^^^such that the distribution of each quantized latent conditioned on C spatially neighboring latents that have already been decoded and selected in a way to introduce as little sequentiality as possible to allow parallel decoding of the L channels, for example in a wavefront-like approach. Theposition of the latent considered must be known by the emitter and the receiver. Thefactorization of ^^(^^) is given by: Where ^^^^y^^^^c^^^ ^ denotes the conditional probability of one latent value at position (i, j, k)conditioned on its spatial context c^^^^ . (i, j) represents the spatial coordinate and k represents alatent feature.FIG. 5 shows an example introducing temporal information in the context by combining thespatial context latents used from the current frame, and others context latents selected from the already encoded / decoded reference frame(s), leading to a better bitrate / distortion trade-off. InFIG. 5, for encoding or decoding a current latent value, at 510, spatial context and temporalcontext latent values are obtained for the current latent value. Spatial context values are obtained from a causal region neighboring the current latent value in the current frame, and temporal context values is selected as latent values reconstructed for the reference frame in aco-located region with the location of the current latent value to encode or decode. Spatial andtemporal context values are combined at 520 and provided at 530 to the probability model estimator (ARM) which outputs at 540 a probability value for the current latent value to encodeor decode. At 550, the current latent value is encoded or decoded using the obtained probabilityvalue.In the case of FIG. 5, the probability distribution over the latent features factorizes as follows: where c^^^^^ denotes the spatio-temporal context.An example of a system such as COOL-CHIC mentioned above is illustrated in FIG. 6. Cool-Chic is a Coordinate-based Low Complexity Hierarchical Image / video codec that proposes analternative way to encode images and videos with less complexity in comparison withautoencoder-based codecs. It is based on the coordinate-based neural representation where an Docket No.2024P00587WO image is represented as a learned function which maps pixel coordinates to RGB values. This is another kind of hybrid INR as mentioned above. The parameters of the mapping function are sent using entropy coding. At the receiver side, the compressed image is obtained by evaluating the mapping function for all pixel coordinates.COOL-CHIC supplements a coordinate-based neural representation with a hierarchical latentrepresentation (610), which contains most of the information about the image. To handle thelatent representation, COOL-CHIC uses an auto-regressive module (620) estimating theparameters of the latent distribution (630) to compress the latent representation using entropy coding (640). The variables of the hierarchical latent representation are upsampled (650) and concatenatedas a dense 3D representation ^̂ = upsample(^^). In image coding, the RGB value of each pixel^^^ at coordinates i,j from the compressed image is reconstructed (660) by feeding featuresdetermined at i,j of the dense latent representation to the synthesis network. For inter coding,the synthesis module generates (660), for every pixel of the upsampled latent, a set of outputssuch as optical flows between reference frames and the current frame, residual and pixel-wiseweighting. These outputs are used by the Inter coding module (670) to generate the currentframe by motion compensation using already decoded reference frames, weighting theresulting motion compensated predictions, and correcting the resulting prediction with theresidual to reconstruct the image at input coordinates.COOL-CHIC for video compression is built on 2 modules: the COOL-CHIC encoder and theInter coding module. The COOL-CHIC encoder is trained to encode a frame by a compressed version composed of 2 neural networks (Auto-regressive probability model, Synthesis model) along with a set of L2- dimensional learned discrete latent variables (the hierarchical latent representation).In COOL-CHIC, entropy coding relies on the discrete distribution p^(^^) extracted from thelearned continuous distribution “g” modeled as a Laplace distribution g ≅ ℒ^μ^^^, σ^^^^ usingintegration. The MLP learns to estimate the proper expectation and scale parametersμ^^^ , σ^^^ of that Laplace distribution over the non-quantized latent y, conditioned on the contextlatents. Such as the probability of latent feature is: Docket No.2024P00587WO ^^^^^^.^ p^^y^^^^c^^^^ = ^ g(y)dy ,  with g ≅ ℒ^μ^^^, σ^^^^ and μ^^^ , σ^^^ = f^^c^^^^^^^^^^.^After entropy decoding the latent representations ^^, the decoded data is upsampled to highestresolution and concatenated into an L dimension representation, which is given as input to the synthesis module to map the values of all the channels of each pixel with an RGB value in caseof image compression. For video compression the output of the synthesis module is a list oftensors representing the input for the inter coding module responsible of the exploitation of motion compensation. Optical flow An optical flow is an estimation of the motion between two images, volumes, surface or other signals. It can be estimated by various approaches, including end-to-end training, the Lucas- Kanade Method, solving optical flow equation etc. Given an optical flow ^, an input, such asan image ^, can be warped to obtain an estimation ^^ = ^^^^(^, ^) of the image after themotion: ^^(^^, ^^) = ^(^^ + ^(^^), ^^ + ^(^^))Such an estimation may involve interpolation. To generate an estimate of the currently decoded frame from a reference frame, COOL-CHIC uses optical flow in the inter coding module. This use is constrained by the coding configuration used, specifically the type of the frame being encoded: I being intra frameswithout references (the first frame), P being inter frame with one reference (used in low delay)and B being inter frame with two references (used in random access).When using two reference frames, the warping predictions obtained from each one of thereference frames are combined: ^= ^ ⊙ warp(^ref1, ^ref1→^) + (1 − ^) ⊙ warp(^ref2, ^ref2→^)According to the formula above a pixel-wise continuous weighting ^ is applied to blend thetwo warpings, yielding the temporal prediction ^. Docket No.2024P00587WO As described with FIG. 5, the temporal context used to entropy code the latent is currentlymade of latent values in a reference frame which are located around the position of the latentvalue currently being encoded. If there is large movement in the frame, these values may notbe very relevant to predict the latent value currently being encoded, leading to suboptimalencoding in terms of bitrate distortion.Some embodiments provide for improving the entropy coding of the latent values usingtemporal context from reference frames. In some embodiments, an optical flow or motioninformation is used to obtain a more relevant context for entropy coding and thus a more efficient encoding. Embodiments provided herein are described in the case of hybrid INR. However, theembodiments could apply to any latent representation entropy encoding as long as motioninformation between a current frame and a reference frame can be derived.In some embodiments, when using temporal context to encode the latent, an optical flow (ormovement field) is used to modify the temporal latent of a reference frame to compensate formovement. In the following, an encoding method and a corresponding decoding method for ahybrid-INR system are provided.An embodiment where the optical flow is obtained from two past reference frames is alsodescribed, for example when another optical flow is used to decode that reference frame fromanother, earlier reference frame, which is typical of low-delay mode. An embodiment wherethe optical flow is obtained from one future reference frame is also described, when anotheroptical flow is used to decode that reference frame from another past reference frame, which istypical of random-access mode.In some embodiments, this feature can be switched on / off using signaling in the bitstream.The latent values used in the hybrid INR system for image or video compression is the costliestin terms of bitrate. These values typically amount for the majority of the bitstream length.Therefore, using a context that is as informative as possible to predict the latent values is desired.Embodiments described herein provide for getting an informative context by exploiting opticalflows. Docket No.2024P00587WO In the following described embodiment, it is assumed the optical flow is already available. Twoother embodiments describe how to obtain such an optical flow from previously decodedframes.Without loss of generality, embodiments are described when the signal is a 2D video, butembodiments can apply to any dynamic signal, for example dynamic 3D objects. Furthermore,in the embodiments described, the parameters of the probability distribution model are of anautoregressive model. It should be understood that the probability distribution can be any off- the-shelf probability estimator, for example a conditional Gaussian distribution as in COOL- CHIC, a mixture of Gaussian distributions, or a Naïve Bayes model. The parameters of these distributions can be estimated by an auto-regressive neural networks or through other means, such as linear regression or decision tree.In an embodiment, the temporal context from a reference frame is warped to the current frameusing motion information, i.e. latent values from a reference frame are warped based on an optical flow to get a better estimate of what those reference values would be like for the current frame and therefore a more informative context.In a variant illustrated in FIG. 7, the position where the values of the temporal context are readin the latent values of the reference frame is warped. In top part of FIG.7(a), the position of thetemporal context for a current latent value to decode is warped using an optical flow to obtainnew coordinates defining the position, in the space of the latent values, of the temporal context values. FIG. 7(b) displays the position from where the temporal context latents are extractedfrom the reference latent values when entropy encoding or decoding a current latent value. Forsimplicity, in the example here, the optical flow is uniform, and all positions are shifted to theright by two steps but typically the flow is not uniform. The temporal context can then becombined with a spatial context (illustrated on FIG. 7(c) and used to predict the distribution ofthe current latent value of a current frame. According to this variant, only a part of the referenceframe is warped to the current frame. This variant can be used for example when the warpingof the temporal context can be switched on or off on a pixel wise manner or on region-wise manner. Only the region of the reference frame which correspond to region of the current frame for which warped temporal context is used are warped.In another variant illustrated in FIG. 9, all the latent values of the reference frame are warped,and the temporal context is obtained from these warped reference latent values. In this variant, all latent values of the reference frame are warped by the optical flow. To obtain the values of Docket No.2024P00587WO the temporal context to predict the distribution of a current latent value to entropy encode or decode, the of the temporal context are read from the warped latent values of the referenceframe at the position of this current latent value. For illustration purposes, the latent values ofthe reference frame are represented by a binary image. It can be seen as representing the position of an edge of a clock. In practice, these latent values may not constitute an image that is meaningful to a human.FIG 8 shows the latent value encoding / decoding method when using both a spatial context(from the same frame as the currently encoded value) and a temporal context (from a reference frame). Alternatively, only the temporal context may be used. The method from FIG. 8 issimilar to the one illustrated on FIG. 5, except that the temporal context that is used for entropyencoding / decoding a current latent value is a warped temporal context obtained according toone of the variants illustrated with FIG. 7 or FIG. 9.FIG. 10 illustrates an example of a method 1000 for encoding a current frame of an input signalusing a hybrid INR encoder, according to an embodiment. At 1010, the encoder obtains anoptical flow and a reference frame which is one frame from which the latent values will beused as temporal context. The reference is a frame for which at least the latent values for thisframe have been previously reconstructed (encoded and decoded). Also, the optical flow ismotion information determined between two frames, for example between the reference frame and another frame previously reconstructed. At 1020, the latent values of the reference frame are warped by the optical flow to obtain warped temporal latent values. These latent values are located on a grid that is defined over the same space as the pixels of the reference frame. Thus, the optical flow may be applied to them.This may involve upsampling of the reference latent values when the latent representation is amultiresolution representation. This upsampling may for example be done by interpolation function or neural network. Alternatively, if the latent values of the reference frames had beenupsampled earlier, for example when decoding the reference frame, those upsampled latentvalues may be directly warped instead. At 1030, the Hybrid INR model is trained, using the warped temporal latent values to constructthe context used to encode each latent value of the hybrid INR model. Typically, the full hybridINR model is trained together, including the synthesis network, the latent values and the autoregressive model (or other model) yielding a probability distribution over the latent values. In one variant, only the autoregressive model is trained to encode a set of precomputed latent Docket No.2024P00587WOvalues. In some variants, training may involve quantization aware training which takes intoaccount quantization of the latent values in the training. The training procedure may also involve steps to account for the coding of the probability model, such as quantization, pruning etc., for example by applying these steps, using additional noise and / or adding specific terms to the loss function such as the rate of the network parameters. At 1040, the latent values of the current frame are coded by an entropy encoder, based on the discrete probability distribution over the quantized latent features given by the autoregressivemodel (or another model). This may involve quantization of the latent features if this was notdone before. The warped latent values are used to determine the temporal context used forpredicting the probability of each latent value of the current frame to entropy encode. At 1050, the probability model sed for entropy encoding the latent values of the current frameis encoded in a bitstream. If the model involves a neural network, this is typically done by aneural compression codec such as Neural Network Coding (NNC) / ISO / IEC 15938-17 or MPEG-7 part 17 or simply by quantizing the weights and / or pruning some neurons from the network. At 1060, the synthesis network is encoded in a bitstream. This is typically done by a neural compression codec such as Neural Network Coding (NNC) / ISO / IEC 15938-17 or MPEG-7 part 17 or simply by quantizing the weights and / or pruning some neurons from the network. Also, any high-level syntax required to set the decoder is encoded in the bitstream, for exampleinformation relating to the mapping functions used to map input coordinates to the latentrepresentation. In some variants, the optical flow may be optimized jointly with the INR model. This may allow for a more efficient encoding. In that case the optical flow needs to be encoded in the bitstream as well. This may also be done by a neural compression codec such as Neural Network Coding (NNC) / ISO / IEC 15938-17 or MPEG-7 part 17 or by quantizing the flow. The flow may be encoded by an INR model. Alternatively, a residual with respect to an input optical flow may be optimized and / or encoded in a bitstream. In another variant, the temporal latent values may be optimized jointly with the model, and an optical flow computed based on this optimized temporal latent value and the latent values of the reference frame. In some variants, warping may involve more than one reference frame. In that case, one or more optical flows may be used. Warping can be done from multiple frames, for example by performing warping individually for each reference frame and then combining the prediction, Docket No.2024P00587WO for example by using a neural network, by doing a weighted average where the weights may be transmitted in the bitstream or computed through other means or by doing interpolation. Another example approach to perform warping from multiple frames is to directly do the warping on multiple frames. For example, to do warping using a weighted average, the following formula may be used: ^^ = ^ ⊙ warp(^^ref1, ^ref1→^) + (1 − ^) ⊙ warp(^^ref2, ^ref2→^)FIG. 11 illustrates an example of a method 1110 for decoding / reconstructing a current frameof a signal from a bitstream according to an embodiment. At 1110, the optical flow isobtained. The optical flow may for example be decoded from the bitstream, obtained from elsewhere or be reused or modified from another frame.At 1120, the auto-regressive model is obtained. It may for example be decoded from thebitstream.At 1130, latent values are obtained for at least one already decoded frame: the reference frame.The reference frame is a frame that is available at the decoder, it is a previously reconstructedframe. These latent values are warped using the obtained optical flow to obtain warped reference latent values. At 1140, latent values of the current frame are then decoded progressively. This may for example be done purely sequentially, or in a wavefront manner. For each latent value, the temporal context for a current latent value to decode is obtained by looking up values of the warped reference latent values. The temporal context may be combined with a spatial context, i.e. already decoded latent values from the current frame. This combination may for examplebe a concatenation, a sum, an average etc. The obtained context is used as input to theautoregressive model. The autoregressive model outputs the parameters of the distribution overthe current latent value to decode. Those parameters are used by the entropy decoder to obtainthe current latent value from the bitstream. In a variant, warping may be done from the latent values of multiple frames rather than one. The most likely case is to perform warping from latent values from one past reference and one future reference frame, as in random access mode. The decoded latent features can then be used to decode the current frame. At 1150, the synthesis network is obtained. This is typically done by decoding it from the Docket No.2024P00587WO bitstream, but it may already be available to the decoder, for example because it has been used for another part of the signal or is transmitted separately.At 1160, the current frame is reconstructed using the decoded latent values and the synthesismodel (INR model). The decoded features (latent values) may be upsampled. This may requireobtaining an upsampling network from the bitstream. The upsampled features are used by thesynthesis network to decode the current frame. For 2D images or video, the output of the synthesis network would typically be pixel colors. For 3D scene, the output would typically be color and density of a voxel. For a 3D surface (e.g. hologram), it may be a signed or unsigned distance to the surface. In a particular variant, the decoded latent features can be used to decode the current frame as in the COOL-CHIC system. In this variant, at 1160, the upsampled features are used by the synthesis network to decode the optical flows, residual and pixels wise weighting. Reconstructing the current frame then comprises motion compensation using already decoded reference frames, weighting the resulting motion compensated prediction, and correcting theobtained result with the residual to generate the current frame.In an embodiment, the optical flow used in the embodiments described in relation with FIG.5and 7-11 is an optical flow from two past reference frames. In this embodiment, the opticalflow is obtained from two past reference frames. This embodiment is particularly suited to low- delay modes, where only past frames are available when decoding the current frame.Given two past frames at times ^ − 1 and ^ − 2, where ^ is the index of the current frame, theoptical flow of frame ^ − 1 with respect to frame ^ − 2 can be computed by any off-the-shelfapproach, including the methods described above. This optical flow can then be used to warpthe latent values of frame ^ − 1 from the frame at ^ − 1 to the frame at ^, to encode the latentvalues of frame ^, as described above. In this variant, the optical flow can be determined on thedecoder side using reconstructed frames at ^ − 1 and ^ − 2.Sometimes, one optical flow may already have been generated between frames ^ − 1 and ^ −2, as part of the decoding of frame ^ − 1, for example to predict the values of frame ^ − 1 fromthese of frame ^ − 2 and possibly to encode the residual only. This is for example the case inthe COOL-CHIC system. In a variant, this optical flow may be used directly to warp the latentvalues of frame ^ − 1 to encode the latent values of frame ^, as described above. Docket No.2024P00587WO Temporal scaling may be needed in case the reference frames are not equally spaced, forexample for past references at times ^ − and ^ − ^^, the optical flow may be scaled by^^ / (^^ − ^^).In another embodiment, the optical flow used in the embodiments described in relation withFIG.5 and 7-11 is an optical flow between two reference frames: one past reference frame and one future reference frame. This embodiment is particularly suited to random access modes, where both future and past frames are available when the current frame being decoded is a B- Frame.Given one past frame at times ^ − 1 and one future frame at ^ + 1 (in terms of display order),where ^ is the index of the current frame, the optical flow of frame ^ + 1 with respect to frame^ − 1 can be computed by any off-the-shelf approach, including the methods described above.This optical flow can then be halved to estimate the optical flow from frame ^ − 1 to frame ^.This flow may then be used to warp the latent values of frame ^ − 1 to encode the latent valuesof frame ^, as described above.In another variant, the optical flow from the latent values of frame ^ + 1 with respect to thelatent values of frame ^ − 1 can be computed by any off-the-shelf approach, including themethods described above. This optical flow can then be halved to estimate the optical flowfrom frame ^ − 1 to frame ^ and used to warp the latent values of frame ^ − 1 to encode thelatent values of frame ^, as described above. Both variants may also be combined.Sometimes, one optical flow may already have been generated between frames ^ + 1 and ^ −1, as part of the decoding of frame ^ + 1, for example to predict the values of frame ^ + 1 fromthese of frame ^ − 1 and possibly to encode the residual only. This is for example the case inthe COOL-CHIC system. In a variant, this optical flow may be halved and used directly towarp the latent values of frame ^ − 1 to encode the latent values of frame ^, as described above.In all the embodiments above, the optical flow may alternatively be available or computed asthe optical flow of frame ^ + 1 with respect to frame ^ − 1 .Any of these two optical flows may also be inverted. For example, the optical flow from thelatent values of frame ^ + 1 with respect to the latent values of frame ^ − 1 may be inverted toobtain an optical flow from the latent values of frame ^ − 1 with respect to the latent values offrame ^ + 1. Docket No.2024P00587WOWhen the reference frames are not located at times ^ − 1 and ^ + 1 but at times ^ − ^^ and ^ +^^, it is possible to create variants of all the above methods. In that case, rather than halvingthe optical flow, it should be multiplied by a factor depending on and ^^, for example^^ / (^^ + ^^) for a flow of frame ^ + ^^ with respect to frame ^ − and ^^ / (^^ + ^^) for theopposite. In the embodiments described above, warping is done using optical flow. In other variants, any motion information that allows to perform motion compensation between the reference frame and the current frame can be used.In an embodiment, the usage of the warped temporal context may be enabled or disabled at aframe level for example. Warping reference latent values may not always lead to a better bitratedistortion tradeoff. This variant allows the encoder to choose whether to use this feature or notfor each frame. This must be signaled in the bitstream, for example using an additional bitwhose value enables or disabled the use of warped temporal context for latent values of thecurrent frame. One possible encoding procedure is as follows:• The current frame is encoded twice: with warped reference latent values and withoutwarping them.• The encoder chooses one of the encodings: typically, the one with the better rate / distortiontrade-off, but other options are possible such as the lowest bitrate.• The encoder adds one bit in the bitstream signaling whether latent values of the currentframe are entropy encoded using warped reference latent values or without using warped reference latent, e.g.1 for warped reference latent values and 0 without. This bit can be denoted the warped temporal context bit.• The frame is encoded in the bitstream using the chosen encoding configuration.In a variant, choosing to warp the temporal feature chosen may be done by a machine learning algorithm or any other heuristic or expert defined algorithm. In a variant, the choice of the solution is optimized over a group of frames rather than a single one, for example a GOP or a full video. In another variant, the choice is made without encoding Docket No.2024P00587WO the frame using both methods. In that case, the choice may for example be predetermined (e.g. only warp the temporal context for frames number 2,3,5,7,8,9 in a GOP), chosen by a machine learning algorithm or any other heuristic. The frame can be decoded as follows:• The decoder reads the warped temporal context bit in the bitstream.• The decoder selects the appropriate decoding algorithm (with or without warping thetemporal context) based on the value of this bit. Using the example given above, if the bit is 1 decoding with warped temporal context is selected and if the bit is 0 decoding without warping them is selected.• The decoder decodes the frame using the appropriate decoding algorithm.In a variant, the warped temporal context configuration can be chosen to be identical for multiple frames at the same time, for example for a GOP, for a full video or for a set numberof frames. In that case, a single bit is used in the bitstream to signal this choice that affectsmultiple frames. In a variant, the number of frames for which this choice is valid can be chosen by the encoder and is also included in the bitstream. The encoding / decoding algorithm can easily be adapted for the variants listed above if necessary.In an embodiment, illustrated in FIG. 12, in a transmission context between two remote devicesA and B over a communication network NET, the device A comprises a processor in relation with memory RAM and ROM which are configured to implement a method for encoding a signal according to any one of the embodiments described herein and the device B comprises a processor in relation with memory RAM and ROM which are configured to implement amethod for reconstructing the signal according to any one of the embodiments described herein.In accordance with an example, the network is a broadcast network, adapted tobroadcast / transmit a coded signal from device A to decoding devices including the device B.FIG. 13 shows an example of the syntax of a signal transmitted over a packet-based transmission protocol. Each transmitted packet P comprises a header H and a payload PAYLOAD. In some embodiments, the payload PAYLOAD may comprise data representativeof an input signal encoded according to any one of the embodiments described above. Thepayload can also comprise any signaling as described above. For example, the signal comprisesthe warped temporal context bit mentioned above. Docket No.2024P00587WO One or more embodiments provide a computer program comprising instructions which when executed by one or more processors cause such processors to perform the encoding and / ordecoding methods according to any of the embodiments described above. One or moreembodiments also provide a computer readable storage medium having stored thereon instructions for encoding or decoding data according to the methods described above. One or more embodiments provide a computer readable storage medium having stored thereon data generated according to the methods described above. One or more embodiments also provide a method and apparatus for transmitting or receiving data generated according to the methods described above. The embodiments described herein may be implemented in, for example, a method or a process, an apparatus, a software program, a data stream, or a signal. Even if only discussed in the context of a single form of implementation (e.g., as a method), the implementation of such features may also be implemented in other forms. An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. Corresponding methods may be implemented in, for example, a processor. The various embodiments described herein provide at least the following features, devices oraspects, alone or on any combination, across various claim categories and types:i. Encoding, into coded data, syntax elements that can enable the decoder to decodethe coded data, according to any of the embodiments described herein. ii. A bitstream that includes one or more of the described syntax elements, orvariations thereof, whether transmitted, stored, or otherwise made available. iii. Creating, transmitting, receiving, and / or decoding of the bitstream.iv. An electronic device (e.g., TV, set-top box, mobile phone, tablet, etc.) that tunesa channel to receive a bitstream or that receives such bitstream over the air. Theelectronic device decodes the syntax elements from the bitstream, and, optionally, displays (e.g., via a monitor or other type of display) a resulting image / video or 3D object or 3D scene. Various numeric values are used in the present application. Such specific values are for example purposes and the embodiments described are not limited to these specific values. Various methods are described herein, and such methods comprise one or more stepsor actions for achieving the described method. Unless a specific order of steps or actions is Docket No.2024P00587WO required for the proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined. Additionally, terms such as “first”, “second”, etc. may be used in various embodiments to modify an element, component, step, operation, etc., for example, a “first decoding” and a “second decoding”. Use of such terms does not imply an order to the operations unless specifically required. The present disclosure may refer to “determining” various pieces of information. Determining information may include one or more of, for example, estimating, calculating,predicting, or retrieving (e.g., from memory) the information.The present disclosure may refer to “accessing” various pieces of information. Accessing information may include one or more of, for example, receiving, retrieving (e.g.,from memory), storing, moving, copying, calculating, determining, predicting, or estimatingthe information. Similarly, the present disclosure may refer to “receiving” various pieces of information. Receiving information may include one or more of, for example, accessing or retrieving (e.g., from memory) the information. “Decoding,” as used herein, encompasses all or part of the processes performed, forexample, on an encoded sequence to produce an output suitable for display. In someembodiments, such processes include one or more of the processes typically performed by a decoder, for example, entropy decoding, inverse quantization, etc. Whether the phrase “decoding process” is intended to refer to a subset of operations or generally to the broader decoding process will be clear based on the context of the specific description and will be well understood by those skilled in the art. “Encoding,” as used herein, encompasses all or part of the processes performed, forexample, on input video data an order to produce an encoded bitstream. Additionally, the terms “reconstructed” and “decoded” may be used interchangeably, the terms “encoded” or “coded” may be used interchangeably, the terms “image,” “picture,” “sub-picture,” “slice,” and “frame” may be used interchangeably, and the terms “pixel” and “sample” may be used interchangeably. The present disclosure refers to information, for example, syntax elements, that can betransmitted or stored. Such information can be packaged or arranged in a variety of manners,including for example manners common in image, video or immersive standards such asputting the information into a sequence parameter set (SPS), a picture parameter set (PPS), anetwork abstraction layer (NAL) unit, a header (for example, a NAL unit header, or a sliceheader), or an SEI message. Other manners are also available, including, for example, manners Docket No.2024P00587WO that are common for system level or application-level standards such as signaling the information into one or more of the following: i. session description protocol (SDP), for example as described in RFCs and / or usedin conjunction with real-time transport protocol (RTP) transmission.ii. hypertext transfer protocol (HTTP) live Streaming (HLS) manifest transmittedover HTTP. iii. dynamic adaptive streaming over HTTP (DASH) media presentation description(MPD) descriptors, for example as used in DASH and transmitted over HTTP.iv. RTP header extensions, for example as used during RTP streaming.v. International Organization for Standardization (ISO) base media file format, forexample, as used in Omnidirectional MediA Format (OMAF).As used herein, “signal” and “signaling” refer to, among other things, indicatinginformation to a decoder. For example, in some embodiments the encoder signals aquantization matrix for de-quantization, whereby the same parameter is used for both encoding and decoding. In some embodiments, the signaling may be explicit, such that information (e.g.,a particular parameter) is transmitted to the decoder enabling the decoder to use the sameparticular parameter. In some embodiments, the signaling may be implicit, in that theinformation (e.g., a particular parameter) is indicated based on other information at ortransmitted to the decoder or derived or selected by the decoder based on information availableat the decoder. By not transmitting the information (e.g., the particular parameter), a bit savingsis thus realized in some embodiments. In some embodiments, one or more syntax elements orflags are used to signal information to a decoder. While the preceding relates to the verb formof the word “signal”, the word “signal” can also be used herein as a noun. In some embodiments, signals may be produced that are formatted to carry informationthat may be stored or transmitted. Such information may include, for example, instructions forperforming a method, or data produced by one of the described implementations (e.g., abitstream of a described embodiment). Such a signal may be formatted, for example, as an electromagnetic wave or as a baseband signal. The formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information that the signal carries may be, for example, analog or digital information. The signal may be transmitted over a variety of different wired or wireless links and may be stored on a processor- readable medium. Docket No.2024P00587WO It is to be understood that use of any of the following “ / ”, “and / or”, and “at least one of” is intended to encompass all possible selections of listed items, taken either individually or in any combination thereof. While specific embodiments have been described in the foregoing description in connection with the accompanying drawings, it should be understood that embodiments described herein are examples only and should not be taken as limiting the scope of the present disclosure or the following claims. Although features and elements are described herein in particular combinations, those of ordinary skill in the art will appreciate that such features or elements may be used alone or in any combination with the other features and elements. It is understood, therefore, that the overall teachings of the present disclosure are not limited to the particular embodiments, implementations, and examples disclosed herein, but are intended to cover variations, modifications, and alternatives as defined by the appended claims and any and all equivalents thereof.

Claims

Docket No.2024P00587WO CLAIMS What is claimed is:

1. A method for encoding a signal using a hybrid Implicit Neural Representation network,the hybrid Implicit Neural Representation network comprising at least one mappingfunction that maps input coordinates of the signal to a latent of features and an ImplicitNeural Representation (INR) network that takes as input features of the latent andoutputs values of the signal at the input coordinates, the method comprises: mapping input coordinates of a current frame of the signal to a current latent offeatures, obtaining parameters of the INR network for the current frame, warping at least one reconstructed part of a reference latent of features of a referenceframe of the signal to the current frame, determining a temporal context for at least of one value of the current latent of features from the warped at least one reconstructed part of the reference latent,entropy encoding the at least one value of the current latent of features based at least on the temporal context, and encoding the parameters of the INR network.

2. A method for decoding a signal using a hybrid Implicit Neural Representation network, thehybrid Implicit Neural Representation network comprising at least one mapping function thatmaps input coordinates of the signal to a latent of features and an Implicit Neural Representation (INR) network that takes as input features of the latent and outputs values of the signal at the input coordinates, the method comprises: obtaining at least one part of a warped reference latent from at least one part of at least one previously reconstructed latent of features of a previously reconstructed frame of asignal to reconstruct, by warping to a current frame of the signal, the at least one part of the at least one previously reconstructed latent,determining a temporal context for at least of one value of a current latent of featuresof the current frame from the at least one part of the warped reference latent,Docket No.2024P00587WO entropy decoding the at least one value of the current latent of features based at least on the temporal context, decoding the INR network for the current frame,reconstructing the current frame using the at least one value of the current latent offeatures, the at least one mapping function and the INR network.

3. An apparatus, comprising one or more processors, wherein the one or more processors isoperable to implement a hybrid Implicit Neural Representation network, the hybrid ImplicitNeural Representation network comprising at least one mapping function that maps inputcoordinates of a signal to a latent of features and an Implicit Neural Representation (INR)network that takes as input features of the latent and outputs values of the signal at the input coordinates, the one or more processors being operable to: map input coordinates of a current frame of the signal to a current latent of features,obtain parameters of the INR network for the current frame, warp at least one reconstructed part of a reference latent of features of a reference frame of the signal to the current frame, determine a temporal context for at least of one value of the current latent of features fromthe warped at least one reconstructed part of the reference latent, entropy encode the at least one value of the current latent of features based at least on thetemporal context; and encode the parameters of the INR network.

4. An apparatus, comprising one or more processors, wherein the one or more processors isoperable to implement a hybrid Implicit Neural Representation network, the hybrid ImplicitNeural Representation network comprising at least one mapping function that maps inputcoordinates of a signal to a latent of features and an Implicit Neural Representation (INR)network that takes as input features of the latent and outputs values of the signal at the input coordinates, the one or more processors being operable to: obtain at least one part of a warped reference latent from at least one part of at least one previously reconstructed latent of features of a previously reconstructed frame of aDocket No.2024P00587WO signal to reconstruct, by warping to a current frame of the signal, the at least one part of the at least one previously reconstructed latent,determine a temporal context for at least of one value of a current latent of features of the current frame from the at least one part of the warped reference latent, entropy decode the at least one value of the current latent of features based at least on the temporal context, decode the INR network for the current frame, reconstruct the current frame using the at least one value of the current latent of features, the at least one mapping function and the INR network.

5. The method of claim 2, further comprising decoding a probability model used in the entropy decoding of the at least one latent value.

6. The method of any one of claim 1, 2 or 5, wherein warping to the current frame, the at least one part of the at least one previously reconstructed latent uses motion information obtainedfrom two previously reconstructed frames.

7. The method of claim 6, wherein motion information is an optical flow between the two previously reconstructed frames obtained from an output of an Implicit NeuralRepresentation model or decoded from a bitstream.

8. The method of any one of claims 1, 2 or 5-7, wherein obtaining the at least one part of a warped reference latent from at least one part of at least one previously reconstructed latentcomprises warping another previously reconstructed latent of features of another previouslyreconstructed frame to the current frame and weighting the warped previouslyreconstructed latent and the other warped previously reconstructed latent.

9. The method of any one of claims 1, 2 or 5-8, wherein warping to a current frame of thesignal, the at least one part of the at least one previously reconstructed latent comprisesupsampling at least one part of the at least one previously reconstructed latent.

10. The method of claim 2, further comprising decoding a syntax element indicating whether warping is done or not.Docket No.2024P00587WO 11. The method of any one of claims 2 or 5-10, wherein entropy decoding the at least one value of the current latent is based on a spatial context determined from previously reconstructed values of the current latent.

12. The apparatus of claim 4, wherein the one or more processors are further configured todecode a probability model used in the entropy decoding of the at least one latent value.

13. The apparatus of any one of claim 3, 4 or 12, wherein warping to the current frame, the atleast one part of the at least one previously reconstructed latent uses motion informationobtained from two previously reconstructed frames.

14. The apparatus of claim 13, wherein motion information is an optical flow between the twopreviously reconstructed frames obtained from an output of an Implicit Neural Representation model or decoded from a bitstream.

15. The apparatus of any one of claims 3, 4 or 12-14, wherein obtaining the at least one part ofa warped reference latent from at least one part of at least one previously reconstructed latent comprises warping another previously reconstructed latent of features of another previously reconstructed frame to the current frame and weighting the warped previouslyreconstructed latent and the other warped previously reconstructed latent.

16. The apparatus of any one of claims 3, 4 or 12-15, wherein warping to a current frame of the signal, the at least one part of the at least one previously reconstructed latent comprises upsampling at least one part of the at least one previously reconstructed latent.

17. The apparatus of claim 4, further comprising decoding a syntax element indicating whetherwarping is done or not.

18. The apparatus of any one of claims 4 or 12-17, wherein entropy decoding the at least onevalue of the current latent is based on a spatial context determined from previously reconstructed values of the current latent.Docket No.2024P00587WO 19. A non-transitory computer readable storage medium having stored thereon a bitstream comprising data encoding parameters of a hybrid Implicit Neural Representation network, the hybrid Implicit Neural Representation network comprising at least one mapping function thatmaps input coordinates of a signal to a latent of features and an Implicit Neural Representation(INR) network that takes as input features of the latent and outputs values of the signal at theinput coordinates, and data encoding values of the latent of features wherein the values areentropy-encoded based on at least one temporal context determined from a warpedreconstructed part of a reference latent of features.

20. The non-transitory computer readable storage medium of claim 19, wherein the bitstreamcomprises a probability model used for entropy encoding of the values.

21. The non-transitory computer readable storage medium of claim 19 or 20, wherein thebitstream comprises a syntax element indicating whether warping is done or not to determinethe temporal context.

22. A non-transitory computer readable medium having stored thereon program code instructions which when executed by one or more processors cause the one or more processorsto perform the method according to any one of claims 1, 2 or 5-11.

Citation Information

Patent Citations

  • Inter coding using deep learning in video compression

    WO2024006167A1