Spatiotemporal context for hybrid INR network

By training INR networks to determine distributions of latent variables within shaped context regions and encoding these distributions, the method enhances data compression efficiency and reduces complexity.

WO2026041528A1PCT designated stage Publication Date: 2026-02-26INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/073324
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-21
Filing Date
2025-08-14
Publication Date
2026-02-26

AI Technical Summary

Technical Problem

Current approaches for encoding and decoding data using implicit neural representation (INR) networks are inefficient in exploiting spatial and temporal redundancies across latent variables, leading to high complexity and bitrate in data compression.

Method used

The proposed method involves training an INR network to produce data regions from latent variables, determining distributions of these variables using a shaped context region, and coding them into a bitstream, while also encoding the network's parameters based on learned distributions to optimize compression efficiency.

Benefits of technology

This approach reduces the bitrate and complexity of data compression by effectively utilizing spatial and temporal redundancies, achieving improved encoding and decoding performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025073324_26022026_PF_FP_ABST
    Figure EP2025073324_26022026_PF_FP_ABST
Patent Text Reader

Abstract

Apparatuses and methods are disclosed for encoding and decoding data. Techniques disclosed provide for the encoding of a data region of a frame. The encoding includes training an INR network to produce the data region from latent variables. And, determining distributions of the latent variables using respective contexts, constructed based on latent variables located within a shaped context region. Then, coding into a bitstream the latent variables based on the determined distributions and further coding into the bitstream the parameters of the trained INR network. Techniques disclosed also provide for decoding the data region. The decoding includes decoding from the bitstream the parameters of the INR network, determining the distributions of the latent variables using the respective contexts, and decoding from the bitstream the latent variables based on the determined distributions. Using the parameters of the INR network, the INR network produces the data region from the latent variables.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SPATIOTEMPORAL CONTEXT FOR HYBRID INR NETWORKCROSS REFERENCE TO RELATED APPLICATIONS[1] This application claims the benefit of European Application No. 24306384.9, filed onAugust 21, 2024, which is incorporated herein by reference in its entirety.BACKGROUND[2] An implicit neural representation (INR) network is a neural network that is trained torepresent a specific signal; the INR network is trained to predict a sample value of the signalwhen presented with the sample’s spatial coordinates. A hybrid INR network improves theINR network’s capability to learn local patterns of the signal. To that end, sample coordinatesof the signal are first mapped into latent variables and the INR network is trained to representthe signal based on these latent variables. When the hybrid INR network is applied to signalcompression, the latent variables represent the signal in the bitstream. To efficiently entropycode the latent variables, their respective distributions should be used. Current approachesestimate the latent variables’ respective distributions based on context. The manner in whichsuch context is constructed is instrumental in exploiting spatial and temporal redundancies across the latent variables. SUMMARY[3] Aspects disclosed in the present disclosure describe methods for encoding data. Thesemethods comprise encoding, into a bitstream, a data region of a currently encoded frame. The encoding includes training an INR network to produce the data region from latent variables, where the training determines parameters of the INR network. Additionally, determining distributions of the latent variables, where a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, coding into the bitstream the latent variables based on thedetermined distributions; and coding into the bitstream the parameters of the INR network.Aspects disclosed in the present disclosure also describe methods for decoding data. These methods comprise decoding, from a bitstream, a data region of a currently decoded frame. The decoding includes decoding from the bitstream parameters of an INR network that istrained to produce the data region from latent variables. Additionally, determiningdistributions of the latent variables, where a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, decoding from the bitstream the latent variables based on thedetermined distributions. Using the parameters of the INR network, the INR networkproduces the data region from the latent variables.[4] Aspects disclosed in the present disclosure describe apparatuses for encoding data.The apparatuses comprise at least one processor and memory storing instructions. Theinstructions, when executed by the at least one processor, cause the apparatuses to encode,into a bitstream, a data region of a currently encoded frame. The encoding includes training an INR network to produce the data region from latent variables, where the training determines parameters of the INR network. Additionally, determining distributions of the latent variables, where a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, coding into the bitstream the latent variables based on the determined distributions; and coding into the bitstream the parameters of the INR network. Aspects disclosed in the present disclosure also describe apparatuses for decoding data. The apparatuses comprise at least one processor and memory storing instructions. The instructions, when executed by the at least one processor, cause the apparatuses to decode, from a bitstream, a data region of a currently decoded frame. The decoding includes decoding from the bitstream parameters of an INR network that is trained to produce the data region from latent variables.Additionally, determining distributions of the latent variables, where a distribution of a latentvariable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, decoding from the bitstream the latent variablesbased on the determined distributions. Using the parameters of the INR network, the INRnetwork produces the data region from the latent variables[5] Aspects disclosed in the present disclosure describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for encoding data. These methods comprise encoding, into a bitstream, a data region of a currently encoded frame. The encoding includes training an INR network to produce the data region from latent variables, where the training determines parameters of the INR network. Additionally, determining distributions of the latent variables, where a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, coding into the bitstream the latent variables based on the determined distributions; and coding into the bitstream the parameters of the INR network. Aspects disclosed in the present disclosure also describe a non-transitory computer-readable medium comprising instructions executable by at least one processor to perform methods for decoding data. These methods comprise decoding, from a bitstream, a data region of a currently decoded frame. The decoding includes decoding from the bitstream parameters of an INR network that is trained to produce the data regionfrom latent variables. Additionally, determining distributions of the latent variables, where adistribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region. Then, decoding from the bitstream the latent variables based on the determined distributions. Using the parameters of the INRnetwork, the INR network produces the data region from the latent variables[6] This Summary is provided to introduce a selection of concepts in a simplified formthat is further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to limitations that solve any or all disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS[7] FIG. 1 is a block diagram of an example system, according to which aspects of thepresent embodiments can be implemented.[8] FIG. 2 is a diagram illustrating an example INR network, according to which aspectsof the present embodiments can be implemented.[9] FIG. 3 is a block diagram of an example video encoder, according to which aspects ofthe present embodiments can be implemented.

[0010] FIG. 4 is a block diagram of an example video decoder, according to which aspects ofthe present embodiments can be implemented.

[0011] FIG. 5 is a diagram illustrating a hybrid INR network, according to which aspects ofthe present embodiments can be implemented.

[0012] FIG. 6 is a block diagram illustrating an example video encoder applying a hybridINR network, according to which aspects of the present embodiments can be implemented.

[0013] FIG. 7 is a block diagram illustrating an example video decoder applying a hybridINR network, according to which aspects of the present embodiments can be implemented.

[0014] FIG. 8 is a block diagram illustrating prediction of a latent variable distribution, usinga spatial context, according to which aspects of the present embodiments can be implemented.

[0015] FIG. 9 is a block diagram illustrating an example video encoder applying a predictivehybrid INR network, according to which aspects of the present embodiments can be implemented.

[0016] FIG. 10 is a block diagram illustrating an example video decoder applying apredictive hybrid INR network, according to which aspects of the present embodiments can be implemented.

[0017] FIG. 11 is a block diagram illustrating prediction of a latent variable distribution,using a spatiotemporal context, according to which aspects of the present embodiments can be implemented.

[0018] FIG. 12 is a diagram illustrating configurations of context region shapes, according towhich aspects of the present embodiments can be implemented.

[0019] FIG. 13 is a diagram illustrating other configurations of context region shapes,according to which aspects of the present embodiments can be implemented.

[0020] FIG. 14 is a flowchart of an example method for encoding data, according to whichaspects of the present embodiments can be implemented.

[0021] FIG. 15 is a flowchart of an example method for decoding data, according to whichaspects of the present embodiments can be implemented. DETAILED DESCRIPTION

[0022] Apparatuses and methods described herein for encoding and decoding data applying ahybrid INR network or a predictive hybrid INR network. A system for processing and displaying content, with which various aspects and examples described herein may be implemented, is generally described in reference to FIG. 1, followed by a description of the aspects of the present disclosure in reference to FIGS.2-15.

[0023] FIG. 1 is a block diagram of an example system, according to which aspects of thepresent embodiments can be implemented. The system 100 may be an electronic device including, for example, a personal computer, laptop computer, mobile phone, tablet computer, multimedia set-top box, digital television receiver, personal video recording system, connected home appliance, vehicle control and / or entertainment system, and server. One or more elements of the system 100, singly or in combination, may be implemented asan integrated circuit (IC), multiple ICs, and / or discrete components. For example, in oneembodiment, the processing, encoding and / or decoding elements of system 100 aredistributed across multiple ICs and / or discrete components. In some embodiments, the system100 is communicatively coupled to and / or in communication with other systems or devices, via, for example, a communications bus or dedicated input / output ports.

[0024] One or more of the elements of system 100 may be provided within an integratedhousing, with such elements being interconnected and able to transmit data therebetween using any suitable connection arrangement 115 generally known in the art, including, for example, an internal bus (e.g., I2C bus), wiring, and printed circuit boards.

[0025] The system 100 includes at least one processor 110 configured to execute instructionsfor implementing the embodiments described herein, including signal / data coding andprocessing. The processor 110 may be a general-purpose processor or microprocessor, digitalsignal processor (DSP), one or more microprocessors in association with a DSP core, a controller, a microcontroller, application specific integrated circuits (ASICs), fieldprogrammable gate arrays (FPGAs), a state machine, and the like. The processor 110 mayinclude at least one central processing unit (CPU), embedded memory, input and output interfaces, and other circuitries.

[0026] The system 100 includes at least one memory 120, for example, a volatile memorydevice and / or a non-volatile memory device. The system 100 includes a storage device 140,that may be or include non-volatile memory and / or dynamic volatile memory, including EEPROM, ROM, PROM, RAM, DRAM, SRAM, DDR, flash, magnetic disk drives, solidstate drives (SSD) and / or optical disk drives. The storage device 140 may be or include, forexample, an internal storage device, an attached storage device, and / or a network accessiblestorage device. Although shown separately, the memory 120 and the storage device 140 maybe collocated, integrated together, or otherwise combined.

[0027] The system 100 includes an encoder / decoder module 130 configured to process videodata and to provide encoded video data or decoded video data. The encoder / decoder module130 may include one or more processors and / or memory (not shown). Although FIG. 1depicts the encoder / decoder module 130 as a separate element of system 100, it will be understood that the processor 110 and the encoder / decoder module 130 may be collocated and / or integrated together as a combination of hardware and / or software, e.g., in an electronicpackage or chip. The encoder / decoder module 130 may be or include one or more modulesthat may be included in one or more separate devices that perform encoding and / or decoding functions.

[0028] Instructions for execution by the processor 110 and / or the encoder / decoder module130 may be stored in the storage device 140 and subsequently loaded into memory 120 forexecution by the processor 110. In some embodiments, one or more of processor 110,memory 120, storage device 140, and encoder / decoder module 130 may store one or moreitems when performing the processes disclosed herein. Such items may include input video,decoded video or portions thereof, bitstreams, matrices, variables, operational logic, and intermediate and / or final results from processing of equations, formulas, or operations.

[0029] In some embodiments, the memory of the processor 110 and / or the encoder / decodermodule 130 is used to store instructions and / or provide working memory for video encodingand decoding functions. In some embodiments, memory external to the processor 110 and / orthe encoder / decoder module 130 (e.g., the memory 120 and / or the storage device 140) is used for one or more of these functions and / or, for example, to store the operating system of a television.

[0030] The system 100 may obtain or receive information via one or more input devices,interfaces, and / or ports as indicated in input block 105. Examples of the input devices includea radio frequency (RF) device for transmitting and / or receiving RF signals over various media, for example, RF signals received over the air from a broadcaster; component video (COMP) inputs; a Universal Serial Bus (USB) input; and / or a High-Definition MultimediaInterface (HDMI) input. Other examples include composite video input (not shown). In someembodiments, the input devices are associated with respective input processing elements,e.g., those generally known in the art. For example, the RF device may be associated withelements suitable for selecting a desired frequency (e.g., selecting or band-limiting a signal)or performing error correction on the signal. The USB and / or HDMI inputs may includerespective interface processors and transceivers (or transmitters and receivers) for couplingthe system 100 to other devices via USB and / or HDMI ports or connections. Various formsof input processing may be implemented, for example, by and / or within a separate input processing device or the processor 110.

[0031] The system 100 includes a communication interface 150 that enables wired and / orwireless communication with other devices, e.g., via a communication channel 190. The communication interface 150 may include one or more transceivers, modems, network cardsand the like. The communication channel 190 may be or include wired and / or wirelessmediums.

[0032] In some embodiments, data may be streamed to the system 100 via wired and / orwireless networks. Examples of such wireless networks include cellular, Bluetooth or Wi-Fi(e.g., IEEE 802.11) networks. The wired and / or wireless networks may include one or morebase stations (e.g., cellular base stations, access points, etc.), and / or user equipment (e.g. cellular user equipment, stations, etc.), and / or other network elements that communicate with the system 100 via the communication interface 150 and communication channel 190, whereby the system 100 may obtain data streamed from streaming applications (e.g., OTTservices) via various networks, including the Internet. In some embodiments, data is streamedto the system 100 via the input block 105 (e.g., using a set-top box that delivers data via theHDMI connection or the RF connection). In some embodiments, data is received by thesystem 100 in a non-streaming manner.

[0033] The system 100 may provide one or more output signals to one or more outputdevices. The output devices may include a display device 165 (e.g., touchscreen display,monitor, etc.), an audio device 175 (e.g., speakers), and other peripheral devices 185, including, for example, a stand-alone DVR, a disk player, a stereo system, a lighting system,and other devices that provide a function based on the output of the system 100. The displaydevice 165 can be for a television, tablet, laptop, mobile phone, head-mounted display, orother device. In some embodiments, control signals are communicated between the system100 and the display device 165, the audio device 175, and / or the peripheral devices 185,enabling device-to-device control with or without user intervention. The output devices maycouple to and / or communicate with the system 100 via dedicated connections via respectivedisplay, audio, and peripheral interfaces 160, 170, 180. Alternatively, the output devices maycouple to and / or communicate with the system 100 via the communication channel 190 and the communication interface 150.

[0034] The display device 165 and the audio device 175 may be collocated, integrated, orotherwise combined with the other components of system 100 in a single unit (e.g., atelevision). Alternatively, the display device 165 and the audio device 175 may be separatefrom one or more of the other components of the system 100. In embodiments in which thedisplay device 165 and the audio device 175 are external components, the output signals may be provided via dedicated outputs and / or connections, including, for example, HDMI ports, USB ports, or COMP outputs.

[0035] Generally, video codecs can be classified into conventional codecs, autoencoders, andoverfitted codecs. Conventional codecs, currently the most common ones, are based onpredictive coding, such as those following the advance video coding (H.264 / AVC), the highefficiency video coding (HEVC), or the versatile video coding (VVC) standards. On the otherhand, an autoencoder is a neural network that is trained to map input images (video frames)into respective sets of latent variables and then map back the sets of latent variables to approximated versions of the input images. During training, the autoencoder parameters are optimized to minimize a rate-distortion cost function. Thus, using the trained autoencoder, encoding an image involves mapping the image into a set of latent variables and decoding the image involves mapping the set of latent variables into a reconstructed version of the image. Since autoencoders are generically trained (that is, their training is based on a large number of images) to compete with the performance of conventional codecs, for example, theircomplexity (the number of parameters and multiplications involved per pixel) is high.Relative to autoencoders, overfitted codecs, using INR networks, offer lower complexitywhile still achieving the performance levels of conventional codecs.

[0036] Representing data via INR networks is a relatively new technology that has onlyrecently been investigated by the computer vision and computer graphics communities. INRnetworks are studied for the applications of compression, where an efficient representation ofsignals (such as, still images, videos, surfaces, and volumes) is required. In contrast to anautoencoder, an INR network is trained to represent a specific signal. Hence, unlikeautoencoders, an INR-based codec is not generic but is adapted (overfitted) to the signal to becoded and is, therefore, more efficient. Moreover, since the INR network is trained to predicta signal based on the signal’s coordinates, the trained network can be applied to progressivelyreconstruct the signal based on any set of coordinates (e.g., a subset of the coordinates used inthe training, or any other set of extrapolated or interpolated coordinates therefrom). Forsimplicity of the presentation, aspects of INR network applications are described herein withrespect to images (e.g., video frames); however, these aspects are extendable to other dataframes, such as those representing surfaces of objects and volumetric data that may bechanging overtime.

[0037] FIG. 2 is a diagram illustrating an example INR network 200. An INR network is aneural network composed of multiple layers, each layer includes multiple nodes (denoted bycircles). Generally, the architecture of a neural network is characterized by the number oflayers, the number of layers’ nodes, and by the way the layers’ nodes are connected. In theexample of FIG. 2, the network 200 has four layers 220, 230, 240, 250 that are fullyconnected. For example, the first layer 220 includes four nodes N11, N12, N13, and N14 thateach receives the coordinate values (^, ^) 210 of a pixel ^ and each outputs an output signalthat, in turn, feeds the nodes of the next layer, N21, N22, N23, and N24. Likewise, the secondlayer 230 includes four nodes N21, N22, N23, and N24, that each receives the output signalsof nodes from the previous layer, N11, N12, N13, and N14, and each outputs an output signalthat, in turn, feeds the nodes of the next layer, N31, N32, N33, and N34. The fourth layer 250includes three nodes N41, N42, and N43 that each receives the output signals of the nodesfrom the previous layer, N31, N32, N33, N34 and each outputs a color component value 260of the pixel ^, respectively, ^, ^, and ^ (or color component values of any other color model,such as ^, ^, and ^).

[0038] Each node in the network 200 represents an operator that generates an output signalbased on the node’s inputs. For example, node N21 of the second layer 230 receives as an input the output signals of nodes N11, N12, N13, and N14, respectively, ^^, ^^, ^^, and ^^. These inputs are translated into an output signal ^^^^that feeds the nodes of the third layer240. A node’s operator can be expressed as follows: where, ^ denotes the number of input signals (i.e., the number of nodes from the previouslayer that connect to the node), ^ = {^^ ∶ ^ = 1 ^^ ^} denotes the node’s input signal vector,^^^^ denotes the node’s output signal, ^ = {^^ ∶ ^ = 0 ^^ ^} denotes the node’s parametervector (or weight vector), and ^ denotes an activation function (e.g., ReLU, Sigmoid, orTanh). The weight vectors (and parameters of the activation functions, if such parametersexist) of respective nodes are collectively referred to as the parameters ^ of the network 200.These parameters ^ are determined through a training process. The network operation,denoted by ^^ , is therefore defined by the network parameters ^.

[0039] Hence, an INR network 200 is trained to predict a pixel value of an image, ^(^, ^),based on the pixel’s coordinates (^, ^) – that is, ^^(^, ^) = (^, ^, ^) (or = (^, ^, ^)).During a training phase of an INR network 200, the parameters ^ (or a subset of them) aredetermined. This is done via an optimization process through which the parameters ^ thatminimize a cost function can be determined. For example, the following cost function can be used:^^^^ = ^(^, ^^) + ^^(^) (2)where, ^ is a distortion measure, measuring the fidelity of the estimated pixel values,provided by ^^, relative to the ground truth, that is, the corresponding pixel values from theoriginal image, denoted by ^. And, where ^ is the resulting bitrate of the encoded parameters^ (e.g., encoded by quantization and entropy coding as discussed with respect to FIG. 3). Atrade-off parameter ^ can be set to determine the balance between ^ and ^. Note that thedistortion measure ^ can be any metric that measure the distance (or similarity) between theoriginal image ^ and its estimated version provided by ^^ , such as a mean squared errormetric or a learned perceptual image patch similarity (LPIPS) metric. For example, a meansquared error metric can be expressed as: where, ^ and ^ are the width and height of the image ^ that the INR network is trained topredict. The optimization of the network parameters ^, according to equation (2), is typicallyperformed by a machine learning optimization technique, applying, for example, a batchgradient descent algorithm. Following the training of the INR network 200 and using theoptimal parameters ^ (obtained via the optimization process), the INR network can beapplied to predict a pixel value based on its corresponding coordinate values.

[0040] Using INR networks for the application of encoding and decoding images is furtherdescribed with respect to FIGS.3 and 4.

[0041] FIG. 3 is a block diagram of an example video encoder 300. In the example of FIG. 3,the video encoder 300 includes an INR-based encoder 320, a quantizer 330, and an entropy-based encoder 340. The INR-based encoder 320 receives an input data 310 to be coded. Theinput data 310 can be data associated with a frame of video (an image), a frame of a surfacerepresentation of an object, or a frame of volumetric data. To code the input data 310, theINR-based encoder 320 trains an INR network (e.g., the INR network 200 described herein).Specifically, based on the input data 310, the INR-based encoder 320 optimizes a costfunction associated with the function ^^, representative of the INR network, to determine theoptimal network parameters ^. For example, for an input image with dimensions ^ and ^,^ times ^ pairs of pixel coordinates (^, ^) and corresponding pixel values ^(^, ^) can be usedto train the INR network according to equation (2). The optimal parameters, generated by theINR-based encoder 320, are then quantized by the quantizer 330, and then the quantized parameters are entropy-coded, by the entropy-based encoder 340, into a bitstream 350. Alternatively, the optimal parameters, generated by the INR-based encoder 320, can be coded using neural compression codecs such as Neural Network Coding (NNC) / ISO / IEC 15938-17or MPEG-7 part 17. The bitstream 350 can be used by a decoder to reconstruct the inputimage 310, as described in reference to FIG.4.

[0042] FIG. 4 is a block diagram of an example video decoder 400. The decoder 400generally reverses the operation of the encoder 300 of FIG. 3. In the example of FIG. 4, the video decoder 400 includes an entropy-based decoder 420, a dequantizer 430, and an INR-based decoder 440. As illustrated, the decoder 400 receives the bitstream 410, 350 (generatedby the encoder 300) and entropy-decodes therefrom the quantized INR network parameters. The dequantizer 430 is then employed to dequantize these quantized INR network parameters,resulting in a restored version of the INR network parameters to be provided to the INR-based decoder 440.

[0043] The INR-based decoder 440 applies the trained INR network, defined by the restoredINR network parameters, to generate the reconstructed data 450. For example, to decode aninput image, the INR-based decoder 440 uses the INR network to predict the value of (or toevaluate ^^using the coordinates of) any pixel of the input image. Thus, the decoder 400 canbe applied to: 1) reconstruct the whole encoded image 310; 2) to reconstruct only a region ofthe encoded image; or 3) to progressively reconstruct the encoded image. For example, at theencoder 300, an INR network may be trained to predict pixel values of an image withdimensions ^ = 256 by ^ = 256 based on the corresponding coordinates. At the decoder400, pixel values of the image can be predicted by evaluating the trained INR network using:1) the full coordinate set used for training, including all pairs of ^ ∈ 0,1, … ,255 and ^ ∈0,1, … ,255; 2) a subset of the full set, including coordinates from a region of the encodedimage; or 3) a first subset of the full set, including coordinates that subsample the image(forming a low-resolution version of the encoded image) and then a second subset including the remaining coordinates. Any set of coordinates can be used to predict the correspondingpixel values, for example, in order to interpolate or to extrapolate the encoded image 310.

[0044] Hybrid INR networks have recently been applied to representing data, includingimages, videos, 3D objects, and volumetric data, among other applications. In a hybrid INRnetwork, the data coordinates are first mapped into latent variables (or a feature vector). Thelatent variables are then used as input for the neural network. An example for a hybrid INRnetwork is described in reference to FIG.5.

[0045] FIG. 5 is a diagram illustrating a hybrid INR network 500. In an aspect, this hybridINR network 500 can be applied by the INR-based encoder 320 of FIG.3 and the INR-baseddecoder 440 of FIG. 4. In the example of FIG. 5, during the encoding of input data (that is,the training of the network 500), the input data’s coordinates 510 are first mapped, by amapping unit 520, into respective latent variables 525. The mapping can be implemented by alookup table or a hash function, for example. The mapping may also involve anytransformation, such as a Fourier transformation, a coordinate transformation, anormalization transformation, or a combination thereof. The latent variables 525 can beupsampled, by an up-sampling unit 530, resulting in upsampled latent variables 535. Theupsampled latent variables are used as an input to an INR network 540 (e.g., such as INRnetwork 200), trained to produce the reconstructed data 550. In a hybrid-INR network 500 thelatent variables 525 are trained together with the parameters ^ of the INR network 540,resulting in optimal network parameters and optimal latent variables. Using such anarchitecture helps in handling the local attributes of the input data. Indeed, a group of latentvariables that correspond to a given part of the data may be uncorrelated with other groups oflatent variables that correspond to other parts of the data, and, thus, groups of latent variablescan be tailored to (or be characteristic of) corresponding parts of the data.

[0046] Following the training of the hybrid INR network 500 (e.g., by the encoder 320), thelearned latent variables 525 and network parameters are quantized (e.g., by the quantizer 330) and coded (e.g., by the entropy-based encoder 340) into the bitstream. Thus, during inference, first the latent variable 525 and the network parameters of the trained INR network 540 aredecoded (e.g., by the entropy-based decoder 420) from the bitstream and dequantized (e.g.,by the dequantizer 430). Then, the decoded and dequantized latent variables are up-sampled530. To reconstruct the data 550, the up-sampled latent variables are fed into the trained INRnetwork 540 using the decoded and dequantized network parameters of the INR network.

[0047] A hybrid INR network, named a coordinate-based low complexity hierarchical imagecodec (COOL-CHIC), was proposed by Ladune et al. (see, Ladune et al., “COOL-CHIC:Coordinate-based low complexity hierarchical image codec,” International Conference on Computer Vision (ICCV), 2023, hereinafter “Ladune”). In Ladune, the latent variables are arranged in hierarchical layers (or channels) ranging from a low-resolution representation(that provides for compact representation of smooth image regions) to a high-resolutionrepresentation (that captures the fine details of the image).

[0048] In hybrid INR networks, the latent variables, denoted by ^, are typically the largestcontributor to the bitrate (several orders of magnitude larger than that contributed by the INRnetwork parameters). One approach to reduce the transmission cost of these latent variables isto entropy code these variables based on their learned distributions, as described in Ladune.Principles of the hybrid INR network proposed in Ladune are described below in reference toFIGS.6-8.

[0049] FIG. 6 is a block diagram illustrating an example video encoder 600 applying a hybridINR network. The encoder 600 includes a probability prediction (PP) network 620, an up-sampling unit 640, an INR network 650, and entropy-based coders 625, 630, 655. Theencoder 600 is configured to process latent variables 610 (e.g., latent variables 525 generatedby the mapping unit 520 of FIG. 5). The encoder 600 up-samples, by the up-sampling unit 640, the latent variables. Based on these upsampled latent variables 645 the INR network 650is trained (overfitted) to produce reconstructed data 660 (e.g., a reconstructed image of avideo frame). The training results in optimal INR network parameters ^ that are coded, by theentropy-based coder 655, into the bitstream 670. In an inference mode, reconstructing thedata (e.g., by a decoder 700) involves feeding the trained INR network, defined by theoptimal network parameters ^, with the upsampled latent variable. And so, in addition to thenetwork parameters ^, the latent variables need to be coded into the bitstream 670. As furtherexplained below, due to their large bit representation, efficient coding of the latent variablescalls for the estimation of their distributions. To that end, the encoder 600 can be furtherconfigured to learn the distributions of respective latent variables using the PP network 620 –that is, a network trained to produce parameters of distributions of respective latent variables. Based on these learned distribution parameters the entropy-based coder 630 codes the latentvariables into the bitstream 670. The PP network is defined by PP network parameters,denoted by ψ, determined during the training of the PP network 620. The entropy-basedcoder 625 codes these PP network parameters into the bitstream 670. Note that in theexample of FIG. 6, while the entropy-based coder 630 that codes the latent variables relies ontheir learned respective distributions, the other entropy-based coders 625, 655 that code thePP network parameters and the INR network parameters rely on respective non-learneddistributions. In an aspect, for some of the latent variables respective non-learneddistributions can be used by the entropy-based coder 630. These non-learned distributionsmay be fixed distributions or may be distributions that was learned with respect to other latent variables (e.g., latent variables representing data from previous frames).

[0050] FIG. 7 is a block diagram illustrating an example video decoder applying a hybridINR network 700. The decoder 700 includes a PP network 720, an up-sampling unit 740, anINR network 750, and entropy-based decoders 715, 730, 755. The PP network 720 producesdistribution parameters of respective latent variables. The PP network 720 operates based onlearned PP network parameters ψ, determined during the training of the PP network 620. Theentropy-based decoder 715 decodes these PP network parameters from the bitstream 710.Based on the produced 720 distribution parameters, the entropy-based decoder 730 decodes the latent variables from the bitstream 710. Already decoded latent variables are provided back to PP network to serve as a context in producing the distribution parameters of thecurrently decoded latent variable (as further described in reference to FIG. 8). The decodedlatent variables are then upsampled by the up-sampling unit 740 (as performed by the up-sampling unit 640 at the encoder 600). Fed by the up-sampled latent variables, the INRnetwork 750 reconstructs the data 760 (e.g., a reconstructed image of a video frame) it istrained to synthesize based on the INR network parameters, decoded from the bitstream 710by the entropy-based decoder 755.

[0051] The operation of the hybrid INR network is further explained with respect to an image^ of a video frame, however, ^ may represent other types of data (such as a surface or avolume) that can be associated with a frame. Note that latent variables representative of data regions (e.g., pixels) from a data frame (e.g., a video frame) referred to herein also as corresponding to that data frame.

[0052] As illustrated, the INR network 650, 750 utilizes a hierarchical representation thatincludes multiple layers of different spatial resolutions 610. Each layer represents an image ^with a width ^ and a height ^ with a corresponding level of detail. Formally, the discretelatent variables, denoted by ^^ , can include ^ layers of latent variables: ^^ = {^^^ , ^ =0: (^ − 1)}. Each layer ^^^ is of width ^ / 2^and of height ^ / 2^. During the up-sampling 640, each layer ^^^may be upsampled by a factor of 2^(using any interpolation method) toobtain the upsampled layer version, denoted by ^^̂ . Together, the upsampled layers, ^̂ ={^^̂ , ^ = 0: (^ − 1)}, result in a dense 3D representation 645 of dimension ^ by ^ by ^.Thus, the INR network 650 can be trained based on ^ by ^ inputs of ^(̂i, j), where eachinput can include up to ^ latent variables, that is, ^(̂i, j) = {^(̂i, j, ^), ^ = 0: (^ − 1))}. Forexample, the trained INR network 750 can be used to predict a reconstructed pixel , ofthe original pixel ^(^, ^), by ^(^, ^) = ^^(^̂(i, j)). In an aspect, depending on the desired bitrate,not all the layers of the latent variables may be used to represent (code) an image ^.

[0053] When compressing an image ^ the goal is to do so while minimizing a cost function,as discussed with respect to equation (2). In the case of a hybrid INR network, the cost ofcoding an image can be expressed as: ^^^^ = ^(^, ^^(^̂)) + ^^(^^, ^, ψ ), (4)where ^ denotes an image to be coded with ^ hight, ^ width, and ^ color channels; where ^^denotes the quantized latent variables and ^̂ denotes their upsampled version; where ^^denotes the INR network 650, 750 and ^ denotes the INR network parameters; denotes the PP network 620, 720 and ψ denotes the PP network parameters; where ^ denotesa distortion metric measuring the distance between the image ^ and its reconstructed version^, as produced by the INR network ^^ from the upsampled latent variables ^̂, that is, ^ =^^(^̂); and where ^ denotes the rate (in bits per pixel) measuring the number of bits that arerequired to represent a pixel in a bitstream, that is, the number of bits that are required torepresent ^^, ^, and ψ. The distortion ^ and the rate ^ are balanced by a scalar value denotedby ^. In a case where the up-sampling unit 640 is implemented by a neural network, the parameters of that network are also learned and coded into the bitstream 670 to be used by the up-sampling unit 740 when used in an inference mode.

[0054] The objective, thus, is to find the latent variables ^^, the INR network parameters ^,and the PP network parameters ψ that minimize the coding cost, as follows:

[0055] Since the contribution of the INR network parameters ^ and the PP networkparameters ψ to the rate ^ is not as significant as that of the latent variables ^^, only the lattercan be considered when minimizing the coding cost, that is, ^(^^, ^, ψ ) ≈ ^(^^). Furthermore,^(^^) can be replaced by the cross entropy. Thus, equation (5) can be replaced by: where ^^(^^) is the joint distribution of the latent variables ^^. According to Equation (6),minimizing the cost involves minimizing the rate associated with the latent variables. Thiscan be achieved by reducing the amount of information contained in the latent variables, atthe price of a less accurate reconstruction, as less information in ^^ is likely to increase thedistortion ^. Alternatively, minimizing the cost can be achieved by obtaining estimates of thedistributions of the respective latent variables, as described herein.

[0056] Due to the high dimensionality of the latent variables, modeling of the jointdistribution of is not tractable. Instead, ^^(^^) can be factorized as follows: where ^^^^^^^^^^c^^^ ^ denotes a discrete conditional probability of a latent variable at position(i, j, k) conditioned on a corresponding spatial context c ^^^^ . Where (i, j) represents the spatialcoordinate of a latent variable in a layer ^ . The spatial context c^^^^may be provided byspatially neighboring latent variables that have already been decoded, and preferably selectedin a way that enables parallel decoding of the different layers of the latent variables (e.g., in awavefront-like approach).

[0057] In practice, the discrete distribution ^^^^^^^^^c ^^^^ ^ can be modeled by integrating thecontinuous distribution of the non-quantized latent variable, denoted by ^(^) and modeled asa Laplacian distribution, for example. Thus, the PP network 620 learns the expectationparameter, μ , and the scale parameter, base^ d on the context c^^^. Accordingly, probability of a latent variable ^^^^^can be expressed as: p^^^^^^^ ^c^^^^ = ∫^^^^^^^.^ ^^^^^^^.^g(y)dy , (8)where g ≅ ℒ^μ^^^, σ^^^^ denotes a Laplacian distribution. Thus, in the case of a Laplaciandistribution, for example, given a context c^^^, the PP network 620 can be trained to producethe corresponding distribution parameters, that is, ^μ^^^, σ^^^^ = f^^c^^^^, as described nextwith respect to FIG.8.

[0058] FIG. 8 is a diagram illustrating prediction of a latent variable distribution, using aspatial context 800. For simplicity of the presentation, only one layer of the latent variables810 is illustrated, however, processes applied to this layer can be similarly applied to theother layers. The example of FIG. 8 illustrates the process of predicting a distribution of acurrent latent variable (indicated by a black square) to be coded (e.g., by encoder 600 or 900)or to be decoded (e.g., by decoder 700 or 1000). However, when coding the latent variable,typically, neighboring latent variables from the current frame are available. While when decoding the latent variable, typically, some of the neighboring latent variables from the current frame are not yet available (decoded). This is indicated by the white squares (available latent variables) and the patterned squares (not yet available latent variables).

[0059] As illustrated, a spatial context 820 can be constructed based on the latent variables inthe spatial neighborhood of the current latent variable 830 – that is, for a current latentvariable at position (^, ^, ^), latent variables can be selected within a neighborhood locatedrelative to position (^, ^, ^) (e.g., latent variables grouped by the gray background in FIG. 8)to form the spatial context c^^^^ . The obtained spatial context, c^^^^, can then be used by the PPnetwork 840 to predict the distribution of the current latent variable. The distribution of thecurrent latent variable is predicted by estimating the parameters of that distribution (e.g.,^μ^^^ , σ^^^^). To that end, in the encoder, the PP network 840 is trained to produce, for eachlatent variable, the distribution parameters 850 based on the respective spatial context. Thistraining can be done by minimizing the coding cost expressed in equation (6). The training ofthe PP network 840 results in the PP network parameters ψ. In the decoder, the trained PPnetwork 840 is operated in an inference mode to produce, for each latent variable, thedistribution parameters 850 from the respective spatial context. The trained PP network 840operates based on the PP network parameters ψ determined by the encoder during trainingand provided to the decoder in the bitstream.

[0060] In the case of a video stream, for example, a hybrid INR network can be used torepresent each image of a video frame by a set of latent variables, ^^, that (together with thePP network parameters ψ and the INR network parameters θ) can be coded into a bitstream.To decrease the bitrate of the compressed video, the set of latent variables can be trained torepresent a group of video frames (see, e.g., Hyunjik et al., “C3: High-Performance and Low- Complexity Neural Compression from a Single Image or Video,” Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp.9347-9358). However, such an approach makes it difficult to randomly access individualframes at the decoder end. In addressing this shortcoming, Leguay et al. propose a predictivehybrid INR network applicable to individual frames (see, Leguay et al., Cool-chic video:Learned video coding with 800 parameters, Data Compression Conference (DCC) 2024, IEEE, pp 23–32, 2024, hereinafter “Leguay”).

[0061] Principles of the predictive hybrid INR network proposed in Leguay are describednext in reference to FIGS. 9 and 10. The operation of the predictive hybrid INR network isexplained herein with respect to an image ^ of a video frame, however, ^ may represent othertypes of data (such as a surface or a volume) that can be associated with a frame. Note that latent variables representative of data regions (e.g., pixels) from a data frame (e.g., a video frame) referred to herein also as corresponding to that data frame.

[0062] FIG. 9 is a block diagram illustrating an example video encoder 900 applying apredictive hybrid INR network. The encoder 900 receives, as an input, latent variables 910and outputs a bitstream 980. The encoder 900 includes a hybrid INR network 920, a motioncompensation unit 930, a multiplier 940, an adder 950, and a decoded frame buffer 960. The hybrid INR network 920 generally operates (in a training mode) as the hybrid INR networkdescribed in reference to FIG. 6. However, in this case the hybrid INR network is trained toproduce two optical flows, ^^^^^ and ^^^^^, a weight mask ^, a prediction mask ∝, and aresidual image ^ . The optical flows represent the pixel-wise motion between respectivereference frames (already reconstructed images 970, stored in the decoded frame buffer 960)and the current frame. And so, the motion compensation unit 930 can use these optical flowsto generate a prediction, denoted by ^, of the currently coded image ^, as follows: where ^^^^ is an operator that warps (i.e., spatially maps) an image into another imagebased on motion vectors given by an optical flow. Specifically, a first reference image, ^^^^^,is warped into a first prediction, ^^ = ^^^^(^^^^^, ^^^^^) and a second reference image,^^^^^ is warped into a second prediction, ^^ = ^^^^(^^^^^ , ^^^^^ ). Using elementwisemultiplication, denoted by “∙”, the two predictions are then blended by the mask yieldingthe prediction image ^. As illustrated in FIG. 9, the prediction image ^ (output of the motioncompensation unit 930) is then corrected by the residual image ^, as follows: ^= ^+∝∙ ^, (10)where, the prediction mask ∝ is a binary mask that can be used to mask out 940 a predictionpixel if it does not reliably predict the corresponding pixel in ^. Adding 950 the maskedprediction image, ∝∙ ^, to the residual image ^ produces the reconstructed image 970, ^.

[0063] Hence, the two optical flows, ^^^^^ and ^^^^^, the weight mask the prediction mask∝, and the residual image ^ can be learned by minimizing the coding cost expressed inequation (4), where ^^(^)̂ = ^ = ^+∝∙ ^ . Note that the predictive hybrid INR network,illustrated in FIG. 9, can be applied using only one reference image or any number of reference images available in the decoded frame buffer 960. In such a case equation (9) can be expressed as: Where, ^ is the number of reference images used; where ∑^ ^^^ ^^is equal to an all-onesmatrix. where ^^^^^ is the optical flow with respect to ^^^^^, that is, the qthreference image. In a variant, the weight mask ^, the prediction mask ∝, or both can be removed from the predictive hybrid INR network.

[0064] FIG. 10 is a block diagram illustrating an example video decoder 1000 applying apredictive hybrid INR network. The decoder 1000 receives, as an input, a bitstream 1010 andoutputs reconstructed images 1070. The decoder 1000 includes a hybrid INR network 1020, amotion compensation unit 1030, a multiplier 1040, an adder 1050, and a decoded frame buffer 1060. The hybrid INR network 1020 generally operates (in an inference mode) as the hybrid INR network described in reference to FIG. 7. However, in this case the hybrid INRnetwork produces two optical flows, ^^^^^ and ^^^^^, a weight mask ^, a prediction mask ∝,and a residual image ^. The optical flows represent the pixel-wise motion between respectivereference frames (already reconstructed images 1070, stored in the decoded frame buffer1060) and the current frame. And so, the motion compensation unit 1030 can use these optical flows to generate a prediction, denoted by ^, of the currently decoded image ^, asexpressed by equation (9). As illustrated in FIG. 10, the prediction image ^ (output of themotion compensation unit 1030) is then corrected by the residual image ^, as shown byequation (10). Adding 1050 the masked prediction image, ∝∙ ^, to the residual image ^produces the reconstructed image 1070, ^.

[0065] Using spatial context (as described in reference to FIG. 8), the entropy coding of thelatent variables is temporal context agnostic. This is because the estimated distributions of thelatent variables are conditioned only on latent variables representing data from the currentframe. Not being conditioned on a temporal context (that is, latent variables representing datafrom other frames) might lead to suboptimal performance in terms of a rate-distortiontradeoff (e.g., see equation (4)). However, although utilizing temporal context often leads to abetter rate-distortion tradeoff, this may not always be the case.

[0066] According to aspects, temporal information can be used to estimate the distribution ofa latent variable. Thus, the distribution of a latent variable can be conditioned on a contextthat is constructed based on a spatial context, a temporal context, or both (namely, aspatiotemporal context). Typically, the spatial context is constructed based on alreadydecoded latent variables representative of data from the current frame (as described inreference to FIG.8), and the temporal context is constructed based on already decoded latentvariables representative of data from reference frame(s). Using a temporal context mayleverage temporal redundancy across the latent variables, resulting in a more accurateprediction of respective distributions, and, thereby, a more efficient entropy coding of these latent variables. The use of a spatial context, a temporal context, or both may be indicated by a flag, coded into the bitstream.

[0067] Generally, a spatiotemporal context, associated with a latent variable representingdata from a current frame, is constructed as follows. To form a temporal context, latentvariables, corresponding to one or more reference frames and located within a neighborhoodrelative to a position collocated with the latent variable, are selected. To form a spatialcontext, latent variables, corresponding to the current frame and located within aneighborhood relative to a position of the latent variable, are selected. A spatiotemporalcontext may be constructed by merging the spatial context and the temporal context. Utilizingsuch a spatiotemporal context in the estimation of a distribution of a latent variable exploitsboth spatial redundancy and temporal redundancy across the latent variables. Principles ofconstructing a temporal context from latent variables representing data from one reference frame will be described herein; extending these principles to constructing a temporal context from latent variables representing data from more than one reference frame should be straightforward. The construction of a spatiotemporal context is further described in reference to FIG.11.

[0068] FIG. 11 is a block diagram illustrating prediction of a latent variable distribution,using a spatiotemporal context 1100. For simplicity of the presentation, only one layer (e.g.,^^^^^ ) of the latent variables is illustrated 1110, 1120. However, processes applied to thislayer can be similarly applied to the other layers ({^^^ , ^ = 1: (^ − 1)}). The example of FIG.11 illustrates the process of predicting a distribution of a current latent variable (indicated bya black square) to be coded (e.g., by encoder 600 or 900) or to be decoded (e.g., by decoder700 or 1000). When coding the latent variable, typically, neighboring latent variables(corresponding to the current frame and reference frame(s)) are available. When decoding thelatent variable, typically, some of the neighboring latent variables (corresponding to thecurrent frame) are not yet available (decoded). This is indicated by the white squares (available latent variables) and the patterned squares (not yet available latent variables).

[0069] As illustrated, a spatial context 1115 can be constructed based on latent variableslocated within a spatial context region 1112 associated with the current latent variable – thatis, for a current latent variable at position (^, ^, ^), latent variables can be selected within aneighborhood located relative to position (^, ^, ^) (e.g., latent variables grouped by the graybackground of region 1112) to form the spatial context c^^^^. Additionally, a temporal context1125 can be constructed based on latent variables located within a temporal context region1122 associated with the current latent variable. That is, for a current latent variable atposition (^, ^, ^), latent variables can be selected within a temporal neighborhood locatedrelative to position (^, ^, ^) (e.g., latent variables grouped by the gray background of region1122) to form the temporal context c^^^^. Note that the manner in which latent variables areselected to form the spatial context 1115 and the temporal context 1125 by the encoder can becommunicated to the decoder by a syntax element coded into the bitstream, may be inferred by the decoder based on other coding parameters, or may be fixed (known to the decoder).

[0070] Next, the spatial context 1115 and the temporal context 1125 can be combined by amerger 1130, forming a spatiotemporal context 1135, c^^^^^ . Combining the spatial context andthe temporal context may be by any function of the underlying latent variables (i.e., variableswithin the spatial context region 1112 and the temporal context region 1122), such as anaverage, a median, or a sum. For example, the underlying latent variables can be combined based on their relative (spatial and / or temporal) distance from the current latent variable (e.g.,weighted average). In an aspect, only a spatial context 1115 or a temporal context 1125 maybe used. In such a case, there is no need to employ the merger 1130 and either context may bedirectly fed to the PP network 1140.

[0071] The obtained spatiotemporal context 1135, c ^^^^^ , can be used by the PP network 1140to predict the distribution of the current latent variable. The distribution of the current latent variable is predicted by estimating the parameters of that distribution. To that end, in theencoder 600, 900, the PP network 1140 can be trained, for each latent variable, to produce thedistribution parameters 1150 based on its respective spatiotemporal context. This can be doneby minimizing the cost expressed in equation (6), using the following term for the joint distribution The training of the PP network 1140 results in the PP network parameters ψ. In the decoder700, 1000, the trained PP network 1140 is operated in an inference mode, using the PPnetwork parameters ψ determined by the encoder during training and provided to the decoderin the bitstream. In an aspect, to predict the distribution of a current latent variable, thedecoder may reuse the PP network parameters ψ provided (in the bitstream or from anothersource) for another latent variable (corresponding to the current frame or a previously codedframe).

[0072] The distribution can follow any distribution function, for example, a conditionalGaussian distribution, a mixture of Gaussian distributions, or a Naïve Bayes model. Theparameters of these distributions can be estimated by the PP network 1140 using an autoregressive module or using a linear regression of a decision tree, for example.

[0073] According to aspects, the use of a certain context can be indicated by a syntax elementcoded into the bitstream. This allows the encoder to choose, based on a bitrate distortiontradeoff for example, whether to use 1) a spatiotemporal context, 2) a spatial context, 3) atemporal context, or 4) neither. For example, the encoder can choose not to use any context topredict the distributions of the latent variables, but instead to use a predetermined (default)distributions for the respective latent variables. Thus, a region of the current frame can becoded using four coding paths: 1) using a spatiotemporal context to predict the distributionsof the latent variables, 2) using a spatial context to predict the distributions of the latentvariables; 3) using a temporal context to predict the distributions of the latent variables; and 4)using predetermined distributions for the latent variables. Next the encoder can select thecoding path that resulted in a lower coding cost (as expressed in equation (4)), the one thatresulted in a lower distortion ^, or the one that resulted in a lower rate ^, for example.

[0074] For example, a syntax element can be coded into the bitstream using one bit toindicate whether a spatiotemporal context is used (e.g., “1”) or not (“0”). Alternatively, the syntax element can be coded into the bitstream using two bits to indicate whether a spatiotemporal context is used (e.g., “11”), a spatial context is used (e.g., “01”), a temporal context is used (e.g., “10”), or neither (e.g., “00”). The syntax element may be determinedwith respect to a region of one frame, a whole frame, a group of frames, or a whole videostream. In an aspect, the choice whether to predict distributions of respective latent variables based on a spatiotemporal context, a spatial context, a temporal context, or neither need notbe coded into the bitstream, but instead can be inferred based on other coding parameters,analysis of already reconstructed content of the frame(s) (e.g., using a machine learningalgorithm), or any other heuristic.

[0075] The spatial context region 1112 and the temporal context region 1122, if having astatic shape, may result in a suboptimal coding performance in terms of the rate-distortioncost (e.g., as expressed by equation (4)). This may be because, depending on the camera’smotion or motion of content at the neighborhood of the current latent variable, the most relevant latent variables to the probability distribution of the current latent variable (i.e., thosemost correlated with the current latent variable) may not be considered in constructing thecontext. For example, to construct the spatial context 1115, more relevant latent variablesmay be residing outside the illustrated spatial context region 1112; and to construct thetemporal context 1125, more relevant latent variables may be residing outside the illustratedtemporal context region 1122. Hence, according to further aspects, an encoder (e.g., 600 or900) can be configured to select the shape of the spatial context region and / or the shape of thetemporal context region. The selected shape(s) can be coded into the bitstream, if used(instead of a default shape for example).

[0076] Hence, some data region of a coded frame may have longer-range dependencies. Insuch a case, to accurately model the probability distribution of a latent variable from thatregion (thereby, achieving a shorter bitstream for the latent variable through entropy coding),neighboring latent variables may be needed that are drawn from a region that is larger.Likewise, to accurately model the probability distribution of a latent variable, neighboringlatent variables may be needed that are drawn from a region that extends in a certain direction.As an example, if a camera is moving towards the left of the scene, latent variablesrepresenting data region from a past reference frame may result in a more informativetemporal context if positioned more to the right of the current latent variable than to the left.

[0077] Therefore, it is beneficial to determine the shape of a context region to improveentropy coding performance of the respective latent variable. According to aspects, during theencoding of data frames (e.g. by encoder 600 or 900) the shape of context regions – that is,the shape of a spatial context region and / or the shape of a temporal context region – can bedetermined and signaled in the bitstream (e.g., received by a decoder 700 or 1000). To thatend, a set of predetermined shapes may be used to select from and the selected shape may besignaled in the bitstream. The selected shape may be with respect to each latent variable, agroup of latent variables, latent variables corresponding to a current frame, or latent variablescorresponding to multiple frames.

[0078] In a first aspect, an encoder may select one of several predetermined shapes for acontext region. Typically, the shape may not be the same for the spatial context region andthe temporal context region. Since the context is used to condition the distribution of acurrent latent variable, the latent variables this context is constructed from need to beavailable, that is, already decoded (as illustrated in FIGS. 8 and 11). Advantageously, theshape of the context region (i.e., the latent variables within that region) may be selected toallow parallel decoding of the layers of the latent variables, for example in a wavefront-like approach. For example, a pair of shapes may be selected, one shape for the spatial contextregion and one shape for the temporal context region, as illustrated by FIGS. 12 and 13.

[0079] FIG. 12 is a diagram illustrating configurations of context region shapes 1200. In afirst configuration 1210, 1220, the temporal context region (gray region in 1220) consists of25 latent variables organized in a square which center is collocated with the position of thecurrent latent variable being coded (denoted by the black square in 1210) and the spatialcontext region (gray region in 1210) consists of 5 latent variables located to the left and to thetop of the position of the current latent variable. In a second configuration 1230, 1240, thetemporal context region (gray region in 1240) consists of 9 latent variables organized in asquare which center is collocated with the position of the current latent variable being coded(denoted by the black square in 1230) and the spatial context region (gray region in 1230)consists of 21 latent variables located to the left and to the top of the position of the currentlatent variable. Furthermore, a third configuration 1210, 1240, may include temporal contextregion and spatial context region both with smaller number of latent variables as well as afourth configuration 1230, 1220, may include temporal context region and spatial contextregion both with larger number of latent variables.

[0080] FIG. 13 is a diagram illustrating other configurations of context region shapes 1300.In the example of FIG. 13, four possible shapes for the temporal context region (gray regionsin 1320, 1340, 1360, 1380) are illustrated that can be selected by an encoder; and fourpossible shapes for the spatial context region (gray regions in 1310, 1330, 1350, 1370) areillustrated that can be further selected by an encoder. The different shapes of a temporalcontext region allow to exploit relative movements of content from the current and referenceframe(s) (e.g., detectable by motion analysis), so that latent variables (representing contentfrom the reference frame(s)) will be better correlated with a currently coded latent variable (representing the content from the current frame), thereby, in a better position to provide temporal context. Likewise, the different shapes of a spatial context region allow to exploitsimilar patterns within the content of the current frame (e.g., detectable by pattern analysis),so that latent variables representing content with similar pattern to the content represented by the current latent variable will be better correlated with that current latent variable, thereby, ina better position to provide spatial context. And so, such better temporal context, betterspatial context, or when combined 1130, better spatiotemporal context, can result in a moreaccurate estimate 1140 of the distribution parameters 1150 of the current latent variable,thereby, a more efficient entropy coding of that latent variable.

[0081] Hence, an encoder 600, 900, may select a shape for the spatial context region and / orfor the temporal context region (such as those shapes illustrated in FIGS. 12 and 13). In a firstaspect, this selection may involve evaluating the coding performance of different pairs ofshapes for the spatial context region and the temporal context region, and the pair of shapesthat yields the least coding cost may be selected. Thus, in evaluating the coding performancefor each pair of shapes, the hybrid INR network (employed by encoder 600) or the predictive hybrid INR network (employed by encoder 900) may be trained as described herein to minimize the coding cost (see, e.g., equation (4)). Alternatively, only the PP network can be used to decide which pair of shapes should be selected, using precomputed latent variables and a coding cost including only ^(^^).

[0082] In a second aspect, the shape(s) may be directly selected (e.g., using expert designedalgorithm or a machine learning model) as follows. One or more shapes for the temporal context region may be directly selected based on motion analysis of content in the temporal neighborhood of the current latent variable to be coded. Likewise, one or more shapes for thespatial context region may be selected based on pattern analysis of content in the spatialneighborhood of the current latent variable to be coded. In the case where more than one shape is directly selected for the spatial context region and / or the temporal context region, determining the final pair of shapes can be done using the first aspect described above.

[0083] According to aspects, the encoder 600, 900 may signal in the bitstream whether ashape for a spatial context and / or a temporal context is selected, or a default shape is being used. This signal may be with respect to latent variable(s) representing all frames, one frame,or a region of a frame. In a case where selected shape(s) is (are) used (with respect to latentvariable(s) representing all frames, one frame, or a region of a frame), the encoder may signala syntax element indicating the selected shape(s). The selected shape for the spatial contextregion and / or for the temporal context region may be signaled in the bitstream using a syntaxelement that identifies the shapes used. Different shapes may be selected and used for thespatial context region and the temporal context region, used to construct the spatiotemporal context of each latent variable. In this case, a mask indicating the used shapes for each latentvariable can be signaled in the bitstream. In some variants, only a spatial context is used andthus only its respective region shape may be signaled, while in other variants a temporalcontext is used (with or without spatial context) and thus its respective region shape may besignaled.

[0084] In other variants, the encoder may be able to select any arbitrary shape rather than oneshape from a predefined set. In that case, the encoder can signal this shape in the bitstream (rather than a syntax element identifying one shape from a predefined set of shapes). Signaling the shape may be performed by coding the locations of the latent variables (enclosed by the shape) relative to the current latent variable being coded.

[0085] When a predictive hybrid INR network is employed, one or more reference framesmay be used (e.g., as described above in reference to FIGS. 9 and 10). In this case, a differentshape of the temporal context region may be selected and used with respect to the latentvariables that correspond to each reference frame. And so, a syntax element may be signaledin the bitstream identifying the shape used with respect to each reference frame. In an aspect,when one shape (e.g., used with respect to one reference frame) can be used to derive othershapes (e.g., used with respect to other reference frames) only that one shape needs to besignaled in the bitstream. For example, when a first reference frame precedes the currentframe and a second reference frame follows the current frame, the shape used with respect tothe preceding reference frame may be spatially related to the shape used with respect to thefollowing reference frame. For example, the spatial relation may be that of symmetry, such aspoint symmetry (over the position of the latent variable currently coded or another point) orvertical / horizontal symmetry. The symmetry may be implicitly known to both the encoder900 and the decoder 1000 or may be signaled in the bitstream.

[0086] FIG. 14 is a flowchart of an example method for encoding data 1400, according towhich aspects of the present embodiments can be implemented. The method 1400 in theexample of FIG. 14 can be applied to encode, into a bitstream, a data region of a currentlyencoded frame. The data region may include one or more data elements, such as pixels orvoxels; and the frame may be of a video image, a surface representation of an object, orvolumetric data. According to aspects described herein, the encoding of the data region mayinclude steps 1410 to 1440. In step 1410, an INR network can be trained to produce the dataregion from latent variables, the training determines the parameters of the INR network. Asdescribed herein, the latent variables can be learned or latent variables that were used toproduce another data region of the currently encoded frame or of a previously encodedframe can be used. In step 1420, the distributions of the latent variables can be determined,where a distribution of a latent variable can be determined using a respective context. Suchcontext can be constructed based on latent variables located within a shaped context region(e.g., the shaped context regions illustrated in FIGS. 12 and 13). These distributions can be modeled, for example, according to a conditional Gaussian distribution, a mixture of Gaussian distributions, or a Naïve Bayes model. These distributions can also be learned. To that end, a neural network (e.g., a PP network) can be trained to produce the distributionsfrom the respective contexts, where each distribution of a latent variable is defined by one ormore distribution parameters, and where the training determines parameters of the neuralnetwork. The method 1400 can then be applied to coding into the bitstream: the latentvariables based on the determined distributions, in step 1430; and the parameters of the INRnetwork, in step 1440. In the case where the distributions of the latent variables are learnedby a neural network, the parameters of the neural network are also coded into the bitstream.

[0087] FIG. 15 is a flowchart of an example method for decoding data 1500, according towhich aspects of the present embodiments can be implemented. The method 1500 in theexample of FIG. 15 can be applied to decode, from a bitstream, a data region of a currentlydecoded frame. The data region may include one or more data elements, such as pixels or voxels; and the frame may be of a video image, a surface representation of an object, or volumetric data. According to aspects described herein, the decoding of the data region mayinclude steps 1510 to 1540. In step 1510, the parameters of an INR network can be decodedfrom the bitstream. This INR network is trained to produce the data region from latentvariables. In step 1520, the distributions of the latent variables can be determined, where adistribution of a latent variable can be determined using a respective context. Such contextcan be constructed based on latent variables located within a shaped context region (e.g., theshaped context regions illustrated in FIGS. 12 and 13). In an aspect, the distributions can beproduced by a neural network (e.g., a PP network), in which case the method 1500 can befurther applied to decode from the bitstream the parameters of the neural network, usedby the network to produce the distributions from their respective spatiotemporal contexts.Next, in step 1530, the method 1500 is applied to decoding from the bitstream the latentvariables based on the determined distributions. Then, the data region can be produced bythe INR network (using the decoded INR network parameters) from the latent variables, instep 1540.

[88] In an aspect, the encoding process 1400 further includes coding into thebitstream a flag that indicates whether the context is a spatial context, a temporal context, or a spatial and temporal context (i.e., a spatiotemporal context). Accordingly,during the decoding 1500, the flag may be decoded from the bitstream and may be usedto inform the decoder whether the context used in determining a distribution of a latentvariable is a spatial context, a temporal context, or a spatiotemporal context. In anotheraspect, the decoded flag is applied to data regions from multiple frames including the currently encoded frame.

[0089] As described herein, a context associated with a current latent variable (e.g., blacksquare in FIG. 11) may comprise a temporal context 1125, a spatial context 1115, or both (aspatiotemporal context). The spatial context can be constructed based on a first set of latentvariables located within a spatial context region 1112, spatially neighboring the latentvariable. The temporal context 1125 can be constructed based on a second set of latentvariables located within a temporal context region 1122, temporally neighboring the latentvariable, where the latent variables from the second set are used by the INR network toproduce a data region from a reference frame previously encoded. In an aspect, aspatiotemporal context can be formed by combining the first set of latent variables and thesecond set of latent variables using a linear or nonlinear function. For example, thespatiotemporal context can be formed by combining the first set of latent variables and thesecond set of latent variables based on their distance to the current latent variable. Suchdistance may be in the coordinate space or the value space of the latent variables.Alternatively, or in addition, the spatiotemporal context can be formed by combining the firstset of latent variables and the second set of latent variables based on the distance between thecurrently encoded frame and the reference frame.

[0090] As described herein with reference to FIGS. 12 and 13, the encoding process 1400may further include coding into the bitstream a flag indicating the shape of a spatial context region and / or the shape of a temporal context region. During the decoding process 1500, such a flag can inform the decoder what latent variables should be used toconstruct a context for determining the distribution of the current latent variable.Accordingly, the encoding process 1400 may include selecting a shape for a spatial contextregion, and coding, into the bitstream, a flag identifying that shape. Selecting the shapemay be from a set of predetermined shapes. Alternatively, latent variables (forming theshape), spatially neighboring the current latent variable, can be selected based on pattern analysis of the currently encoded frame. Furthermore, the encoding process 1400 mayinclude selecting a shape for a temporal context region, and coding, into the bitstream, aflag identifying that shape. Selecting the shape may be from a set of predetermined shapes. Alternatively, latent variables (forming the shape), temporally neighboring the current latent variable, can be selected based on motion analysis of the currently encoded frame and the reference frame.

[0091] The illustrations of the aspects described herein are intended to provide a generalunderstanding of the structure, function, and operation of the various aspects. The illustrations are not intended to serve as a complete description of all of the elements and features of apparatuses and systems that utilize the structures or methods described herein. Many other aspects may be apparent to those of skill in the art upon reviewing the disclosure. Other aspects may be utilized and derived from the disclosure, such that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, the disclosure and the figures are to be regarded as illustrative rather than restrictive.

[0092] The description of the aspects is provided to enable the making or use of the aspects.Various modifications to these aspects will be readily apparent, and the generic principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.

Claims

1. CLAIMS 1. A method, comprising:encoding, into a bitstream, a data region of a currently encoded frame, theencoding comprises: training an implicit neural representation (INR) network to produce the dataregion from latent variables, the training determines parameters of the INR network,determining distributions of the latent variables, wherein a distribution of alatent variable is determined using a respective context, constructed based on latentvariables located within a shaped context region, coding, into the bitstream, the latent variables based on the determined distributions, and coding, into the bitstream, the parameters of the INR network.

2. A method, comprising:decoding, from a bitstream, a data region of a currently decoded frame, the decoding comprises: decoding, from the bitstream, parameters of an INR network, the INRnetwork is trained to produce the data region from latent variables,determining distributions of the latent variables, wherein a distribution of alatent variable is determined using a respective context, constructed based on latent variables located within a shaped context region, decoding, from the bitstream, the latent variables based on the determined distributions, and producing, by the INR network, the data region from the latent variables,using the parameters of the INR network.

3. An apparatus, comprising:at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to encode, into a bitstream, a data region of a currently encoded frame, the encoding comprises: training an INR network to produce the data region from latent variables, thetraining determines parameters of the INR network, determining distributions of the latent variables, wherein a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region, coding, into the bitstream, the latent variables based on the determined distributions, and coding, into the bitstream, the parameters of the INR network.

4. An apparatus, comprising:at least one processor; and memory storing instructions that, when executed by the at least one processor, cause the apparatus to decode, from a bitstream, a data region of a currently decoded frame, the decoding comprises: decoding, from the bitstream, parameters of an INR network, the INR network is trained to produce the data region from latent variables,determining distributions of the latent variables, wherein a distribution of a latent variable is determined using a respective context, constructed based on latent variables located within a shaped context region, decoding, from the bitstream, the latent variables based on the determined distributions, and producing, by the INR network, the data region from the latent variables, using the parameters of the INR network.

5. The method according to claim 1 further comprising, or the apparatusaccording to claim 3 wherein the instructions further cause the apparatus to perform:coding, into the bitstream, a flag indicating whether the context is a spatialcontext, a temporal context, or a spatiotemporal context.

6. The method according to claim 1 or 5, or the apparatus according toclaim 3 or 5, wherein the context comprises a spatial context, and wherein:the spatial context is constructed based on a first set of latent variables locatedwithin a spatial context region, spatially neighboring the latent variable.

7. The method according to claim 6 further comprising, or the apparatusaccording to claim 6 wherein the instructions further cause the apparatus to perform:selecting a shape for the spatial context region; and coding, into the bitstream, a flag identifying the shape.

8. The method according to claim 7, or the apparatus according to claim 7,wherein the selecting of the shape for the spatial context region comprises: selecting the shape from a set of predetermined shapes, or selecting latent variables, spatially neighboring the latent variable, based on pattern analysis of the currently encoded frame.

9. The method according to any one of claims 1 or 5-8, or the apparatusaccording to any one of claims 3 or 5-8, wherein the context further comprises a temporal context, and wherein: the temporal context is constructed based on a second set of latent variables located within a temporal context region, temporally neighboring the latent variable, wherein latent variables from the second set are used by the INR network to produce a data region from a reference frame previously encoded.

10. The method according to claim 9 further comprising, or the apparatusaccording to claim 9 wherein the instructions further cause the apparatus to perform:selecting a shape for the temporal context region; and coding, into the bitstream, a flag identifying the shape.

11. The method according to claim 10, or the apparatus according to claim 10,wherein the selecting of the shape for the temporal context region comprises: selecting the shape from a set of predetermined shapes, or selecting latent variables, temporally neighboring the latent variable, based on motion analysis of the currently encoded frame and the reference frame.

12. The method according to claim 9 further comprising, or the apparatusaccording to claim 9 wherein the instructions further cause the apparatus to perform:combining the first set of latent variables and the second set of latent variables basedon their distance to the latent variable.

13. The method according to claim 9 or 12 further comprising, or theapparatus according to claim 9 or 12 wherein the instructions further cause theapparatus to perform: combining the first set of latent variables and the second set of latent variables basedon the distance between the currently encoded frame and the reference frame.

14. The method according to any one of claims 1 or 5-13, or the apparatusaccording to any one of claims 3 or 5-13, wherein the determining of the distributionscomprises: training a neural network to produce the distributions of the latent variables from respective contexts, wherein the training determines parameters of the neural network; and coding, into the bitstream, the parameters of the neural network.

15. The method according to claim 2 further comprising, or the apparatusaccording to claim 4 wherein the instructions further cause the apparatus to perform: decoding, from the bitstream, a flag indicating whether the context is a spatial context, a temporal context, or a spatiotemporal context.

16. The method according to claim 2 or 15 further comprising, or theapparatus according to claim 4 or 15 wherein the instructions further cause theapparatus to perform: decoding, from the bitstream, a flag identifying a shape of the shaped context region.

17. The method according to any one of claims 1-2 or 5-16, or the apparatusaccording to any one of claims 3-16, wherein the distributions are modeled according to a conditional Gaussian distribution, a mixture of Gaussian distributions, or a Naïve Bayes model.

18. The method according to any one of claims 1-2 or 5-17, or the apparatusaccording to any one of claims 3-17, wherein the training of the INR network furthercomprises learning the latent variables.

19. The method according to any one of claims 1-2 or 5-17, or the apparatusaccording to any one of claims 3-17, wherein the latent variables are latent variables used toproduce another data region of the currently encoded frame or of a previously encodedframe.

20. The method according to any one of claims 1-2 or 5-19, or the apparatusaccording to any one of claims 3-19, wherein the frame is of a video image, a surfacerepresentation of an object, or volumetric data.

Citation Information

Patent Citations

  • Hybrid neural network based end-to-end image and video coding method

    US20230096567A1