Hybrid non-event and event video communications

Hybrid non-event and event video communication systems enhance video quality by combining event and non-event data through dual-layer encoding and neural fields, addressing the limitations of traditional cameras in temporal resolution and dynamic range.

WO2026055360A1PCT designated stage Publication Date: 2026-03-12DOLBY LABORATORIES LICENSING CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing video cameras struggle to capture light intensity changes with high temporal resolution and dynamic range, limiting the quality and frame rate of generated videos.

Method used

Hybrid non-event and event video communication systems that combine event camera data with non-event image data using dual-layer encoding architectures, neural fields, and artificial neural networks to enhance video frame interpolation and reconstruction.

Benefits of technology

The system achieves high-quality video streams with improved temporal resolution and dynamic range, enabling frame rates higher than traditional cameras by integrating event and non-event data effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000009_0001
    Figure IMGF000009_0001
  • Figure IMGF000009_0002
    Figure IMGF000009_0002
  • Figure IMGF000010_0001
    Figure IMGF000010_0001
Patent Text Reader

Abstract

Non-event camera data is derived from first and second non-event images generated from a scene for first and second image frame time points, respectively, by a non-event camera. Event camera data is derived from a sequence of raw events generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor. The sequence of event time points is between the first and second time points. Event related image data is encoded into an overall image container. The event related image data is derived from a combination of both the event camera data and the non-event camera data. A recipient device of the overall image container is caused by the event related image data of the overall image container to render a display image on an image display. The display image is derived from the event related image data.
Need to check novelty before this filing date? Find Prior Art

Description

D24045W001HYBRID NON-EVENT AND EVENT VIDEO COMMUNICATIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 691,206, filed on 5 September 2024, and from European Patent Application No. 24 210 655.7, filed on 4 November 2024, which are both incorporated by reference herein in their entirety.TECHNOLOGY

[0002] The present disclosure relates generally to visual computing and more particularly to hybrid non-event and event video communications.BACKGROUND OF THE INVENTION

[0003] An event camera with an imaging sensor that responds to local luminance changes can be used to generate event camera data from a physical environment or visual scene by an array or spatial distribution of pixels or sensor elements of the image sensor inside the event camera without using a shutter. Each pixel or sensor element in the image sensor can operate in parallel independently and asynchronously with other pixels or sensor elements in the same image sensor.

[0004] The event camera data as generated from the physical environment or visual scene may include a stream of asynchronously occurring pixel-specific events at a temporal resolution much higher than an image refresh rate at which a typical non-event camera used to acquire RGB or YCbCr images / videos operates. In addition, the event camera data can sensitively respond to a wide range of light intensities present in the physical environment or visual scene in a luminance or brightness range much higher than most if not all non-event cameras.

[0005] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.BRIEF DESCRIPTION OF DRAWINGS

[0006] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0007] FIG. 1 A illustrates an example dual-layer encoding architecture; FIG. IB illustratesD24045W001 an example fixed time interval encoder / decoder architecture;

[0008] FIG. 2 illustrates an example backwards compatible dual layer codec architecture;

[0009] FIG. 3A and FIG. 3B illustrate example dual layer encoder / decoder architectures or frameworks;

[0010] FIG. 4A and FIG. 4B illustrate example process flows; and

[0011] FIG. 5 illustrates an example hardware platform on which a computer or a computing device as described herein may be implemented.DETAILED DESCRIPTION OF THE INVENTION

[0012] Example embodiments, which relate to hybrid non-event and event video communications, are described herein. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present disclosure.

[0013] Example embodiments are described herein according to the following outline:1. GENERAL OVERVIEW2. EVENT DATA3. EVENT REPRESENTATIONS4. EVENT FORMATS5. EVENT-GUIDED OPERATIONS6. EVENT AND NON-EVENT DATA MUXING7. DUAL-LAYER ENCODING ARCHITECTURE8. EVENT BITSTREAM9. EVENT POINT CLOUD CODING10. RESIDUAL BITSTREAM11. IMAGE RECONSTRUCTION QUALITY12. EVENT METADATA13. ARCHITECTURAL EXTENSIONS AND ALTERNATIVES14. EXAMPLE ALTERNATIVE ONE15. EXAMPLE ALTERNATIVE TWO16. OTHER EXAMPLE ALTERNATIVES17. EXAMPLE PROCESS FLOWS18. IMPLEMENTATION MECHANISMS - HARDWARE OVERVIEWD24045W00119. EQUIVALENTS, EXTENSIONS, ALTERNATIVES ANDMISCELLANEOUS1. GENERAL OVERVIEW

[0014] This overview presents a basic description of some aspects of an example embodiment of the present invention. It should be noted that this overview is not an extensive or exhaustive summary of aspects of the example embodiment. Moreover, it should be noted that this overview is not intended to be understood as identifying any particularly significant aspects or elements of the example embodiment, nor as delineating any scope of the example embodiment in particular, nor the disclosure in general. This overview merely presents some concepts that relate to the example embodiment in a condensed and simplified format, and should be understood as merely a conceptual prelude to a more detailed description of example embodiments that follows below.

[0015] Event cameras are bio-inspired sensors that asynchronously capture per-pixel intensity changes and output a stream or sequence of events each of which encodes the time, location, and polarity of a brightness change.

[0016] Since an event stream can be used to capture light intensity changes or information in a relatively wide dynamic range at a relatively high temporal resolution such as at the order or granularity of milliseconds or microseconds, event data in the event stream can be used with non- event image data such as RGB data captured in a relatively narrow dynamic range at a relatively low temporal resolution with RGB image sensors of non-event cameras (e.g., based on shutter, lens, and image sensor, CCD based, etc.) to generate a relatively high quality interpolated video stream in a dynamic range (e.g., much, twice, three times or more, etc.) higher than that supported by the non-event cameras at a temporal resolution (e.g., much, twice, three times or more, etc.) higher than that supported by the non-event cameras.

[0017] As used herein, a non-event camera refers to any of a wide range of cameras that do not process light or luminance data into asynchronous events but rather sample and / or integrates light or luminance sensory responses over a specific time interval (e.g., corresponding to an image refresh rate of 30 frames per seconds or fps, 60 fps, 90 fps, 120 fps, tens of milliseconds or longer, at a relatively coarse temporal resolution as compared with light change event detection by an event camera, etc.) into an image that includes pixel values in some or all color channels in a representative color space for pixels represented in the image. An image generated by such a camera may be referred to as a non-event image. Pixel values or coefficients derived therefrom in a non-event image may be referred to as non-event image data.

[0018] Examples of a color space in which an image (e.g., a non-event image, a residualD24045W001 image, a constructed or reconstructed image, an interpolated image, a neural field reconstructed image, etc.) as described herein may be represented may include, but are not necessarily limited to only, one of: an RGB color space, an YCbCr color space, a linear color space, a non-linear color space, a non-perceptually quantized space, a perceptually quantized (or PQ) color space, an input color space, an intermediate color space, an output color space, etc.

[0019] Event data or events captured with event cameras may be represented in any event representation format among a variety of lossless and / or lossy event representation formats (including but not limited to those commercially available an event camera maker or provider such as Prophesee or other event camera providers / vendors).

[0020] In some operational scenarios, a framework may be implemented for representing, encoding, transmitting, receiving, decoding and / or reconstructing a hybrid or mixture of both non-event image / video data (e.g., generated with a regular shutter-based RGB camera, etc.). The framework may be used as a video data delivery pipeline or architecture to generate a bitstream or image file (container) that carries a specific (e.g., video coding, etc.) representation of the hybrid or mixture of the non-event image / video data and the event image / video data.

[0021] As one of many specific applications, the hybrid or mixture of the non-event image / video data and the event image / video data can be delivered from an upstream encoding device to a downstream recipient or decoding device to enable one or both of the upstream encoding device and the downstream recipient / decoding device to perform video frame interpolation tasks. Additionally, optionally or alternatively, the framework, pipeline or architecture can be used for any other event-guided tasks in addition to or other than the video frame interpolation tasks.

[0022] The framework, pipeline and / or architecture can be used to provide or support an event-guided workflow for transmitting, receiving, processing, encoding and / or decoding a hybrid or mixture of event and non-event image / video data in connection with still images or videos. In some operational scenarios, the event image / video data can be represented as fixed length events encoded in NAL units of a (e.g., H.264, H.265, HEVC, etc.) coded bitstream that is encoded with both the event and non-event image / video data. In some operational scenarios, the event image / video data can be encoded in one or more tracks of an (e.g., ISO-BMFF, an image or media file, etc.) image / video container.

[0023] The framework as described herein may a backwards compatible dual layer framework in which the non-event image / video data (e.g., RGB images at a relatively coarse temporal resolution such as 30 or 60 fps, etc.) may be carried in a first layer or base layer of a video signal or bitstream, whereas event-derived or event-dependent image / video data may beD24045W001 carried in a second layer or non-base layer of the video signal or bitstream.

[0024] Example event-derived or event-dependent image / video data as described herein may include, but are not necessarily limited to only, event-derived or event-dependent residual image / video data used for video frame interpolation tasks and / or other event-guided tasks.

[0025] By way of illustration, an interpolated non-event (e.g., RGB, YCbCr, etc.) image sequence covering a sequence of time points / instants may be generated from two temporally successive (e.g., RGB, YCbCr, etc.) images in the non-event image / video data. A corresponding event image sequence covering the same sequence of time points / instants may be interpolated, constructed or generated from events in the event image / video data. Residuals may be generated as differences between the interpolated non-event image sequence and the corresponding event image sequence. The residuals (or simply residue) may be modeled with artificial neural networks using operational parameters such as weights, biases, non-linear continuous activation functions, etc. The artificial neural networks may be implemented with a combination of one or more blocks of multilayer perceptrons or MLPs, convolutional neural networks or CNNs, feedforward neural network layers, etc.

[0026] In some operational scenarios, the artificial neural networks used to model the residues or residue may be implemented as neural fields each with a combination of MLP and CNN blocks, etc. As activation functions used in the neural fields or the MLP (and / or CNN) blocks are non-linear continuous functions, the neural fields are implicit neural representations which are continuous functions. Hence the neural fields can be used to incorporate the event data with the non-event data and generate or predict images at an arbitrary or finer frame (refresh) rate as compared with a frame (refresh) rate represented in the images in the non-event image / video data.

[0027] Additionally, optionally or alternatively, event metadata - including but not limited to some or all event camera configuration data of event cameras used to capture or generate the event image / video image data - may be specified, generated, and / or included in a video signal (e.g., a HEVC or H.265 or H.264 bitstream) or image file (e.g., ISO-BMFF file, etc.) as described herein.

[0028] In some operational scenarios, a framework for processing, transmitting, receiving, encoding and / or decoding a hybrid or combination of non-event image / video data and event image / video data may be implemented with a non-backwards compatible architecture using neural fields. The non-backwards compatible architecture may be implemented as a base layer alone or a base layer in addition to a residual layer.

[0029] As used herein, a neural field - such as a neural radiance field (NeRF) modified orD24045W001 enhanced based on event data - can be used as a coordinate based neural network that parameterizes event data / measurements across space (e.g., event locations, etc.) and time (e.g., event time points / instants at a relatively high event temporal resolution, etc.). After modeling or fitting event data with corresponding neural field operational parameters such as weights or biases in MLP or CNN layers / blocks, the neural field can be used as an implicit image function to represent an image as set of latent code to predict pixel (e.g., RGB, YCbCr, etc.) values of a given pixel at a given coordinate or spatial location.

[0030] Example neural fields and / or related operations are described in U.S. Provisional Patent Application No. 63 / 611,975, titled “CODING EVENT CAMERA DATA USING NEURAL FIELD,” by Anustup Kumar Atanu Choudhury, Guan-Ming Su, filed on December 19, 2023; Yiheng Xie et al., “Neural Fields in Visual Computing and Beyond.'' Eurographics / CGF State-of-the-Art Report (2022); Chen et. al., “NeRV: Neural Representation for Videos,” NeurlPS (2021), all of which are incorporated herein by reference in entirety.

[0031] Example embodiments described herein relate to encoding non-event and event camera data. Non-event camera data is derived from a first non-event image and a second non- event image. The first non-event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera. Event camera data is derived from a sequence of raw events. The sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor. The sequence of event time points is between the first time point and the second time point. Event related image data is encoded into an overall image container. The event related image data is derived from a combination of both the event camera data and the non-event camera data. A recipient device of the overall image container is caused by the event related image data of the overall image container to render at least one display image on an image display. The at least one display image is derived from the event related image data.

[0032] Example embodiments described herein relate to decoding non-event and event camera data. Event related image data is decoded from an overall image container. The event related image data has been derived and encoded into the overall image container by an upstream device from a combination of both event camera data and non-event camera data. The non-event camera data has been derived by the upstream device from a first non-event image and a second non-event image. The first non-event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera. The event camera data is derived by the upstream deviceD24045W001 from a sequence of raw events. The sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor. The sequence of event time points is between the first time point and the second time point. At least one display image is generated from the event related image data. The at least one display image is rendered on an image display.

[0033] In some example embodiments, mechanisms as described herein form a part of a media processing system, including but not limited to any of: an image processing system, a video or image encoder, a video or image decoder, a media content provider system, a media content streaming system, a mobile device, a wearable device, a computer server, a media processing system, a cloud based computing system, laptop computer, netbook computer, tablet computer, desktop computer, computer workstation, or various other kinds of computing devices and media processing units.

[0034] Various modifications to the embodiments and the generic principles and features described herein will be readily apparent to those skilled in the art. Thus, the disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.2. EVENT DATA

[0035] Event cameras are asynchronous (physical) sensors that have a different representation of a scene as compared with other types of image sensors such as CCD. The event cameras can generate event camera data with a relatively high sensitivity to low light as well as light changes in a physical environment or visual scene, a relatively high temporal resolution, a relatively low latency (in the order of microseconds), a relatively (e.g., very, etc.) high dynamic range (e.g., a range of 140dB versus a range of 60dB of a camera with a different type of image sensor, etc.), and a relatively low power consumption.

[0036] A non-event camera may acquire full images at an image acquisition or refresh rate specified by or in reference to an external clock (e.g., 30fps, etc.). In contrast, an event camera or an event sensor therein such as a dynamic vision sensor (DVS) respond to luminance or brightness changes in a physical environment or visual scene instantly, asynchronously and independently for every pixel or sensor element in the event camera or sensor. A per-pixel reference log intensity (e.g., of received photons or light, etc.) may be maintained for each pixel or sensor element in the event camera or sensor. A pixel or sensor element fires or outputs a binary output / event if the magnitude of the log intensity changes beyond a maximum (log) intensity change threshold for positive or negative change. Mathematically, this event firing process can be expressed as follows:D24045W001wherelog (I(xk, tk- Atk) (2) where I(xfe, tfe) represents the intensity (e.g., of photons or light, etc.) at the pixel location xk=(%fe, yk), and time tk. Atkrepresents the time elapsed since the last event at the same pixel location or sensor element. I(xk, tk— Atk) represents the intensity (e.g., of photons or light, etc.) at the pixel location xk= (x^y^. Also, eprepresents the maximum (log) intensity change threshold for the positive value change and enrepresent the maximum (log) intensity change threshold for the magnitude of the negative value change. Here, e(xk, tk) outputs or represents the polarity of the event data or (log) intensity value change at the given pixel location xkand the given time point tk, which may be of one of the values 1, -1 and 0. If e(xfe, tfe) = 0, then no events are generated. Regardless of whether an event is generated for a pixel position or sensor element at a given time point, the polarity of the pixel position or sensor element at the given time point may be denoted as pkfor simplicity. Events with pk= 0 for pixel positions or sensor elements may not be stored or explicitly represented in the event camera data or a stream thereof and may be implicitly assumed in the absence of positive or negative events for these pixel positions or sensor elements. Here k represents an event index, for example, used to identify or order events belonging to the same sequence of events. Different indexes k may refer to multiple events occurred at the same time or at different times.

[0037] An event ek(t) - if fired - at a pixel location or sensor element may be defined as a 4- parameter tuple efe(t) = (xk, yk, tk,pk) containing coordinate information about its pixel location xk,yk). the exact trigger time tk, and the 1 -bit polarity pk. As the event ek(t) is only stored if fired, only two states of pk, namely positive and negative events, need to be stored for which 1 bit is sufficient. As used throughout the specification, a reference to pkis not limited to a particular representation or coding of pk. Unless explicitly stated, pk, may either refer to a 3- states-representation or 2-states-representation. These representations may be associated with individual codings of these representations as how to map positive events, negative events, and no events (if part of the representation) to the respective value of pk. In preferred embodiments, the 1 -bit polarity pkis set to 1 for positive events and set to 0 for negative events. In some variants of event cameras / sensors, instead of merely denoting whether a pixel at which an event is fired has increased / decreased in brightness or luminance, the event also contains information about or specifies how much intensity has changed. Other pixel locations are inferred to have theD24045W001 same intensity values as before and are thus not represented in the event camera data or the stream thereof.

[0038] Note that the event camera data or stream may be composed of sparse points or pixels as it may only explicitly represent - or contain information - these sparse points or pixels at each of which a sufficiently large intensity change as compared with a maximum intensity change threshold is detected or sensed by a corresponding sensor element.3. EVENT REPRESENTATIONS

[0039] Event (camera) data generated with an event camera can be represented in a wide variety of event (data) representations.

[0040] In a first example, the event camera data may be represented or defined in a relatively straight-forward way by a sequence of events (denoted as e , tj ) between an initial time (point) ti and a final time (point)= (e (t) 11 E [tq, tj]). A time point t within the interval between and tj may also be referred to as a time instance. Similarly, the sequence of events may also be referred to as an event stream. The polarity and the timestamp information for events in the event camera data is retained in this (e.g., lossless, etc.) representation.

[0041] In a second example, the event camera data may be represented or defined in an event frame representation, which may be alternatively referred to as an event image representation or a 2D histogram of events. In this representation, all the events within a spatiotemporal neighborhood are accumulated or aggregated together. This representation quantizes event timestamps from the original temporal resolution of the event camera into a coarser temporal resolution corresponding to a unit of time (denoted as A = [tj, tq]) covered with the spatiotemporal neighborhood.

[0042] By way of illustration, an event camera may comprise or include a plurality of event sensor elements covering an overall spatial size or area with a spatial resolution. In some operational scenarios, the overall spatial size or area may be comparable to what covered by a non-event camera in an overall non-event and event image / video processing or acquisition system.

[0043] A spatiotemporal neighborhood for event accumulation or aggregation may be formed as the (Cartesian) product - W xH x A - of a spatial neighborhood denoted as W x H (e.g., covered by W x H event sensor elements, etc.) and the aforementioned unit of time A.

[0044] Within this spatiotemporal neighborhood W x H x A , various event frame representation may be implemented. For instance, for each pixel location (x, y) corresponding to an individual event sensor element within A, only the latest polarity of the events is considered. Or, for each pixel location (x, y), the polarity of all the events occurring at (x, y) within A isD24045W001 summed up, and the (e.g., final, etc.) polarity may be set as or based on the sign of the sum. Or, for each pixel location (x, y), the polarity of all the events occurring at (x, y) within A is summed up and normalized between [0, 1] such that a value of 0.5 indicates no events occurred at that location and a value > 0.5 indicates more positive events occurred and vice versa.

[0045] In a third example, the event camera data may be represented or defined in a time surface (or surface of active events) representation. In this representation, each pixel stores a single time value and the intensity of the pixel is a function of motion history at that location represented as follows:where T is a tunable parameter that depends on the motion of the scene and t is the reference time.

[0046] In a fourth example, the event camera data may be represented or defined in an event voxel grid representation. In this representation, a 3D histogram of the individual events (corresponding to both space and time) is created, generated or derived based on the event camera data or an event sequence corresponding thereto. This representation can support or maintain a relatively high temporal resolution for relatively high quality temporal information of the events. Each event’s polarity can spread among its closest voxels using a kernel represented as follows:E(x, y, t) =lipikb(x - xi)kb(y - yt)kbt - tf) (4-2) kb(a) = max(0, 1 — |a|) (4-3) where B represents the number of bins to discretize the time dimension, whereas the event camera data or the event sequence has N events within the time interval of [tt, tN].

[0047] Example event aggregations with Time Surface and Event Voxel are described in Lagorce et. al., “HOTS: A hierarchy of event based time surfaces for pattern recognition, ” IEEE TP AMI (2017); and Rebecq et. al., “Events-to-video: Bringing modern computer vision to event cameras f CVPR (2019), all of which are incorporated herein by reference in entirety.

[0048] In a fifth example, the event camera data may be represented or defined in an event point cloud representation. In this representation, individual events within a spatiotemporal neighborhood of VI / x H x A can be treated as individual points in 3D space formed by two spatial coordinates and one temporal coordinate. In some operational scenarios, depending on the polarity of the events, two point clouds Poand may be created - one for each of the two polarity values (or positive and negative polarity values) represented as follows:D24045W001where Porepresents the point cloud corresponding to the positive polarity, whereas Prrepresents the point cloud corresponding to the negative polarity.

[0049] In a sixth example, the event camera data may be represented or defined in an event neural field representation. In this representation, a neural field (which typically acts as a regressor) may be modified or enhanced to act as a classifier and to model raw event data and / or an event frame representation of the raw event data. The neural field is represented or implemented using a multi-layer perceptron in which an activation function such as the sigmoid function may be replaced at the last layer with a non-sigmoid function such as the Softmax function using a cross-entropy loss function instead of an MSE loss function.

[0050] Additionally, optionally or alternatively, various intermediate event representations such as DAT or uncompressed binary representation may be used or implemented for representing or buffering intermediate event data generated in event processing operations in one or more event data processing systems or devices as described herein.4. EVENT FORMATS

[0051] Event (camera) data generated with an event camera can be transferred or delivered between upstream and downstream event (and / or non-event image / video) data processing systems or devices in a wide variety of lossless or lossy event formats.

[0052] Example lossless event formats may include, but are not necessarily limited to only, any of: Prophesee RAW formats such as (32 bits, 4-bit Event type, 28-bit payload) EVT-2, 64-bit EVT-2.1, compact 16-bit EVT-3; IniVation raw event data formats such as AEDAT versions 1.0, 2.0, 3.0, 3.1 and (64-bit data, compression optional) 4.0; HDF5 / H5 (Hierarchical Data Format version 5) with general-purpose library and file format for storing data with variations among different implementors / providers / vendors (e.g., Prophesee’ s implementation of HDF5, etc.) and / or shared event databases; comma-separated- value or CSV formats with asynchronous raw events stored in a csv file indicating or identifying corresponding their timestamps, spatial coordinates and polarity values; TXT format with asynchronous raw events stored in a text file indicating or identifying corresponding their timestamps, spatial coordinates and polarity values; etc.

[0053] For example, in the compact 16-bit EVT-3 event format, each 16-bit word representing an event may include the four (4) MSBs to define the word or event type. Example event types are illustrated in TABLE 1 below.D24045W001TABLE 1

[0054] Some event processing operations as described herein can be implemented to perform lossless compression or operating with event camera data in a lossless event representation / format. Additionally, optionally or alternatively, some event processing operations as described herein can be implemented to perform lossy compression or operating with event camera data in a lossy event representation / format.

[0055] Example lossy event formats may be used for lossy event representations of event camera data as discussed herein.

[0056] In some operational scenarios, event camera data can be represented in a lossy event format associated with the point cloud representation. The event camera data in this point cloud representation may be coded using MPEG Point Cloud Coding (G-PCC). Specifically, V-PCC (Video-based PCC) and / or G-PCC (Geometry-based PCC) may be used to compress independently both point clouds Poand P Example point cloud coding operations are described in MPEG Systems, “Text of ISO / IEC DIS 23090-18 Carriage of Geometry based Point Cloud Compression Data”, ISO / IEC JTC1 / SC29 / WG03 Doc. N0075 (Nov. 2020), the contents of which are incorporated herein by reference in entirety.

[0057] In some operational scenarios, event camera data can be represented in a lossy event format associated with the event frame representation. The event camera data in this event frame representation may be coded or compressed using video coding standards such as HEVC and vvc.

[0058] In some operational scenarios, event camera data can be represented in a lossy event format associated with a neural network or neural field representation. The event camera data in this neural network / field representation - or the event field - may be coded or compressed using MPEG Neural Network Compression (NNC) including but not limited to compression efficient quantization and deep context adaptive binary arithmetic coding (DeepCABAC). Example NNCD24045W001 operations are described in Heiner Kirchhoffer, et al, “Overview of the Neural Network Compression and Representation (NNR) Standard,” IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, VOL. 32, NO. 5, MAY 2022, pp.3203~3216, the contents of which are incorporated herein by reference in entirety. Additionally, optionally or alternatively, additional processing operations of the neural network / field operational parameters such as weights like sparsification, pruning and so on may be implemented. Some or all these operations may help compress the neural field (operational parameters) relatively efficiently but may also result in lossy compression or lossy representation of the event camera data.5. EVENT-GUIDED OPERATIONS

[0059] Event data derived with an event camera can be used to support a wide variety of event-guided operations including but not limited to event-guided high frame rate (e.g., greater than 30 or 60 fps, RGB, YCbCr, etc.) image reconstructions.

[0060] For example, frame rate interpolation may be performed to construct or reconstruct intermediate (e.g., RGB, etc.) image frames between two successive (e.g., non-interpolated, RGB, available, etc.) image frames. Additional information from the event data can be used to enhance or improve image qualities of the interpolated image frames constructed or reconstructed from the two successive image frames.

[0061] For example, interpolated RGB images may be generated from two available RGB images in a sequence of (non-event) RGB images (generated or derived from non-event image / video data acquired by a non-event camera) based at least in part on event data (generated or derived from events acquired by a event camera) or a portion thereof representing events between the successive RGB images. Whereas the sequence of (non-event) RGB images acquired with the non-event camera supports a relatively temporal resolution such as a relatively low image refresh rate such as 30 or 60 fps, the interpolated RGB images generated in part from the event data, along with the sequence of (non-event) RGB images, may be used to support a relatively high temporal resolution such as a relatively high image refresh rate of greater than 60 fps.

[0062] Since raw events in original event camera data are sparse and asynchronous, the raw events in its original representation (e.g., Prophesee EV2.1, etc.) used by an event camera become relatively difficult to use directly. Hence, these raw events may be transformed or converted into events in different event camera data representations other than the original representation.

[0063] In some operational scenarios, the event data used for performing event-guided operations - such as (e.g., RGB, etc.) image construction or reconstruction tasks, event-drivenD24045W001 frame rate interpolation tasks, etc. - may be represented with a voxel grid representation of events. The voxel grid representation may be more advantageous over many other event camera data representations such as an event frame representation in that the voxel grid representation can contain or preserve more temporal information represented in the original event camera data that gives rise to the event data.

[0064] Event-assisted (or event-guided) image construction / reconstruction / interpolation may be carried out with a TimeLens approach, which is described in S. Tulyakov et al., “TimeLens: Event based video frame interpolation,” CVPR (2021), the contents of which are incorporated herein by reference in entirety. Under this approach, assuming that two successive or consecutive non-event (e.g., RGB, intensity values, etc.) images corresponding to two respective time points / instants - referred to as a first time point / instant and a second time point / instant, respectively - are available or given along with event data representing events from the first time point / instant indexing the first frame to the second time point instant indexing the second frame, the event-assisted video frame interpolation (VFI) may be performed to generate interpolated images for intermediate time points or instants between the first and second time points.

[0065] Bidirectional events representing both past and future events relative to a given intermediate time point / instant may be used in the VFI operation to generate an intermediate or interpolated image for the given intermediate time point / instant. The past and future events may also be referred to as left and right events, respectively. Notably, reconstructing intermediate frame using event data is non-trivial, since event data encodes image differences (or intensity change events) in a binarized and spatially sparse way. The TimeLens approach may be implemented to use a data-driven, deep-leaming-based method to solve this problem. For example, the TimeLens approach may act as a multi-branch approach under which a neural network-based image synthesis module and an optical flow-based warping module can be used to refine aggregated results used to generate the intermediate or interpolated image.

[0066] Event-assisted (or event-guided) image construction / reconstruction / interpolation may be carried out with other approaches other than the TimeLens approach. These other approaches may include, but are not necessarily limited to only, SuperFast, EVDI, and REFID, CBMNet, etc. Example SuperFast operations are described in Y. Gao et al., “SuperFast: 200x video frame interpolation via event camera”, TP AMI (2023), the contents of which are incorporated herein by reference in entirety. Example EVDI operations are described in X. Zhang and L. Yu, “Unifying motion deblurring and frame rate interpolation with events,” CVPR (2022), the contents of which are incorporated herein by reference in entirety. Example REFID operations are described in L. Sun et al., “Event based frame interpolation with ad-hoc deblurring,” CVPR (2023), the contentsD24045W001 of which are incorporated herein by reference in entirety. Example CBMNet operations are described in T. Kim et al., “Event-based Video Frame Interpolation with Cross-Modal Asymmetric Bidirectional Motion Fields,” CVPR (2023), the contents of which are incorporated herein by reference in entirety.

[0067] The SuperFast approach divides the frame rate interpolation task into two modules or pathways - a fast synthesis pathway and slow synthesis pathway. The slow synthesis pathway uses the same module as in the TimeLens approach operating with events or event data represented in the event voxel grid representation. In comparison, the fast synthesis module uses the raw events (in the original event camera data generated by the event camera or event sensor elements therein) as input in addition to the two consecutive or successive non-event (e.g., RGB, etc.) image frames.

[0068] The raw event stream in the original event camera data can be passed through an event feature extraction module implemented based at least in part on a Spiking neural network, followed by an event encoder (neural network) to extract (e.g., relatively high level as compared with the features extracted from the raw events, etc.) features related to the events. The fast synthesis module may include or contain an image encoder (neural network), a feature fusion module and a decoder (neural network). A refine module may be used to process or refine the output from the fast Synthesis module to generate, enhance or output synthesized intermediate image frames.

[0069] Finally, the output (or the synthesized intermediate image frames) of the refine module and the fast synthesis module / pathway and the output (or the interpolated images from the TimeLens approach) of the slow synthesis module / pathway can pass through an (intermediate image) fusion module in which learned or optimized weights for the fusion module or neural networks therein can be applied to generate or construct the (final output) interpolated images or image frames.

[0070] Additionally, optionally or alternatively, the EVDI or REFID approach can use an event voxel grid representation for joint deblurring and frame rate interpolation.6. EVENT AND NON-EVENT DATA MUXING

[0071] Event data in an event data representation as described herein can be multiplexed with non-event image / video data into the same bitstream or image file that is transferred or delivered between upstream and downstream event (and / or non-event image / video) data processing systems or devices in a wide variety of lossless or lossy event formats.

[0072] In some operational scenarios, the event data can be encoded into or carried in one or more NAL (Network Abstraction Layer) units as defined or specified in one or more videoD24045W001 coding specification such as the H.264 and H.265 video coding standards. These NAL units can be transmitted with data packets each of which contain an integer number of bytes, for example as data payloads. These NAL units may include or contain specific event information about an event stream representing the event data.

[0073] In some operational scenarios, the event data can be encoded into or carried in one or more tracks of an ISO-BMFF (ISO Base Media File Format) image file. The ISO-BMFF image file represents an image / video container file in a format or structure for carrying multimedia content. The ISO-BMFF image file may include or contain multiple tracks, among which Track 1 may be used to carry or contain a base layer (e.g., RGB, YCbCr, etc.) video stream carrying non- event image / video data and other tracks such as one or more subsequent tracks following Track 1 may be used to carry or contain event data as described herein.7. DUAL-LAYER ENCODING ARCHITECTURE

[0074] FIG. 1A illustrates an example dual-layer encoding architecture. This encoding architecture may be implemented by an upstream device in an non-event and event image / video data pipeline. The upstream device encodes non-event image / video data in such as non-event (RGB) images and corresponding event (camera) data into an overall bitstream or image / video container file. The overall bitstream or image / video container file may be implemented with a multi-layer structure that includes at least one base layer and one non-base layer such as an event layer.

[0075] The encoding architecture as illustrated in FIG. 1 A includes or comprises first processing blocks - such as RGB sensor, RGB image signal processor or ISP, image processor, video enhancing, video encoder, etc. - to encode the non-event (RGB) image data into a video stream. This architecture also includes or comprises second processing blocks - event sensor, event signal processor or ESP, spatial alignment, event preprocessor, event encoder, etc. - to encode the event data into event metadata and an event bitstream.

[0076] The video bitstream generated by the first processing blocks of the encoding architecture of FIG. 1 A, and the event metadata and event bitstream generated by the second processing blocks of the encoding architecture of FIG. 1 A may be inputted to a multiplexer or muxer to generate an overall a hybrid or combination of non-event and event image / video data.

[0077] More specifically, in connection with the above-mentioned first processing blocks as illustrated in FIG. 1A, original or raw non-event (RGB) image / video data may be acquired or generated with the RGB sensor of a non-event camera, which may operate in conjunction with the upstream encoding device as a part of the upstream encoding device or as a separate device. The original or raw non-event (RGB) image / video data generated by the RGB sensor may beD24045W001 received and processed by the non-event (RGB) ISP into RGB camera images. The RGB camera images may be received by the image processor of the upstream encoding device. Regions of interest (ROIs) may be identified by the image processor based at least in part on processing image content or pixel values of the RGB camera images or intermediate RGB images generated therefrom using object detection techniques implemented by the upstream encoding device or the image processor. The upstream device or the video enhancing processing block (e.g., reshaping, display management, etc.) therein may enhance the intermediate or processed RGB images generated by the image processor into enhanced RGB images (e.g., reshaped images, etc.). The enhanced RGB images can be received, processed, compressed or encoded by the video encoder into the video bitstream to be received by the MUX for multiplexing into the overall bitstream or image / video container file.

[0078] Furthermore, in connection with the above-mentioned second processing blocks as illustrated in FIG. 1A, original or raw (light intensity change) events may be acquired or generated with the event sensor (or a plurality of spatially distributed event sensor elements) of an event camera, which may operate in conjunction with the upstream encoding device as a part of the upstream encoding device or as a separate device. The original or raw events generated by the event sensor may be received and processed by the (event) ESP into the original event (camera) data. The upstream encoding device or the event camera may perform spatial alignment operation on the original event (camera) data to spatially align a first X-Y space represented by pixels in the non-event RGB images with a second X-Y space represented by events in the original event (camera) data. For example, the first X-Y space and the second X-Y space may be spatially aligned (e.g., spatial transformation, spatial rotation, spatial translation, spatial scaling, etc.) into a common X-Y space such as the first X-Y space represented by the pixels in the non- event RGB images. The spatially aligned original event (camera) event may be received and processed by the event preprocessor into events in a specific event data representation such as the event voxel grid representation. In some operational scenarios, the ROIs may be generated by the image processor of the first processing blocks and used by the event preprocessor to preserve or generate relatively high quality event data in the ROIs and relatively low quality event data - which may even be excluded from being encoded into the event bitstream or the overall bitstream - outside the ROIs. The events in the specific event data representation as generated or outputted by the event preprocessor may be received, processed, compressed or encoded by the event encoder into the event metadata and event bitstream.

[0079] The multiplexed hybrid or combination of the non-event and event image / video data as generated by the muxer may be used to generate the overall bitstream or image / videoD24045W001 container file. The overall bitstream or image / video container file may be formatted in a specific non-event and event hybrid image / video data format in accordance with one or more non-event and event image / video data coding formats as described herein.

[0080] The overall bitstream or image / video container file in the specific non-event and event hybrid image / video data format may be stored in non-transitory tangible computer readable media or transmitted / delivered to downstream devices such as a downstream recipient device in the present example.

[0081] The downstream recipient device (not shown) receives the overall bitstream or image / video container file directly or indirectly from the upstream (encoding) device, decodes the non-event (RGB) images the base layer of the overall bitstream or image / video container file, and decodes the corresponding event data from the event layer of the overall bitstream or image / video container file. The decoded event data - which may be the same as or approximates the event data at the upstream device - may be used by the downstream recipient device to perform event-assisted or event guided operations on or with the decoded non-event (RGB) images, which may be the same as or approximates the non-event (RGB) images at the upstream device.

[0082] As described herein, each of the upstream encoding device and the downstream (recipient or decoding) device may be implemented with one or more computing processors and other software and hardware implemented specific logics or computing system resources.

[0083] The video stream in the base layer of the overall bitstream or image / video container file may be encoded with a sequence of non-event (RGB) image frames at a relatively low frame rate such as 30fps or 60fps. In some operational scenarios, the relatively low frame rate may be set or configured depending on specific image / video related applications to be supported by the architecture of FIG. 1 A. This (RGB) video stream or the base layer of the overall bitstream or image / video container file may be encoded using available video codecs (e.g., already deployed in the field, to be deployed in the field, etc.) such as H.264, H.265 and so on.

[0084] The non-event (RGB) images encoded in the base layer or video stream of the overall bitstream or image / video container file may be derived based on original non-event camera data captured or generated by an non-event camera from a scene.

[0085] In comparison, the event data such as the event metadata and event bitstream encoded in the event layer of the overall bitstream or image / video container file may be derived based on the raw or original event camera data captured or generated by the event camera from the same scene.

[0086] For the purpose of illustration only, it has been described that a bitstream comprisingD24045W001 two layers may be used to carry non-event and event image / video data. It should be noted that, in other operational scenarios, instead of or in addition to a bitstream, an image / video container file may be used to carry non-event and event image / video data as described herein. For example, in some operational scenarios, one or more main or primary file components or tracks in the image container may be used to carry encoded non-event image / video data, whereas one or more attendant or secondary file components or tracks in the image container may be used to carry encoded event data.8. EVENT BITSTREAM

[0087] The event layer or the event stream therein may comprise, include or contain NAL units used to store the encoded event data in the hybrid or combination of the non-event and event image / video data in the overall bitstream or image / video container file. These NAL units used to carry the event data may be particularly identified by a data field value therein as one or more specific types in an applicable video coding standard. By way of example but not limitation, the NAL units used to store the encoded event data as described herein may be specifically designated or identified (e.g., in a header or data type field, etc.) as the “Unspecified” type in the H.264 or H.265 video coding specifications.

[0088] Events represented in the encoded event data in the event layer may be represented in a specific event data representation - such as the voxel grid representation - that is supported by the downstream recipient device (or its video synthesis logic) of the overall bitstream or image / video container file. For example, in response to determining that downstream devices such as the downstream recipient device in the present example are to support the voxel grid representation for events, the upstream encoding device may encode or pack the events included in the event data into the voxel grid representation to represent the events. Additionally, optionally or alternatively, the upstream encoding device may encode or pack some or all raw events generated by the event camera into the overall bitstream or image / video container file.

[0089] In some operational scenarios, each event in the event data of the event bitstream may be of a fixed data length such as a fixed 32-bit or 16-bit data length per event - e.g., based on a Prophesee event data format. In some operational scenarios, some or all events occurred in a given or fixed time interval (e.g., between two successive non-event RGB images acquired with the non-event camera, etc.) may be packed into a single NAL unit.

[0090] In an example, as shown in TABLE 1, timing information specifying a time point / instant at which an event occurs or is detected by an event sensor element of the event camera may be represented by a timestamp at a relatively high temporal resolution such as 24 bits.D24045W001

[0091] The timing information of the event or the 24-bit timestamp may be represented or denoted using two parts: EVT_TIME_LOW and EVT_TIME_HIGH, each of which is 12 bits. More specifically, the 12 LSBs of the timestamp (e.g., in micro- seconds, etc.) are encoded in an EVT_TIME_LOW event data portion, whereas the 12 MSBs of the timestamp are encoded in EVT_TIME_HIGH event data portions.

[0092] To construct the (full event) timestamp, a recipient device can concatenate the EVT_TIME_LOW event data portion with the last EVT_TIME_HIGH event data portion received or decoded before the EVT_TIME_LOW event data portion.

[0093] The overall concatenated time information or timestamp may be denoted as EVT_TIME, and may be used to find or identify events that lie within the same time interval demarcated by the two successive or consecutive non-event (RGB) images to be encoded in NAL units of the base layer. All these identified events within the same time interval are gathered and packed into an (event) NAL unit (e.g., of the “Unspecified” type in H.264 and H2.65, etc.) along with other event data portions illustrated in Table 1 above.

[0094] FIG. IB illustrates an example fixed time interval encoder / decoder architecture that may be implemented with an upstream encoding device and a downstream recipient / decoding device as described herein. As illustrated, the upstream encoder encodes non-event image / video data into a sequence of consecutive non-event (e.g., RGB, etc.) image frames - denoted as . . . VFn-i, VFn, VFn+i, etc. - at a plurality of consecutive or sequential time points or instants (logically indexed by frame indexes . .., n-1, n, n+1, . .., etc.). Each non-event (RGB) image or image frame - e.g., VFn, etc. - in the sequence of consecutive non-event (RGB) images or image frames corresponds to a respective time point - e.g., n, etc. - in the plurality of consecutive time points or instants.

[0095] Each pair of adjacent time points or instants in the plurality of consecutive time points or instants may be separated by a fixed time interval such as 1 / 30 second, 1 / 160 second, etc. Event data may be partitioned into a plurality of event data portions in a plurality of fixed time intervals demarcated or delimited by a plurality of pairs of two consecutive time points or instants in the plurality of time points or instants. As illustrated on the left side of FIG. IB, a first event data portion may include or represent a portion of event data or events that occur between a first pair of consecutive time points (n-1) and (n), whereas a second event data portion may include or represent a portion of event data or events that occur between a second pair of consecutive time points (n) and (n+1).

[0096] For each fixed interval such as a fixed time interval between time points / instants (n-1) and (n), the upstream encoder or a video encoder therein may encode a non-event (RGB) imageD24045W001 or image frame into a corresponding image / video frame or image / video NAL unit(s), for example to be carried in a base layer or video stream of an overall bitstream or image / video container file.

[0097] The upstream encoder or an event encoder therein may also encode a corresponding event data portion for events within the fixed time interval up to the time point / index of the non- event (RGB) image into a corresponding event frame or event NAL unit(s) in a plurality of event frames (e.g. , . . . , Efn-i , Efn, Efn+i, . . . , etc.), for example to be carried in an event layer or event stream of the overall bitstream or image / video container file.

[0098] Additionally, optionally or alternatively, event metadata generated for or from the event data may also be partitioned, similar to events represented in the event data. The upstream encoder or an event encoder therein may encode a corresponding event metadata portion in connection with the events within the fixed time interval up to the time point / index of the non- event (RGB) image into event metadata NAL unit(s), for example to be carried in the event layer with the event stream of the overall bitstream or image / video container file.

[0099] As a result, the overall bitstream or image / video container file is generated with the non-event and event fused data by the upstream encoder, as illustrated in FIG. IB.

[0100] As noted, the upstream encoder can perform a method to pack or encode all events between two video frames (corresponding to a fixed time interval by a pair of two consecutive points or instants) into an event frame or event NAL unit(s). For example, events between video frames n- 1 and n can be collected and encoded to form an event stream portion or an event frame or corresponding event NAL unit(s) for video frame n.

[0101] The downstream decoder or a video access unit can access, decode or retrieve - from the overall bitstream or image / video container file - video (bitstream) NAL unit(s), event (bitstream) NAL unit(s) and event metadata NAL unit(s). The downstream decoder or a video decoder therein can decode, construct or generate a sequence of non-event (RGB) image frames from the video (bitstream) NAL unit(s) carried in the overall bitstream or image / video container file. During decoding and processing operations of non-event (RGB) images or videos, the downstream decoder or a video decoder therein can decode, construct or generate corresponding sequences of event data portions and / or event metadata portions from the event (bitstream) NAL unit(s) and event metadata NAL unit(s) carried in the overall bitstream or image / video container file.

[0102] The downstream decoder that receives the overall bitstream or image / video container file with the combination of non-event and event image / video data can perform image event fusion operations or event-assisted or event-guided operations. For example, event data portions in the specific event frames can be fused with non-event (e.g., RGB image / video, etc.) data in theD24045W001 specific video frames to generate multiple intermediate RGB image frames between two adjacent non-event (RGB) image frames decoded from the video stream of the overall bitstream or image / video container file.

[0103] In some operational scenarios, video codec such as H.264 and H.265 video codecs can be used to encode events or event data in a specific event data representation (e.g., voxel based, etc.) into an event frame as described herein. An event frame as described herein can carry or contain events occurring within a (e.g., fixed, NAL, non-event image frame, etc.) time interval.

[0104] Additionally, optionally or alternatively, other encoding methods or approaches may be used to encode raw or processed event data in specific event data representations into event frames. The raw event data may be (e.g., lossless, etc.) raw camera events. The processed event data may be (e.g., lossless, lossy, etc.) events or event data derived from raw camera events. Example encoding operations for encoding lossless raw camera events in event frames are described in I. Schiopu et al., “Lossless compression of event camera frames,” IEEE Signal Processing letters (2022), the contents of which are incorporated herein by reference in entirety.

[0105] It should be noted that encoding events or event data into event frames as described herein may be a lossy process that incurs a loss of event information in the output (e.g., the event frames, etc.) of the lossy process as compared with the input (e.g., raw camera events, processed events or event data derived from raw camera events, etc.) to the lossy process.9. EVENT POINT CLOUD CODING

[0106] In some operational scenarios, G-PCC related encoding or compression operations may be implemented or performed to encode event data into NAL units of an overall bitstream or image / video container file as described herein. For example, all raw individual events covering a time period may be partitioned into a plurality of sets of raw individual events for a plurality of (e.g., fixed, NAL, non-event image frame, etc.) time intervals. For the purpose of illustration only, a time interval as described herein may be an NAL time interval between two successive or consecutive non-event image frames generated by a non-event camera. Each set of raw individual events in the plurality of sets of raw individual events includes or contains all the raw individual events occurring within a respective NAL time interval in the plurality of NAL time interval.

[0107] Each event with its spatial location (corresponding to where an event sensor element detecting the event is located) and its time point / instant (corresponding to when the event sensor element detecting the event; with a much higher temporal resolution of the event camera than the temporal resolution used to acquire or capture non-event image frames by the non-event camera) may be represented as a point in a three dimensional (3D) space. The 3D space has two spatial dimensions such as X and Y axes or Cartesian coordinates for (event) spatial locations of eventsD24045W001 and one temporal dimension such as a Z axis or Cartesian coordinate for (event) time points / instants of the events.

[0108] As a result, for a given NAL time interval, point cloud(s) are formed by points representing some or all events in a corresponding set of raw individual events within the NAL time interval. This point cloud representation of events can then be encoded accordingly using point cloud encoding methods / approaches such as G-PCC. Example G-PCC point cloud coding operations are described in B. Huang et al., “Evaluation of the impact of lossy compression on event camera based computer vision tasks”, SPIE Application of Digital Image Processing XLVI (2023); B. Huang and T. Ebrahimi, “Event data stream compression based on point cloud representation,” ICIP (2023); M. Martini et al., “Lossless compression of neuromorphic vision sensor data based on point cloud representation,” IEEE Access (2022), the contents of all of which are incorporated herein by reference in entirety.

[0109] For the purpose of illustration, it has been described that a hybrid, combination or fusion of non-event image frames, events (or event data in a specific event data representation) and event metadata as described herein can be carried in NAL units. It should be noted that, in other operational scenarios, other types of data units / elements / structures, data formats, file components, protocols, data coding syntaxes, etc., may be used to carry a hybrid, combination or fusion of non-event image frames, events (or event data in a specific event data representation) and event metadata as described herein. For example, in some operational scenarios, instead of or in addition to NAL units, main and / or attendant tracks in an ISO-BMFF file may be used to carry, store, transmit, receive, or deliver, a hybrid, combination or fusion of non-event image frames, events (or event data in a specific event data representation) and event metadata as described herein.10. RESIDUAL BITSTREAM

[0110] In some operational scenarios, instead of transmitting events or event data, an event- aided interpolated image can be generated and used to calculate residual values (or residue) between a non-event aided interpolated image and the event-aided interpolated image. The event- aided interpolated image may be an interpolated image in a sequence of event-aided interpolated image frames interpolated based at least in part on event data for a relatively high frame rate higher than a frame rate at which a sequence of non-event images is generated with a non-event camera. The non-event-aided interpolated image may be an interpolated image in a sequence of non-event-aided interpolated image frames interpolated independent of the event data for the same relatively high frame rate. The residual values can then be encoded into an overall bitstream using available video codecs (e.g., H.264, H265, etc.) or an artificial or advanced neural network,D24045W001 such as a neural field.

[0111] FIG. 2 illustrates an example backwards compatible dual layer encoder / decoder architecture or framework in which an encoder architecture (portion) may be implemented with an upstream encoding device as described herein, whereas a decoder architecture (portion) may be implemented by a downstream recipient / decoding device as described herein.

[0112] As illustrated, in the encoder architecture, the upstream encoder may receive or access two non-event (e.g., RGB, etc.) image frames denoted as IQand / 1;respectively. The non-event images Ioandmay be two successive or consecutive non-event images in a sequence of non- event images generated from image data acquired by a non-event camera.

[0113] In some operational scenarios, the sequence of non-event images that includes the non-event images IQmay have been tone-mapped, enhanced or reshaped from a corresponding sequence of original or input non-event images, for example with video enhancing operations as illustrated in FIG. 1A.

[0114] In addition to the two non-event (RGB) image frames, the upstream encoder can also receive or access corresponding events (originated or derived from event camera data acquired by an event camera from the same scene as the non-event camera) between those two non-event (RGB) image frames.

[0115] The sequence of non-event (RGB) images or image frames including the two non- event (RGB) image frames in the present example can be encoded into, or used to form, a base layer video stream in an overall bitstream or image / video container file.

[0116] Non-event-assisted video frame interpolation (VFI) techniques may be applied to the two non-event (RGB) image frames to generate a sequence of intermediate non-event-assisted interpolated image frames at a relatively high target frame rate as compared with the frame rate at which the sequence of non-event (non-interpolated) image frames is acquired with the non-event camera.

[0117] For example, non-event-assisted image frame interpolation techniques including but not limited to those relating to real-time intermediate flow estimation (RIFE) or FFMPEG may be implemented or applied by the upstream encoder to generate a sequence of intermediate non- event-assisted interpolated image frames denoted as It(also referred to as 7T) at the target rate. Example RIFE operations are described in Z. Huang et al., “Real-time intermediate flow estimation for video frame interpolation,” in Computer Vision - ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part XIV, Berlin, Heidelberg, 2022, p. 624-642, Springer- Verlag, the contents of which are incorporated herein by reference in entirety.D24045W001

[0118] In the meantime, event-assisted or event-guided video frame interpolation techniques may be applied to the events to generate a sequence of intermediate event-assisted interpolated image frames at the target frame rate.

[0119] For example, event-assisted image frame interpolation techniques including but not limited to those relating to TimeLens may be implemented or applied by the upstream encoder to generate a sequence of intermediate event-assisted interpolated image frames denoted as Jt(or interchangeably denoted as JTas shown in FIG. 2) at the target rate.

[0120] A residual image sequence denoted as Rt(or interchangeably denoted as RTas shown in FIG. 2) can be generated to contain residual values (or residue), which are obtained by subtracting each event-assisted or event-based high frame rate image in the interpolated event- assisted image sequence Jtwith a respective non-event-assisted or non-event-based high frame rate image in the interpolated non-event-assisted image sequence It, as follows:Rt= Jt- / tfor t = 0, 1, ..., T-1 (6) where t represents a timestamp or time index (in a time sequence 0, 1, . . . T-l collectively denote as r in FIG. 2) for the t-th image in the image sequence / tor It.

[0121] The residual values can have different signs, namely, positive or negative. A rescaling operation may be performed to adjust, place or limit these residual values in a (positive) data value range such as between 0 and 1, as follows:Rt= 0.5 * (Rt+ 1) (7) where Rtrepresents the t-th image in a positively valued residue image sequence derived with the rescaling operation from the residual image sequence Rt.

[0122] In an example, this residue image sequence (Rtfor the t-th image therein) can then be encoded using available video codecs (e.g., already deployed in the field, to be deployed in the field, etc.) such as those implementing H. 264, H. 265 or the like.

[0123] In another example, as illustrated in FIG. 2, the residue image sequence (Rtfor the t- th image therein) can be (e.g., indirectly, etc.) coded using the neural field approach. For instance, NeRV techniques can be applied or implemented by the upstream encoding device and / or the downstream decoding device as described herein to model the residual image sequence. The NeRV approach may utilize both MLP and CNN blocks to parametrize a video such as a residual video represented by the residual image sequence (Rtfor the t-th image therein) , as follows:where and 9 (which may be collectively referred to as co) represent neural network operationalD24045W001 parameters of the NeRV blocks or the MLP and CNN blocks, respectively. These neural network operational parameters such as weights and / or biases may be optimized or trained using the residual image sequence (e.g., as training data, as to-be-modeled data, etc.) as ground truth. In expression (8) above, y represents the application of positional encoding to the timestamps or time indexes; Rt(or interchangeably denoted as RTas shown in FIG. 2) represents the t-th modeled residual (or residue) image frame in a modeled residual or residual image sequence generated or predicted by the neural field with the optimized neural network operational parameters. Example positional encoding is described in B. Mildenhall et al., “Nerf: Representing scenes as neural radiance fields for view synthesis,” in European conference on computer vision, pp. 405-421. Springer, Cham, 2020, the contents of which are incorporated herein by reference in entirety. The set of parameters and 9 are indicated using a> in FIG. 2.

[0124] While the (original) non-interpolated non-event RGB image frames can be encoded into the video bitstream or base layer of the overall bitstream or image / video container file, the operational parameters co of the neural field ( and 0) and the timestamps r (or the interpolated time sequence at the target frame rate) at which the intermediate or interpolated image frames were generated can be encoded and transmitted in the non-base layer (e.g., event layer, residual layer, etc. ) or a neural field bitstream of the overall bitstream or image / video container file.

[0125] As illustrated in FIG. 2, at the decoder side, the downstream decoding device receives the overall bitstream or image / video container file and decode the base layer or video bitstream in the overall bitstream or image / video container file to construct or reconstruct the sequence of non-event (RGB) images including but not necessarily limited to (e.g., a coded version of, subject to coding errors introduced in the video encoding and decoding operations, etc.) the non- event images IQand I .

[0126] For the purpose of eventually reconstructing or generating the sequence of event- assisted interpolated images between (the respective time points / instances of) the non-event images / 0and / 15the sequence ltof non-event-assisted interpolated images between (the respective time points / instances of) the non-event images IQand 1^ may be first reconstructed or generated, using the same non-event- assisted or non-event-based VFI technique such as RIFE or the like.

[0127] Furthermore, the downstream device decodes the compressed or encoded neural field bitstream in the overall bitstream or image / video container file to construct or reconstruct the optimized neural field operational parameters such as weights and / or biases used in the MLP and CNN blocks of the neural field used to model the residual images. Additionally, based on image and / or event metadata carried in the overall bitstream or image / video container file or the neuralD24045W001 field bitstream therein, the downstream device establishes or determines the sequence of timestamps r representing the sequence of intermediate time points / instants - between the two successive non-event (RGB) images IQand- corresponding to the target temporal resolution at which the event-assisted and non-event-assisted interpolated image sequences were generated.

[0128] Using the optimized neural field (e.g., NeRV, etc.) operational parameters retrieved from the non-base layer (e.g., event layer, residual layer, etc.) or the neural field bitstream, a reconstructed version (denoted as i?T; or interchangeably denoted as Rt) of the rescaled modeled residual image sequence RT(or Rt) for the time sequence or timestamps r may be generated by the neural as follows:where 9 and < > represent a decoded version of the neural network operational parameters cp and 9 such as the weights and / or biases of the neural field or neural network blocks therein. Due to the encoding (or compression) operations at the encoder side and the decoding (or decompression) operations at the decoder side of the neural field parameters, there may be a loss of information or precision and / or an introduction of coding or quantization errors, unless these operations are implemented as lossless compression / decompression operations. Hence, the neural field parameters 0 and cp such as weights and / or biases as decoded by the downstream decoding device may or may not be identical to the corresponding neural field parameters 9 and cp obtained from training or modeling the neural field by the upstream encoding device. As a result, the decoded version RTof the rescaled modeled residual image sequence RT- as generated by the downstream decoding device - may or may not be identical to the rescaled modeled residual image sequence RTas trained or modeled with the neural field by the upstream encoding device. For example, RTmay approximate, or may be the same as, RT. subject to the loss of information due to compression / decompression and / or quantization / coding errors incurred in these operations.

[0129] As this decoded version RTof the modeled residual image sequence Rtwas generated with a rescaling operation, a decoded version of the (pre-rescaled) residual image sequence Rt(or interchangeably RT)may be generated with an inverse rescaling operation as follows:

[0130] This residual image sequence Rt- decoded by way of the neural field and rescaling operation - can then be added back to the non-event-assisted interpolated image sequence generated from applying non-event-assisted VFI techniques to the two successive non-event (RGB) images Ioand I in the base layer or non-event image (RGB) sequence to generate,D24045W001 construct or reconstruct an event- assisted interpolated image sequence, as follows:Jt= Rt+ Itfor t = 0, 1, ..., T-l (11) where Itrepresents the t-th image in the non-event-assisted interpolated base layer image sequence; and / t(or interchangeably / ,.) represents the t-th image (at time t) in the (final reconstructed) event-assisted image sequence.11. IMAGE RECONSTRUCTION QUALITY

[0131] To validate techniques as described herein, experiments have been provided with an image dataset such as the HS-RGB datasets. The value of the proposed techniques has also been validated by experimental results.

[0132] In connection with these experiments, non-event-assisted or non-event-guided interpolated images may be generated by applying non-event- assisted image interpolation techniques such as the VFI or RIFE techniques to non-event non-interpolated images.

[0133] Reconstructed event- assisted or event-guided interpolated images may also be generated based at least in part on reconstructed residual images generated from a neural field with optimized neural field operational parameters.

[0134] As noted, the optimized neural field operational parameters of the neural field can be obtained or generated by first training or modeling the neural field with non-reconstructed residual images.

[0135] In turn, the non-reconstructed residual images can be obtained or computed as differences between the non-event- assisted or non-event-guided interpolated images and nonreconstructed event-assisted or event-guided interpolated images. While the non-event-assisted or non-event-guided interpolated images are generated by applying the non-event-assisted image interpolation techniques such as the VFI or RIFE techniques, the non-reconstructed event- assisted or event-guided interpolated images are generated by applying event-assisted image interpolation techniques such as the TimeLens techniques to non-event (non-interpolated) images.

[0136] The neural field may be implemented in different models with different total numbers of (e.g., 1.5 million, 3 million, etc.) neural field operational parameters.

[0137] In some operational scenarios, given two successive or consecutive non-event images in the non-event non-interpolated image sequence, non-event-assisted intermediate images or image frames may be generated using RIFE techniques.

[0138] The non-event-assisted intermediate images or image frames together with the non- event non-interpolated images retained or unchanged by the RIFE techniques may correspond to a relatively high temporal resolution as compared with a relatively low temporal resolution atD24045W001 which the non-event non-interpolated image sequence is generated, for example with a non-event camera.

[0139] By way of example but not limitation, the relatively high temporal resolution may be eight (8) times the relatively low temporal resolution. Hence, for fifteen (15) successive or consecutive images in the non-event non-interpolated image sequence, a resultant non-event- assisted interpolated image sequence may be composed of 120 images, fifteen (15) of which are retained or unchanged non-event non-interpolated images.

[0140] An event-assisted interpolated image sequence - e.g., 120 event- assisted interpolated images - at the same relatively high temporal resolution may be generated by applying the eventbased interpolation techniques or the TimeLens techniques to event data.

[0141] A residual image sequence generated by subtraction between the non-event-assisted interpolated image sequence and the event-assisted interpolated image sequence can then be used as training data or ground truth for the neural field (e.g., as a flat field, etc.) with its neural networks to model or learn. In an example, only the interpolated images (e.g., 104 interpolated images, etc.) are used to train, model or fit the neural field. In another example, both the retained / unchanged and interpolated images (e.g., all 120 images, etc.) are used to train, model or fit the neural field.

[0142] The event-assisted interpolated image / video data can be used to contain additional image details that may be absent from non-event-assisted interpolated image / video data.

[0143] To measure effectiveness of the event-assisted or event-guided image interpolation techniques, image / video quality measurements such as peak-signal-noise-ratio (PSNR) may be computed between the reconstructed interpolated image frames constructed from residual images modeled with the neural field (e.g., NeRV, etc.) and the ground truth or GT (used to train the neural field) generated at least in part using the event-based interpolation techniques TimeLens. It is observed that the PSNR values are relatively high indicating the effectiveness of the techniques as described herein.

[0144] Additionally, optionally or alternatively, other VFI interpolation techniques such as FFMPEG may be used in place of RIFE. It is observed that the PSNR values are still relatively high with image reconstruction results indicating the effectiveness and robustness of the techniques as described herein.12. EVENT METADATA

[0145] For the purpose of illustration, it has been described that NAL units may be used to carry event data and event metadata generated from raw event camera data acquired with an event camera from a scene. The event data or events represented therein may be partitioned intoD24045W001(e.g., fixed, etc.) time intervals synchronized or aligned with non-event (e.g., RGB) images generated from non-event camera images acquired with a non-event camera from the same scene.

[0146] It should be noted that in various other operational scenarios, other types of event and / or non-event data units or files or formats other than NAL units may be used to carry event data and event metadata as described herein. For example, in some operational scenarios, one or more specific tracks - e.g., separate from tracks carrying non-event images, etc. - in an ISO- BMFF compliant bitstream or image / video container file may be used to carry event data and / or event metadata.

[0147] The event data and / or the event metadata or portions thereof in the specific tracks of the ISO-BMFF compliant bitstream or image / video container file may carry their respective specific time codes or timestamps. Using these specific time codes or timestamps represented in the event data and / or event metadata in the specific tracks, specific event data portions or events that occur with a given time interval such as between two successive or consecutive non-event non-interpolated images carried in separate tracks of the ISO-BMFF compliant bitstream or image / video container file may be identified.

[0148] Additionally, optionally or alternatively, in place of or in addition to event metadata or portions, a residual event based layer may be carried in one or more specific tracks of an ISO- BMFF compliant bitstream or image / video container file. For example, the residual event based layer in the specific tracks of an ISO-BMFF compliant bitstream or image / video container file may include a residual image sequence or neural field operational parameters of a neural field that model or train on the residual image sequence.

[0149] Event metadata as described herein may be of a wide variety of event metadata (information) types or fields.

[0150] A first example of event metadata type or field is event picture spatial resolution, which may specify a spatial resolution of (event sensor elements in) an event camera / sensor. The event picture spatial resolution may identify or specify the width and height of (an array or an imaging plane area formed by the event sensor elements in) the event camera / sensor. This field can be used in the downstream recipient device or decoder to generate other event data representations or formats such as a voxel grid representation.

[0151] A second example of event metadata type or field is event temporal information, which may specify a sampling rate or temporal resolution of (the event sensor elements in) the event camera / sensor. The event temporal information may identify or specify how fast the event sensor or an event sensor element therein can capture the event data. Additionally, optionally, or alternatively, this field may identify or specify temporal sampling rates that are used to createD24045W001 various event representations supported by the event sensor / camera. This field can be used in the downstream recipient device or decoder to estimate, determine or identify a maximum frame rate (possible), for example in event-assisted or event-guided image interpolation operations.

[0152] A third example of event metadata type or field is event representation information, which may specify an event representation format used to represent the events or event data. An enumerated value (Enum) may be used for each event representation format among candidate event representation formats. Additionally, optionally, or alternatively, in some operational scenarios, in which an event voxel grid representation format is used, this field may (e.g., further, etc.) contain or specify the total number of histogram bins used in operations for event voxel grid generation. This field can be used in the downstream recipient device or decoder to a specific type of event representation used during encoding operations performed by the upstream encoding device.

[0153] A fourth example of event metadata type or field is event contrast threshold settings, which may specify or store positive and negative (e.g., contrast, intensity change, luminance change, etc.) threshold settings (positive and negative) at the time of event capture by an event sensor element in the event sensor / camera. This field may be set with specific value(s) configured by the vendor, user, or operator of the event sensor / camera and / or used by specific image reconstruction algorithms. Like non-event camera ISO, the values of this field or type for the event sensor / camera may be specified according to industry standards or proprietary specifications.

[0154] A fifth example of event metadata type or field is scene luminance, which may specify or store a measured average scene luminance. This field may be used by specific image reconstruction algorithms.

[0155] A sixth example of event metadata type or field is camera calibration parameters, which may specify or store camera intrinsic and extrinsic parameters (or camera intrinsics and extrinsics). These camera parameters may be set, measured or adjusted based on pre-calibrated event sensor / camera and lens arrangements therein. Camera calibration parameters in the event metadata may be used to help with spatial alignment of the event sensor / cameras with the non- event camera.13. ARCHITECTURAL EXTENSIONS AND ALTERNATIVES

[0156] For the purpose of illustration, it has been described that event related data or event- related image / video data may be encoded, transmitted, received, decoded, or processed, in a backward compatible encoder and decoder architecture or framework.

[0157] It should be noted that, in various operational scenarios, event related data or eventD24045W001 related image / video data may be encoded, transmitted, received, decoded, or processed, in backward compatible or non-backwards compatible encoder and decoder architectures or frameworks.

[0158] For example, a neural field may be implemented to capture some or all event-related data or event related image / video data. Some neural field approaches such as those related to HNeRV - which is described in H. Chen et al., “HNeRV: A Hybrid Neural Representation for Videos,” arXiv, Apr. 05, 2023. doi: 10.48550 / arXiv.2304.02633, the contents of which are incorporated herein by reference in entirety - may perform better in representing non-residual image sequences than representing residual image sequences. Since the residual image sequences are obtained from taking differences between two image sequences with a large amount of similar or same image / video data, the residual image sequence may contain images of pixel values mostly zeros (0s) or close to zeros. It may be difficult or inefficient for an upstream device or encoder to use these neural field approaches performing better in representing non-residual images to handle or encode residual images with sparse information.

[0159] In some operational scenarios, non-residual event related image / video data may be processed or modeled with a neural field instead of residual event related image / video data. As the non-residual event related image / video data is used by the upstream device or encoder to model or train the neural field or to generate optimized neural field operational parameters, the neural field can be used by the downstream device or decoder that receives the optimized neural field operational parameters to generate or predict or reconstruct the non-residual event-related image / video data that were used to train or model the neural field.

[0160] As the non-residual event image / video data can be transferred to or reconstructed by the downstream device by way of the optimized neural field operational parameters, it may not be necessary for the upstream device or encoder to provide non-event images in some operational scenarios. Hence, in these operational scenarios, in place of a base layer encoded with non-event images, the base layer of an overall bitstream or image / video container file may be directly encoded with the optimized neural field operational parameters to be used by the downstream device or decoder to generate, predict or reconstruct the event-related residual images.

[0161] In some other operational scenarios, the base layer can still be reserved or used for carrying or sending non-event images for the purpose of backwards compatibility and use another layer to carry the optimized neural field operational parameters at the expense of sending largely redundant image / video data in both layers.14. EXAMPLE ALTERNATIVE ONE

[0162] FIG. 3A illustrates an example dual layer encoder / decoder architecture or frameworkD24045W001 in which an encoder architecture (portion) may be implemented with an upstream encoding device as described herein, whereas a decoder architecture (portion) may be implemented by a downstream recipient / decoding device as described herein.

[0163] As illustrated, at the encoder side, the upstream encoder may receive or access two non-event (e.g., RGB, etc.) image frames denoted as l0and / 1?respectively. The non-event images Ioand I±may be two successive or consecutive non-event images in a sequence of non- event images generated from image data acquired by a non-event camera.

[0164] In some operational scenarios, the sequence of non-event images that includes the non-event images 1Qandmay have been tone-mapped, enhanced or reshaped from a corresponding sequence of original or input non-event images, for example with video enhancing operations as illustrated in FIG. 1 A.

[0165] In addition to the two non-event (RGB) image frames, the upstream encoder can also receive or access corresponding events (originated or derived from event camera data acquired by an event camera from the same scene as the non-event camera) between those two non-event (RGB) image frames.

[0166] The sequence of non-event (RGB) images or image frames including the two non- event (RGB) image frames in the present example can be encoded into, or used to form, a base layer video stream in an overall bitstream or image / video container file.

[0167] Non-event-assisted video frame interpolation (VFI) techniques may be applied to the two non-event (RGB) image frames to generate a sequence of intermediate non-event-assisted interpolated image frames at a relatively high target frame rate as compared with the frame rate at which the sequence of non-event (non-interpolated) image frames is acquired with the non-event camera.

[0168] For example, non-event-assisted image frame interpolation techniques including but not limited to those relating to real-time intermediate flow estimation (RIFE) or FFMPEG may be implemented or applied by the upstream encoder to generate a sequence of intermediate nonevent-assisted interpolated image frames denoted as IT(also referred to as / t) at the target rate.

[0169] In the meantime, event-assisted or event-guided video frame interpolation techniques may be applied to the events to generate a sequence of intermediate event-assisted interpolated image frames at the target frame rate.

[0170] For example, event-assisted image frame interpolation techniques including but not limited to those relating to TimeLens may be implemented or applied by the upstream encoder to generate a sequence of intermediate event-assisted interpolated image frames denoted as / T(also referred to as Jt) at the target rate.D24045W001

[0171] A residual image sequence denoted as RTcan be generated to contain residual values (or residue), which are obtained by subtracting each event-assisted or event-based high frame rate image in the interpolated event-assisted image sequence JTwith a respective non-event-assisted or non-event-based high frame rate image in the interpolated non-event-assisted image sequence IT

[0172] For the purpose of illustration, a neural field such as relating to the HNeRV method or approach may be used to model images that incorporate event-related data. As the neural field may perform relatively efficiently at modeling non-residual natural image sequences such as nonresidual RGB images or pixel values, the upstream device or encoder can generate three pseudo image sequences from the non-event non-residual interpolated image sequence ITand the event- related residual image sequence RTas follows.

[0173] More specifically, to generate the first pseudo image sequence, the upstream device or encoder takes two color channels such as blue (B) and green (G) (among the three color channels of a color space such as the RGB color space) from the t-th non-event interpolated image 1Tand concatenates or combines these two channels (denoted as IBTand IGT, respectively) of the t-th non-event interpolated image ITwith one color channel (among the three color channels of a color space such as the RGB color space) such as red (R) denoted as RRTof the t-th residue image RT.

[0174] Likewise, to generate the second pseudo image sequence, the upstream device or encoder takes two color channels such as blue (B) and red (R) from the t-th non-event interpolated image ITand concatenates or combines these two channels ( / Btand IRr, respectively) of t-th non-event interpolated image ITwith one color channel green (G) RGTfrom the t-th residue image RT.

[0175] To generate the third pseudo image sequence, the upstream device or encoder takes two color channels such as green (G) and red (R) from the t-th non-event interpolated image ITand concatenates or combines these two channels (ICTand IRT, respectively) of the t-th non-event interpolated image ITwith one color channel blue (B) RBTfrom the t-th residue image RT.

[0176] As a result, three non-residual pseudo image sequences are generated from the non- event interpolated image sequence ITand the event-assisted or event-related residual image sequence RTAs illustrated in FIG. 3A, the upstream device or encoder can use three different HNeRV neural fields to be trained to model the three pseudo image sequences, respectively.

[0177] It should be noted that in various operational scenarios, these and other combinations of color channels in RGB color spaces and / or non-RGB color spaces may be used to concatenate into pseudo image sequences with which the neural field techniques can perform operationsD24045W001 relatively efficiently.

[0178] Three sets of neural field / network operational parameters (denoted as S)1, ai2,respectively) and the timestamps r (or the interpolated time sequence at the target frame rate) at which the intermediate or interpolated image frames were generated can be encoded and transmitted in an overall bitstream or image / video container file.

[0179] In some operational scenarios, to support backwards compatibility, the non-event image sequence including the non-event images IQand may be encoded in a base layer of the overall bitstream or image / video container file, while the three sets of neural field / network operational parametersM2, m3and the timestamps r may be encoded in a non-base layer of the overall bitstream or image / video container file.

[0180] In other operational scenarios, the non-event image sequence including the non-event images / 0may be omitted or prevented from being encoded in the overall bitstream or image / video container file. The three sets of neural field / network operational parameters a)1, M2, <o3and the timestamps r may be encoded in the overall bitstream or image / video container file.

[0181] At the decoder side, three different HNeRV decoders or neural fields may be used to operate with the three sets of the optimized neural field operational parametersG)2. m3. respectively, and the timestamps r to generate, predict or reconstruct the three event-related pseudo image sequences respectively.

[0182] Each of these three event-related pseudo image sequences respectively contains one channel of the event-assisted or event-related residual image sequence. Hence, the event-assisted or event-related residual image sequence (denoted as RT) can be reconstructed from these channels in the three event-related pseudo image sequences.

[0183] In some operational scenarios, in which the non-event images Ioandare carried in the base layer of the overall bitstream or image / video container file, the non-event-assisted interpolated image sequence (denoted as / ) between (the respective time points / instances of) the non-event images 1Qand l may be first reconstructed or generated, using the same non- event-assisted or non-event-based VFI technique such as RIFE or the like.

[0184] In other operational scenarios, in which the non-event images Ioand / qare not carried in the overall bitstream or image / video container file, the three event-related pseudo image sequences can be used to generate or reconstruct the non-event interpolated image sequence.

[0185] The event-assisted or event-related residual image sequence can be combined with the non-event interpolated image sequence in addition operations to generate or reconstruct the event-assisted interpolated image sequence.

[0186] Image reconstruction quality of the event-assisted interpolation image sequence canD24045W001 be measured, for example using the PSNR value. It is observed that the image reconstruction quality is relatively high and similar to what has been observed with the backwards compatible approach.15. EXAMPLE ALTERNATIVE TWO

[0187] FIG. 3B illustrates another example dual layer encoder / decoder architecture or framework in which an encoder architecture (portion) may be implemented with an upstream encoding device as described herein, whereas a decoder architecture (portion) may be implemented by a downstream recipient / decoding device as described herein.

[0188] In the architecture of FIG. 3A, there are three neural fields used to model three pseudo image sequences, respectively. As a result, the total number of neural field operational parameters that are to be transferred between the upstream device or encoder and the downstream device or decoder in FIG. 3 A may be three times of the total number of neural field operational parameters of a single neural field.

[0189] In comparison, in the architecture of FIG. 3B, there are two neural fields used to model two pseudo image sequences, respectively. As a result, the total number of neural field operational parameters that are to be transferred between the upstream device or encoder and the downstream device or decoder in FIG. 3B may be two times of the total number of neural field operational parameters of a single neural field.

[0190] As illustrated in FIG. 3B, at the encoder side, a residual image sequence RTcan be generated to contain residual values (or residue), which are obtained by subtracting each event- assisted or event-based high frame rate image in an interpolated event-assisted image sequence JTwith a respective non-event-assisted or non-event-based high frame rate image in an interpolated non-event-assisted image sequence IT.

[0191] For the purpose of illustration, a neural field such as relating to the HNeRV method or approach may be used to model images that incorporate event-related data. As the neural field may perform relatively efficiently at modeling non-residual natural image sequences such as nonresidual RGB images or pixel values, the upstream device or encoder can generate two pseudo image sequences from the non-event non-residual interpolated images ITand the event-related residual images RTas follows.

[0192] More specifically, to generate the first pseudo image sequence, the upstream device or encoder takes one color channel such as blue (B) (among the three color channels of a color space such as the RGB color space) from the t-th non-event interpolated image ITand concatenates or combines the channel (denoted as IBT) of the t-th non-event interpolated image ITwith two color channels (among the three color channels of a color space such as the RGBD24045W001 color space) such as green (G) and red (R) denoted as RGT, RRT, respectively, of the t-th residue image RT.

[0193] To generate the second pseudo image sequence, the upstream device or encoder takes two color channels such as green (G) and red (R) from the t-th non-event interpolated image ITand concatenates or combines these two channels (IGTand IRT, respectively) of t-th non-event interpolated image ITwith one color channel blue (B) RBTfrom the t-th residue image RT.

[0194] As a result, two non-residual pseudo image sequences are generated from the non- event interpolated image sequence ITand the event-assisted or event-related residual image sequence RT. As illustrated in FIG. 3B, the upstream device or encoder can use two different HNeRV neural fields to be trained to model the two pseudo image sequences, respectively.

[0195] It should be noted that in various operational scenarios, these and other combinations of color channels in RGB color spaces and / or non-RGB color spaces may be used to concatenate into pseudo image sequences with which the neural field techniques can perform operations relatively efficiently.

[0196] Two sets of neural field / network operational parameters (denoted as to,, irrespectively) and the timestamps r (or the interpolated time sequence at the target frame rate) at which the intermediate or interpolated image frames were generated can be encoded and transmitted in an overall bitstream or image / video container file.

[0197] In some operational scenarios, to support backwards compatibility, the non-event image sequence including the non-event images l0and may be encoded in a base layer of the overall bitstream or image / video container file, while the two sets of neural field / network operational parameters M1,and the timestamps r may be encoded in a non-base layer of the overall bitstream or image / video container file.

[0198] In other operational scenarios, the non-event image sequence including the non-event images Iomay be omitted or prevented from being encoded in the overall bitstream or image / video container file. The two sets of neural field / network operational parameters S)1, m2and the timestamps r may be encoded in the overall bitstream or image / video container file.

[0199] At the decoder side, two different HNeRV decoders or neural fields may be used to operate with the two sets of the optimized neural field operational parameters ai1, m2, respectively, and the timestamps r to generate, predict or reconstruct the three event-related pseudo image sequences respectively.

[0200] Collectively, the two event-related pseudo image sequences contain the event-assisted or event-related residual image sequence. Hence, the event-assisted or event-related residual image sequence (denoted as RT) can be reconstructed from these channels in the two event-D24045W001 related pseudo image sequences.

[0201] In some operational scenarios, in which the non-event images l0andare carried in the base layer of the overall bitstream or image / video container file, the non-event-assisted interpolated image sequence (denoted as / T’) between (the respective time points / instances of) the non-event images 70and I may be first reconstructed or generated, using the same nonevent-assisted or non-event-based VFI technique such as RIFE or the like.

[0202] In other operational scenarios, in which the non-event images Zoandare not carried in the overall bitstream or image / video container file, the two event-related pseudo image sequences can be used to generate or reconstruct the non-event interpolated image sequence.

[0203] The event-assisted or event-related residual image sequence can be combined with the non-event interpolated image sequence in addition operations to generate or reconstruct the event-assisted interpolated image sequence.

[0204] Image reconstruction quality of the event-assisted interpolation image sequence can be measured, for example using the PSNR value. It is observed that the image reconstruction quality is relatively high and similar to what has been observed with the backwards compatible approach.16. OTHER EXAMPLE ALTERNATIVES

[0205] In some operational scenarios, instead of using multiple neural fields, a non-event and event related image / video encoding and decoding architecture may be implemented with a single neural field.

[0206] For example, a single neural field based on the HNeRV architecture may be implemented at the encoder side to model a pseudo image sequence represented with six (6) color channels instead of three (3) color channels of a color space. The six (6) color channels include three color channels from three channels in which a non-event-assisted or non-event related (e.g., VFI, RIFE, etc.) interpolated image sequence is represented, concatenated with additional three (3) color channels in which an event-assisted or event-related residual (or residue) image sequence is represented. After modeling or training the single neural field with the single pseudo image sequence represented in the six (6) color channels, optimized neural field operational parameters of the neural field can be carried or sent in an overall image bitstream or image / video container file from the upstream device or encoder to the downstream device or decoder.

[0207] Correspondingly, at the decoder side, a single neural field based on the HNeRV architecture may be implemented to operate with the optimized neural field operational parameters decoded from overall image bitstream or image / video container file to reconstruct theD24045W001 single pseudo image sequence represented in the six (6) color channels. The pseudo image sequence may be used to generate or reconstruct the non-event interpolated image sequence (e.g., at a relatively high temporal resolution or frame rate, etc.) from three color channels of the pseudo image sequence. Additionally, optionally or alternatively, the upstream device can reconstruct or generate the event- assisted interpolated image sequence by adding the non-event interpolated image sequence with the event-assisted interpolated residual image sequence from the other three color channels of the pseudo image sequence.17. EXAMPLE PROCESS FLOWS

[0208] FIG. 4A illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., an upstream device, an encoding device / module, a transcoding device / module, a media device / module, etc.) may perform this process flow.

[0209] In block 402, an upstream device such as an upstream encoding device derives non- event camera data from a first non-event image and a second non-event image. The first non- event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera.

[0210] In block 404, the upstream device derives event camera data from a sequence of raw events. The sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor. The sequence of event time points is between the first time point and the second time point.

[0211] In block 406, the upstream device encodes event related image data into an overall image container. The event related image data is derived from a combination of both the event camera data and the non-event camera data. A recipient device of the overall image container is caused by the event related image data of the overall image container to render at least one display image on an image display. The at least one display image is derived from the event related image data.

[0212] In an embodiment, the first non-event image and the second non-event image represent one of: two still images, or two successive non-event images in a sequence of non- event images forming a video.

[0213] In an embodiment, the event related image data is encoded into one of: one or more network abstract layer (NAL) units, or one or more tracks in an ISO-BMFF file.

[0214] In an embodiment, the image container contains a base layer that carries a video stream encoded with non-event image including the first and second non-event images. TheD24045W001 image container contains a non-base layer that carries the event camera data.

[0215] In an embodiment, the image container contains a base layer that carries a video stream encoded with non-event images including the first and second non-event images. The image container contains a non-base layer that carries neural field operational parameters of a neural field that is trained with a residual image sequence generated as differences between a non-event interpolated image sequence and an event-assisted interpolated image sequence. The non-event interpolated image sequence is generated by applying non-event video frame interpolation operations to the first and second non-event images. The event-assisted interpolated image sequence is generated by applying event-based interpolation techniques based at least in part on the event camera data.

[0216] In an embodiment, the neural field represents a combination of one or more multilayer perceptron (MLP) blocks, one or more convolutional neural network (CNN) blocks, etc.

[0217] In an embodiment, the neural field represents an implicit function that is continuous over time. The neural field is used to generate predicted residual images at a target frame rate that is higher than a non-event image frame rate supported by the non-event camera.

[0218] In an embodiment, the event-assisted interpolated image sequence is generated at least in part by applying one or more of: image interpolation techniques that utilize a voxel grid representation of the event camera data in combination with the first and second non-event images, TimeLens image interpolation techniques, SuperFast image interpolation techniques, EVDI, REFID, CBMNet, or other event-assisted image interpolation techniques.

[0219] In an embodiment, the non-event- assisted interpolated image sequence is generated at least in part by applying one or more of: RIFE, FFMPEG, or other non-event-assisted image interpolation techniques.

[0220] In an embodiment, the image container includes event metadata in the event related image data separate from event data in the event related image data.

[0221] In an embodiment, the image container is excluded from any layer encoded with non- event images.

[0222] In an embodiment, the image container carries neural field operational parameters of one or more neural fields that are trained with one or more pseudo image sequences, respectively. Each of the one or more pseudo image sequences for a respective neural field of the one or more neural fields is generated by combining one or more first color channels of a non-event interpolated image sequence generated at least in part from the first and second non-event images and one or more second color channels of an event-assisted interpolated image sequence generated at least in part on the event camera data.D24045W001

[0223] In an embodiment, the one or more pseudo image sequences includes three pseudo image sequences each of which is generated by combining two color channels of the non-event interpolated image sequence and one color channel of the event- assisted interpolated image sequence.

[0224] In an embodiment, the one or more pseudo image sequences includes two pseudo image sequences a first one of which is generated by combining two color channels of the non- event interpolated image sequence and one color channel of the event-assisted interpolated image sequence, and a second one of which is generated by combining one color channel of the non- event interpolated image sequence and two color channels of the event-assisted interpolated image sequence.

[0225] In an embodiment, the one or more pseudo image sequences includes only one pseudo image sequence that is generated by combining three color channels of the non-event interpolated image sequence and three color channels of the event-assisted interpolated image sequence.

[0226] In an embodiment, the event camera data is represented in one of: individual luminance change events forming an event sequence, an event frame in which events in a spatiotemporal neighborhood are accumulated as a whole, an event time surface representation, an event voxel grid representation, an event point cloud, an event neural field, or another event data representation.

[0227] In an embodiment, at least one of the event camera data or the sequence of raw events is specified in one of: a proprietary event format defined by a provider of the event camera, an industry standard event format, a fixed length event data format, a DAT format, a Prophesee event data format, an IniVation event data format, a Hierarchical Data Format version 5, a Comma Separated Value (CVS) format, a compressed point cloud representation format, a compressed data format in accordance with one of H.264, H.265 or another video data coding standard, an event field using MPEG Neural Network Compression (NNC), a lossless event data format, or a lossy event data format.

[0228] In an embodiment, the event camera data includes a specific portion derived from events occurring between the first and second image frame time points. The specific portion of the event camera data is encoded into a single network abstract layer (NAL) unit.

[0229] FIG. 4B illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., a downstream device, one or more event data codecs, a decoding device / module, a transcoding device / module, a media device / module, etc.) may perform this process flow.

[0230] In block 452, a downstream device such as a downstream decoding device decodesD24045W001 event related image data from an overall image container. The event related image data has been derived and encoded into the overall image container by an upstream device from a combination of both event camera data and non-event camera data. The non-event camera data has been derived by the upstream device from a first non-event image and a second non-event image. The first non-event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non- event camera. The event camera data is derived by the upstream device from a sequence of raw events. The sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor. The sequence of event time points is between the first time point and the second time point.

[0231] In block 454, the downstream device generates at least one display image from the event related image data.

[0232] In block 456, the downstream device renders the at least one display image on an image display.

[0233] In an embodiment, image container represents one of: an image signal that is exclusive of a base layer encoded with the first and second non-event images, or an image signal that includes a base layer encoded with the first and second non-event images.

[0234] In an embodiment, a computing device such as a display device, a mobile device, a set-top box, a multimedia device, etc., is configured to perform any of the foregoing methods. In an embodiment, an apparatus comprises a processor and is configured to perform any of the foregoing methods. In an embodiment, a non-transitory computer readable storage medium, storing software instructions, which when executed by one or more processors cause performance of any of the foregoing methods.

[0235] In an embodiment, a computing device comprising one or more processors and one or more storage media storing a set of instructions which, when executed by the one or more processors, cause performance of any of the foregoing methods.

[0236] Note that, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments.18. IMPLEMENTATION MECHANISMS - HARDWARE OVERVIEW

[0237] Embodiments of the present invention may be implemented with a computer system, systems configured in electronic circuitry and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), anD24045W001 application specific IC (ASIC), and / or apparatus that includes one or more of such systems, devices or components. The computer and / or IC may perform, control, or execute instructions relating to the adaptive perceptual quantization of images with enhanced dynamic range, such as those described herein. The computer and / or IC may compute any of a variety of parameters or values that relate to the adaptive perceptual quantization processes described herein. The image and video embodiments may be implemented in hardware, software, firmware and various combinations thereof.

[0238] Certain implementations of the inventio comprise computer processors which execute software instructions which cause the processors to perform a method of the disclosure. For example, one or more processors in a display, an encoder, a set top box, a transcoder or the like may implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processors. Embodiments of the invention may also be provided in the form of a program product. The program product may comprise any non-transitory medium which carries a set of computer- readable signals comprising instructions which, when executed by a data processor, cause the data processor to execute a method of an embodiment of the invention. Program products according to embodiments of the invention may be in any of a wide variety of forms. The program product may comprise, for example, physical media such as magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAM, or the like. The computer-readable signals on the program product may optionally be compressed or encrypted.

[0239] Where a component (e.g. a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including a reference to a "means") should be interpreted as including as equivalents of that component any component which performs the function of the described component (e.g., that is functionally equivalent), including components which are not structurally equivalent to the disclosed structure which performs the function in the illustrated example embodiments of the invention.

[0240] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination. Such special-purposeD24045W001 computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices or any other device that incorporates hard-wired and / or program logic to implement the techniques.

[0241] For example, FIG. 5 is a block diagram that illustrates a computer system 500 upon which an embodiment of the invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general purpose microprocessor.

[0242] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions.

[0243] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions.

[0244] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane.

[0245] Computer system 500 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques as described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may beD24045W001 read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In other embodiments, hard-wired circuitry may be used in place of or in combination with software instructions.

[0246] The term “storage media” as used herein refers to any non-transitory media that store data and / or instructions that cause a machine to operation in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH- EPROM, NVRAM, any other memory chip or cartridge.

[0247] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications.

[0248] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone line using a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infra-red transmitter to convert the data to an infra-red signal. An infra-red detector can receive the data carried in the infra-red signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504.

[0249] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) cardD24045W001 to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0250] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet” 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are example forms of transmission media.

[0251] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518.

[0252] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution.19. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS

[0253] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is claimed embodiments of the invention, and is intended by the applicants to be claimed embodiments of the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.Enumerated Exemplary Embodiments

[0254] The invention may be embodied in any of the forms described herein, including, but not limited to the following Enumerated Example Embodiments (EEEs) which describe structure,D24045W001 features, and functionality of some portions of embodiments of the present invention.

[0255] EEE1. A method, comprising: deriving non-event camera data from a first non-event image and a second non-event image, wherein the first non-event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera; deriving event camera data from a sequence of raw events, wherein the sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor, wherein the sequence of event time points is between the first time point and the second time point; encoding event related image data into an overall image container, wherein the event related image data is derived from a combination of both the event camera data and the non-event camera data, wherein a recipient device of the overall image container is caused by the event related image data of the overall image container to render at least one display image on an image display, wherein the at least one display image is derived from the event related image data.

[0256] EEE2. The method as recited in EEE1 , wherein the first non-event image and the second non-event image represent one of: two still images, or two successive non-event images in a sequence of non-event images forming a video.

[0257] EEE3. The method as recited in EEE1 or EEE2, wherein the event related image data is encoded into one of: one or more network abstract layer (NAL) units, or one or more tracks in an ISO-BMFF file.

[0258] EEE4. The method as recited in any of EEE1-EEE3, wherein the image container contains a base layer that carries a video stream encoded with non-event image including the first and second non-event images, wherein the image container contains a non-base layer that carries the event camera data.

[0259] EEE5. The method as recited in any of EEE1-EEE4, wherein the image container contains a base layer that carries a video stream encoded with non-event images including the first and second non-event images, wherein the image container contains a non-base layer that carries neural field operational parameters of a neural field that is trained with a residual image sequence generated as differences between a non-event interpolated image sequence and an event-assisted interpolated image sequence; wherein the non-event interpolated image sequence is generated by applying non-event video frame interpolation operations to the first and second non-event images; wherein the event- assisted interpolated image sequence is generated byD24045W001 applying event-based interpolation techniques based at least in part on the event camera data.

[0260] EEE6. The method as recited in EEE5, wherein the neural field operational parameters comprise weights and biases of the neural field that are optimized by training the neural field using the residual image sequence as ground truth.

[0261] EEE7. The method as recited in EEE5 or EEE6, wherein the neural field represents a combination of one or more multi-layer perceptron (MLP) blocks and one or more convolutional neural network (CNN) blocks.

[0262] EEE8. The method as recited in any of EEE5- EEE7, wherein the neural field represents an implicit function that is continuous over time, wherein the neural field is used to generate predicted residual images at a target frame rate that is higher than a non-event image frame rate supported by the non-event camera.

[0263] EEE9. The method as recited in any of EEE5- EEE8, wherein the event-assisted interpolated image sequence is generated at least in part by applying one or more of: image interpolation techniques that utilize a voxel grid representation of the event camera data in combination with the first and second non-event images, TimeLens image interpolation techniques, SuperFast image interpolation techniques, EVDI, REFID, CBMNet, or other event- assisted image interpolation techniques.

[0264] EEE10. The method as recited in any of EEE5- EEE9, wherein the non-event- assisted interpolated image sequence is generated at least in part by applying one or more of: RIFE, FFMPEG, or other non-event-assisted image interpolation techniques.

[0265] EEE11. The method as recited in any of EEE1- EEE 10, wherein the image container includes event metadata in the event related image data separate from event data in the event related image data.

[0266] EEE12. The method as recited in any of EEE1- EEE11, wherein the image container is excluded from any layer encoded with non-event images.

[0267] EEE13. The method as recited in any of EEE1- EEE 12, wherein the image container carries neural field operational parameters of one or more neural fields that are trained with one or more pseudo image sequences, respectively; wherein each of the one or more pseudo image sequences for a respective neural field of the one or more neural fields is generated by combining one or more first color channels of a non-event interpolated image sequence generated at least in part from the first and second non-event images and one or more second color channels of an event-assisted interpolated image sequence generated at least in part on the event camera data.

[0268] EEE14. The method as recited in EEE13, wherein each neural field comprisesD24045W001HNeRV neural network blocks comprising MLP and CNN blocks and corresponding neural field operational parameters which are trained to model the respective pseudo image sequence.

[0269] EEE15. The method as recited in EEE 13 or EEE 14, wherein the one or more pseudo image sequences includes three pseudo image sequences each of which is generated by combining two color channels of the non-event interpolated image sequence and one color channel of the event-assisted interpolated image sequence.

[0270] EEE16. The method as recited in any of EEE13- EEE15, wherein the one or more pseudo image sequences includes two pseudo image sequences a first one of which is generated by combining two color channels of the non-event interpolated image sequence and one color channel of the event-assisted interpolated image sequence, and a second one of which is generated by combining one color channel of the non-event interpolated image sequence and two color channels of the event-assisted interpolated image sequence.

[0271] EEE17. The method as recited in any of EEE13- EEE16, wherein the one or more pseudo image sequences includes only one pseudo image sequence that is generated by combining three color channels of the non-event interpolated image sequence and three color channels of the event-assisted interpolated image sequence.

[0272] EEE18. The method as recited in any of EEE1- EEE17, wherein the event camera data is represented in one of: individual luminance change events forming an event sequence, an event frame in which events in a spatiotemporal neighborhood are accumulated as a whole, an event time surface representation, an event voxel grid representation, an event point cloud, an event neural field, or another event data representation.

[0273] EEE19. The method as recited in any of EEE1- EEE18, wherein at least one of the event camera data or the sequence of raw events is specified in one of: a proprietary event format defined by a provider of the event camera, an industry standard event format, a fixed length event data format, a DAT format, a Prophesee event data format, an IniVation event data format, a Hierarchical Data Format version 5, a Comma Separated Value (CVS) format, a compressed point cloud representation format, a compressed data format in accordance with one of H.264, H.265 or another video data coding standard, an event field using MPEG Neural Network Compression (NNC), a lossless event data format, or a lossy event data format.

[0274] EEE20. The method as recited in any of EEE1- EEE19, wherein the event camera data includes a specific portion derived from events occurring between the first and second image frame time points; wherein the specific portion of the event camera data is encoded into a single network abstract layer (NAL) unit.

[0275] EEE21. A method, comprising:D24045W001 decoding event related image data from an overall image container, the event related image data being derived and encoded into the overall image container by an upstream device from a combination of both event camera data and non-event camera data, the non-event camera data being derived by the upstream device from a first non-event image and a second non-event image, the first non-event image and the second non-event image being generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera, the event camera data being derived by the upstream device from a sequence of raw events, wherein the sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor, wherein the sequence of event time points is between the first time point and the second time point; generating at least one display image from the event related image data; rendering the at least one display image on an image display.

[0276] EEE22. The method as recited in EEE21 , wherein image container represents one of: an image signal that is exclusive of a base layer encoded with the first and second non-event images, or an image signal that includes a base layer encoded with the first and second non-event images.

[0277] EEE23. An apparatus performing any of the methods as recited in EEE1-EEE22.

[0278] EEE24. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any of the methods as recited in EEE1-EEE22.

Claims

D24045W001CLAIMS1. A method, comprising: deriving non-event camera data from a first non-event image and a second non-event image, wherein the first non-event image and the second non-event image are generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera; deriving event camera data from a sequence of raw events, wherein the sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, by an event sensor, wherein the sequence of event time points is between the first time point and the second time point; encoding event related image data into an overall image container, wherein the event related image data is derived from a combination of both the event camera data and the non-event camera data, wherein a recipient device of the overall image container is caused by the event related image data of the overall image container to render at least one display image on an image display, wherein the at least one display image is derived from the event related image data.

2. The method as recited in claim 1 , wherein the first non-event image and the second non- event image represent one of: two still images, or two successive non-event images in a sequence of non-event images forming a video.

3. The method as recited in claim 1 or 2, wherein the event related image data is encoded into one of: one or more network abstract layer (NAL) units, or one or more tracks in an ISO-BMFF file.

4. The method as recited in any one of claims 1-3, wherein the image container contains a base layer that carries a video stream encoded with non-event image including the first and second non-event images, wherein the image container contains a non-base layer that carries the event camera data.

5. The method as recited in any one of claims 1-4, wherein the image container contains a base layer that carries a video stream encoded with non-event images including the first and second non-event images, wherein the image container contains a non-base layer thatD24045W001 carries neural field operational parameters of a neural field that is trained with a residual image sequence generated as differences between a non-event interpolated image sequence and an event-assisted interpolated image sequence; wherein the non-event interpolated image sequence is generated by applying non-event video frame interpolation operations to the first and second non-event images; wherein the event-assisted interpolated image sequence is generated by applying event-based interpolation techniques based at least in part on the event camera data.

6. The method as recited in any one of claims 1-5, wherein the image container includes event metadata in the event related image data separate from event data in the event related image data.

7. The method as recited in any one of claims 1-6, wherein the image container is excluded from any layer encoded with non-event images.

8. The method as recited in any one of claims 1-7, wherein the image container carries neural field operational parameters of one or more neural fields that are trained with one or more pseudo image sequences, respectively; wherein each of the one or more pseudo image sequences for a respective neural field of the one or more neural fields is generated by combining one or more first color channels of a non-event interpolated image sequence generated at least in part from the first and second non-event images and one or more second color channels of an event-assisted interpolated image sequence generated at least in part on the event camera data.

9. The method as recited in any one of claims 1-8, wherein the event camera data is represented in one of: individual luminance change events forming an event sequence, an event frame in which events in a spatiotemporal neighborhood are accumulated as a whole, an event time surface representation, an event voxel grid representation, an event point cloud, an event neural field, or another event data representation.

10. The method as recited in any one of claims 1-9, wherein at least one of the event camera data or the sequence of raw events is specified in one of: a proprietary event format defined by a provider of the event camera, an industry standard event format, a fixed length event data format, a DAT format, a Prophesee event data format, an IniVationD24045W001 event data format, a Hierarchical Data Format version 5, a Comma Separated Value (CVS) format, a compressed point cloud representation format, a compressed data format in accordance with one of H.264, H.265 or another video data coding standard, an event field using MPEG Neural Network Compression (NNC), a lossless event data format, or a lossy event data format.

11. The method as recited in any one of claims 1-10, wherein the event camera data includes a specific portion derived from events occurring between the first and second image frame time points; wherein the specific portion of the event camera data is encoded into a single network abstract layer (NAL) unit.

12. A method, comprising: decoding event related image data from an overall image container, the event related image data being derived and encoded into the overall image container by an upstream device from a combination of both event camera data and non-event camera data, the non-event camera data being derived by the upstream device from a first non-event image and a second non-event image, the first non-event image and the second non-event image being generated at least in part from a scene for a first image frame time point and a second image frame time point, respectively, by a non-event camera, the event camera data being derived by the upstream device from a sequence of raw events, wherein the sequence of raw events is generated in response to luminance changes of the same scene at a sequence of event time points, respectively, hy an event sensor, wherein the sequence of event time points is between the first time point and the second time point; generating at least one display image from the event related image data; rendering the at least one display image on an image display.

13. The method as recited in Claim 12, wherein image container represents one of: an image signal that is exclusive of a base layer encoded with the first and second non-event images, or an image signal that includes a base layer encoded with the first and second non-event images.D24045W00114. An apparatus performing any of the methods as recited in claims 1-13.

15. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any of the methods as recited in claims 1-13.

Citation Information

Patent Citations

  • Multi-mode video frame insertion method based on event camera reconstruction reference

    CN118590665A

  • Method, device, and computer program for encapsulating hevc layered media data

    GB2535453A

  • US202363611975P