Coding event camera data using neural field

Neural fields, implemented with MLP networks, effectively address the challenges of processing and compressing event camera data by modeling and transforming event data, resulting in high rate-distortion performance and efficient compression.

WO2025136863A1PCT designated stage expired Publication Date: 2025-06-26DOLBY LABORATORIES LICENSING CORP

Patent Information

Application Number
PCT/US2024/060311
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-24
Filing Date
2024-12-16
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing technologies face challenges in efficiently processing and compressing event camera data, which is generated by sensors responding to local luminance changes without using a shutter, leading to asynchronous and asynchronous event streams.

Method used

The use of neural fields, implemented with multi-layer perceptron (MLP) networks, to model and transform event camera data. This approach involves encoding event data with positional information, training neural fields with ground truth event polarities, and using the trained neural fields to predict and reconstruct events, thereby enabling efficient compression and representation of event camera data.

Benefits of technology

The neural field approach achieves high rate-distortion performance and effective compression of event camera data, allowing for accurate reconstruction of events and improved handling of asynchronous event streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024060311_26062025_PF_FP_ABST
    Figure US2024060311_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Event camera data containing raw events is received. Neural field training data is generated from the raw events in the event camera data. In a training phase, optimized values for operational parameters of a neural field are generated by training the neural field with the neural field training data based on a specific loss function. Neural field operational parameter values are encoded in a coded bitstream to enable a recipient device to use the neural field operating with the neural field operational parameter values to generate reconstructed events.
Need to check novelty before this filing date? Find Prior Art

Description

CODING EVENT CAMERA DATA USING NEURAL FIELD CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 611,975, filed December 19, 2023, and European Patent Application No.24159531.3, filed February 24, 2024, each of which is incorporated by reference herein in its entirety. TECHNOLOGY

[0002] The present disclosure relates generally to visual computing and more particularly to coding event camera data using neural field. BACKGROUND

[0003] An event camera with an imaging sensor that responds to local luminance changes can be used to generate event camera data from a physical environment or visual scene. For example, the event camera data may be generated by an array or spatial distribution of pixels or sensor elements of the image sensor inside the event camera without using a shutter. Each pixel or sensor element in the image sensor can operate independently and asynchronously with other pixels or sensor elements in the same image sensor. The event camera data may include a stream of asynchronously occurring pixel-specific events in a plurality of time points covering a time interval or duration. Each such event indicates a polarity (increase or decrease) or an instantaneous measurement in luminance or brightness at a specific spatial location or direction – in the physical environment or visual scene – corresponding to a specific pixel at a specific time point in the time points represented in the event camera data.

[0004] Techniques are being developed to process event camera data generated using different types of event cameras. For example, temporal smoothing may be applied to the event camera data to construct images of physical environments or visual scenes with a relatively wide luminance or brightness range or with a relatively dark luminance or brightness. The event camera data may also be processed or filtered for the purpose of recognizing stationary or moving objects or pedestrians in the physical environments or visual scenes relatively efficiently and responsively.

[0005] The approaches described in this section are approaches that could be pursued, but not necessarily approaches that have been previously conceived or pursued. Therefore, unless otherwise indicated, it should not be assumed that any of the approaches described in this section qualify as prior art merely by virtue of their inclusion in this section. Similarly, issues identified with respect to one or more approaches should not assume to have been recognized in any prior art on the basis of this section, unless otherwise indicated.BRIEF DESCRIPTION OF DRAWINGS

[0006] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings and in which like reference numerals refer to similar elements and in which:

[0007] FIG.1A illustrates an example multi-layer perceptron (MLP) network that may be used to implement a neural field; FIG.1B illustrates example basic building blocks of an MLP layer in the MLP network; FIG.1C and FIG.1D illustrate example neural field encoder and decoder structure;

[0008] FIG.2A illustrates an example image containing aggregated events; FIG.2B through FIG.2D illustrate example reconstructed event frames with prediction accuracies; FIG.2E illustrates example modeling of raw events by a neural field with predicted events;

[0009] FIG.3A through FIG.3F illustrate example MLP networks with different numbers of neurons in MLP layers therein;

[0010] FIG.4A and FIG.4B illustrate example process flows; and

[0011] FIG.5 illustrates an example hardware platform on which a computer or a computing device as described herein may be implemented.DESCRIPTION OF EXAMPLE EMBODIMENTS

[0012] Example embodiments, which relate to coding or modeling event camera data using neural field, are described herein. In the following description, for the purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. It will be apparent, however, that the present disclosure may be practiced without these specific details. In other instances, well-known structures and devices are not described in exhaustive detail, in order to avoid unnecessarily occluding, obscuring, or obfuscating the present disclosure.

[0013] Example embodiments are described herein according to the following outline: 1. GENERAL OVERVIEW 2. NEURAL FIELD AND MLP NETWORK 3. POSITIONAL ENCODING 4. EVENT DATA 5. NEURAL FIELD CODING FRAMEWORK 6. MODELING AGGREGATED EVENTS WITH NEURAL FIELD 7. RECONSTRUCTING AGGREGATED EVENTS 8. MODELING RAW EVENTS 9. MODELING RAW EVENT DATA WITH NON-EVENTS 10. COMPRESSING RAW EVENT CAMERA DATA USING NEURAL FIELDS 11. EXAMPLE PROCESS FLOWS 12. IMPLEMENTATION MECHANISMS – HARDWARE OVERVIEW 13. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS 1. GENERAL OVERVIEW

[0014] This overview presents a basic description of some aspects of an example embodiment of the present invention. It should be noted that this overview is not an extensive or exhaustive summary of aspects of the example embodiment. Moreover, it should be noted that this overview is not intended to be understood as identifying any particularly significantaspects or elements of the example embodiment, nor as delineating any scope of the example embodiment in particular, nor the disclosure in general. This overview merely presents some concepts that relate to the example embodiment in a condensed and simplified format, and should be understood as merely a conceptual prelude to a more detailed description of example embodiments that follows below.

[0015] Neural field can be used as coordinate based neural network that parameterizes physical properties of scenes or objects across space and time. Neural field has applications in visual computing problems such as 3D image reconstruction, image synthesis, etc. In a neural field framework, field quantities are produced by sampling coordinates and fed to a neural network. Field quantities may represent samples in a target (or desired) reconstruction domain.

[0016] For example, a neural field such as a neural radiance field (NeRF) may be used to provide a 3D scene representation. The NeRF can take a specific input from or at a coordinate such as a specific spatial location (x, y, z) and / or a specific viewing direction (φ, ∅) and uses the specific input to generate a specific output such as a volume density and / or a view dependent emitted radiance at or of that coordinate. The output volume density and RGB color corresponding to the view dependent emitted radiance are followed by volume rendering operations to construct a 2D projected image of the 3D scene. Hence, a neural field may be used as a local implicit image function to represent an image as a set of latent code to predict RGB values of a given pixel at a given coordinate or spatial location (x, y) and / or 2D features.

[0017] Techniques as described herein can be implemented to provide a novel framework using neural fields (e.g., as implicit functions, etc.) to model or transform one or more sequences of event camera data originated or generated from one or more event cameras at one or more views in relation to a physical environment or visual scene. Each of the sequence of event camera data may comprise sensory data collected from sensor elements of a respective event camera among the one or more event cameras.

[0018] Under some approaches, neural networks could function as a filter to accept or take one or two dimensional (1D / 2D) sensor captured data as input. The neural network can then generate filtered results as output.

[0019] In contrast, a neural field as described herein accepts or takes coordinate or spatial location information relating to a signal (e.g., a sequence of event camera data generated by an event camera, etc.) as input. Through training with sensor captured data such as polarities in a sequence of events as ground truths, the neural field “memorizes” the signal or the sequence of events itself with the coordinate or spatial location information. The neural field can then be used to predict or generate a sequence of reconstructed events approximating the sequence of events used as the input and ground truths for the neural field.

[0020] A neural field may be referred to as an implicit neural function containing hidden layers operating with weights, biases, operational parameters, activation functions. The trained neural field or the fitted implicit neural function represented therein maps the input sequence of events to the sequence of reconstructed events.

[0021] Under techniques as described herein, neural fields may be used to implement (e.g., backbone, non-backbone, etc.) multimedia encoders or codecs that encode and / or represent event camera data by way of trained or fitted operational parameters used in the neural fields. These operational parameters can be delivered from an upstream device to a downstream recipient device. The downstream device can use the trained or fitted operational parameters to operate corresponding neural fields to predict or generate reconstructed events relatively accurately approximating events represented in or derived from the event camera data. The neural field multimedia codecs as described herein can help achieve relatively high rate- distortion (R-D) performance as well as relatively high compression performance in coding operations.

[0022] In some example operational scenarios, a neural field as described herein may be implemented with a multi-layer perceptron (MLP) network. The MLP network may be relativelyefficiently trained or fitted through machine learning based training to describe, encode and / or represent multimedia data such as trained or fitted operational parameters of the MLP network or layers therein, which can then be used to predict or generate reconstructed event camera data corresponding to event camera data collected with one or more event cameras.

[0023] In comparison with a deep learning architecture which takes multimedia raw data as input and perform filtering like operations, the MLP network can take coordinate or spatial location information relating to or embedded with the event camera data as input and memorizes or preserves the coordinate or spatial location information in output – reconstructed events – generated from the MLP network.

[0024] Under some approaches, an MLP represents a regressor or an implicit continuous function outputting predicted output values covering all contiguous values (e.g., contiguous luminance or RGB valued, etc.) of a value range or space such as all byte values, all integer values, etc.

[0025] In contrast, an MLP network as described herein operates or works as a classifier (e.g., to predict a positive or negative polarity in luminance change, etc.), rather than as a regressor. While other layers of the MLP network may output contiguous values, the last layer – or output layer – of the MLP network uses an activation function that outputs a limited number (e.g., 2, 3, etc.) of discrete values representing a proper subset of all contiguous values of a value range or space or selected values in all byte values, all integer values, etc.

[0026] In some operational scenarios, a neural field or MLP network as described herein may be used to model a sequence of aggregated event camera data generated from aggregating a sequence of raw or pre-aggregated event camera data collected from an event camera. The aggregated events may be generated or aggregated from raw events in the event camera data based at least in part on (e.g., timestamp-based, etc.) user input specifying time intervals used for aggregation. Additionally, optionally or alternatively, the neural field can model multiplesequences of event camera data based on multiple views of multiple event camera or camera elements in relation to a physical environment or visual scene.

[0027] In some operational scenarios, a neural field or MLP network as described herein may be used to model one or more sequences of raw or pre-aggregated event camera data collected by one or more event cameras from a physical environment or visual scene. Additionally, optionally or alternatively, the neural field can model non-events data corresponding to coordinates or spatial locations not represented by the raw events in the raw event camera data.

[0028] The neural field – along with trained or fitted operational parameters including but not limited to weights, biases, activation function parameters, etc. – may be deemed as an implicit function that logically approximates or provides a representation of aggregated or raw event camera data used as input to train or fit (e.g., generate optimized values for, etc.) the neural field or the operational parameters. Under techniques as described herein, compression of the aggregated or raw event camera data may be implemented or accomplished by way of compression of the neural field or the operational parameters defining or specifying the neural field.

[0029] In some operational scenarios, event locational data used to train the neural field may be coded into a bitstream with the optimized values of the operational parameters.

[0030] Additionally, optionally or alternatively, while the optimized values for the operational parameters may be represented in a relatively high precision numeric format such as floating point, double, etc., the optimized values for the operational parameters may be quantized into a relatively low precision numeric format (in comparison with the high precision format) such as (e.g., linearly quantized, non-linearly quantized, etc.) integers. The quantized optimized values for the operational parameters of the neural field may be coded into the bitstream instead of the pre-quantized optimized values used in training.

[0031] Example embodiments described herein relate to encoding event camera data by way of neural field operational parameters. Event camera data having a sequence of raw events is received. Each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element in a plurality of sensor elements of an event image sensor. Each sensor element in the plurality of sensor elements of the event image sensor corresponds to a respective pixel position in a plurality of pixel locations. Neural field training data is generated from the sequence of raw events in the event camera data. The neural field training data includes a plurality of training instances. Each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance. In a training phase, optimized values for operational parameters of a neural field are generated by training the neural field with the neural field training data based at least in part on a specific loss function selected for the neural field. The neural field is trained to output respective predicted classification categories for pixel locations in the plurality of pixel locations at a plurality of time points. Neural field operational parameter values are encoded in a coded bitstream. The neural field operational parameter values are generated from the optimized values for the operational parameters of the neural field. The coded bitstream causes a recipient device to use the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data.

[0032] Example embodiments described herein relate to decoding event camera data by way of neural field operational parameters. Neural field operational parameter values are decoded from a coded bitstream generated by an upstream device. The neural field operational parameter values are generated from optimized values for operational parameters of a neural field. The optimized values for the operational parameters of the neural field were generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field. The neural field is trained to outputpredicted classification categories. The neural field training data includes a plurality of training instances. Each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance. The neural field training data is generated from a sequence of raw events in event camera data. Each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element in a plurality of sensor elements of an event image sensor. Each sensor element in the plurality of sensor elements of the event image sensor corresponds to a respective pixel position in a plurality of pixel locations. The neural field operating with the neural field operational parameter values is used to generate a plurality of reconstructed events approximating a plurality of events represented in the training data.

[0033] In some example embodiments, mechanisms as described herein form a part of a media processing system, including but not limited to any of: a handheld device, game machine, television, laptop computer, netbook computer, tablet computer, desktop computer, computer workstation, computer kiosk, or various other kinds of computing devices and media processing units.

[0034] Various modifications to the embodiments and the generic principles and features described herein will be readily apparent to those skilled in the art. Thus, the disclosure is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein. 2. NEURAL FIELD AND MLP NETWORK

[0035] FIG.1A illustrates an example multi-layer perceptron (MLP) network that may be used to implement a neural field as described herein. The MLP network may include MLP layers such as an input layer, an output layer and one or more hidden layers between the input layer and the output layer. These MLP layers are fully connected – each neuron in a subsequent layer among the MLP layers is connected with all neurons of an immediately preceding layer (ifexists) among the MLP layers and hence receives outputs generated from all the neurons of the immediately preceding layer.

[0036] Activation functions are used in neurons of every hidden layer in the MLP network. An activation function of a neuron of a given layer of the MLP network receives from outputs of all neurons of an immediately preceding layer (relative to the given layer) as input and generates an output.

[0037] Instead of acting as a regressor or an implicit function generating continuous output values covering all contiguous values of a specific value range or value space, the MLP network as described herein acts as a classifier generating sparse or a proper subset – as opposed to all contiguous values – of a specific value range or value space such as all byte values, all integer values, etc. In some operational scenarios, some or all of the input layer or the hidden layers before the output layer in the MLP network may use activation functions that generate continuous output values. In comparison, the last layer – or the output layer – in the MLP network as describe herein uses an activation function that outputs discrete sparse values (not covering all contiguous values of a value range or space or selected values in all byte values, all integer values, etc.).

[0038] As illustrated in FIG.1A, the input layer of the MLP network comprises neurons (represented as circles) receiving input to the neural field or MLP network such as input0, input 1, etc., and generating intermediate output to be received as input by the one or more hidden layers or neurons therein. The output layer of the MLP network comprises neurons (represented as circles) receiving neuron outputs from the immediately preceding layer (relative to the output layer) as input and generating predictions (or MLP predicted values) as output from the neural field or MLP network such as output0.

[0039] Different layers of a MLP network as described herein can have different total numbers of neurons. A deep or multi-layer MLP network that includes multiple MLP layers canbe constructed by stacking or arranging multiple building blocks of the multiple MLP layers therein sequentially.

[0040] FIG.1B illustrates example basic building blocks involving a neuron of an MLP layer in the MLP network. Each of the other neurons in the same MLP layer and / or neurons of other MLP layer(s) in the MLP network may have the same or similar structure.

[0041] As shown, the neuron in the MLP layer receives multiple inputs (fully connected), which are in turn generated as outputs from all neurons of an immediately preceding layer and generates an output. The neuron may include one linear operation followed by a rectifier linear unit (ReLU) activation function.

[0042] The linear operation may receive inputs x1, x2, x3, … xm, where m represents an integer no less than one (1) representing the total number of neurons of the immediately preceding layer. The linear operation multiplies these inputs with respective weights (or weight factors) w1, w2, w3, … wm – which may be collectively denoted as ^ – to generate a weighted sum. The linear operation further adds to the weighted sum a bias (or bias factor) ^ to generate or output an output value x of the linear operation. The ReLU activation function f(x) (e.g., max (0, x), etc,) receives the output value x from the linear operation as input and generates an output value ^^based on the input x received from the preceding linear operation.

[0043] More generally, denote operational parameters such as weights and bias at (e.g., each neuron of, etc.) the ^^-th MLP layer – among ^^MLP layers forming the MLP network or neural field, wherein ^^represents an integer no less than one (1) – as weights (^^^) and bias (^^^),which may be collectively denoted as ^ = {{^^^ ^, {^^^ ^^. For simplicity, indexes for neurons areomitted, but it should be noted thatand biases may be used for different neurons of the same MLP layer.

[0044] In a model training phase, these operational parameter can be trained or fitted to optimize or minimize prediction errors as measured or computed with a selected loss function for the MLP (or neural) network implementing the neural field. An output value from the ^^-thMLP layer or an output neuron therein can be generated or adjusted by an activation function such as a Rectified Linear Unit (ReLU) – it should be noted that, for the output layer, a different activation function such as Softmax function may be used to generate discrete sparse values representing outputs of the overall MLP network as a classifier. The ReLU activation function may be given as follows: ^^^^^^^ = max ^0, ^^ (1)

[0045] In some operational scenarios, the final MLP network (or neural field) output value may be represented or limited in a specific value range such as [0, 1]. A sigmoid layer can then be added as the final layer of the MLP network which squishes every (input) real number from the immediately preceding layer (e.g., output from the last ReLU activation, etc.) in the MLP network to in the specific range between 0 and 1. An activation function for the sigmoid layer may be given as a sigmoid function as follows: ^^^^ = ^^^^ ! (2)

[0046] consisting of ^^MLP layers with respective operational parameters weights {^^^^ and bias {^^^^or ^ = {{^^^ ^, {^^^ ^^. This MLP network takes input x and outputs ^^ as^^ = "^#$^%^ (3)

[0047] Given a ground truth signal ^&'(e.g., based on samples in the event camera data, etc.), the formal problem formulation for fitting or training the MLP network or neural field – e.g., in the training phase – may be defined or specified as optimizing values for the operationalparameters ^ to minimize the loss function (^, ^ as follows:^∗ = *+, min (^ "^# ^% &'$ $ ^, ^ ^ (4)

[0048] locations at which sensor elements are disposed, etc.) in the training data may be used as input to the MLP or neural fields to generate predictions of polarities for different positions. The (single) loss function may be used to measure prediction errors between the predicted polarities andground truth polarities in the training data. Given that activation functions in the MLP network or neural field are continuous and differentiable, gradients or differential values may be computed based at least in part on the prediction errors and back propagated to iteratively adjust weights and biases until the prediction errors as measured by the loss function is minimized, for example, below a prediction error threshold and / or until all training instances of the training data are processed. A variety of candidate loss functions may be considered for being selected as a loss function to train the MLP network or neural field. In some operational scenarios, a loss function may be specifically selected from candidate loss functions or used to measure accuracies or to compute errors of predicted event polarities and / or non-events in reference to sensor-generated event polarities and / or non-events implied by the events. In some operational scenarios, a task-specific (e.g., application-specific, task-oriented, application-oriented, etc.,) loss function may be selected from candidate loss functions or used to support a specific type of downstream applications or tasks of the reconstructed events. Different (types of) loss functions may be used as task-specific loss functions for different (types of) downstream applications or tasks. Example (types of) downstream applications or tasks as described herein may include image / video generation, object recognition, and so on. By way of illustration but not limitation, in some operational scenarios in which the events are used to generate or reconstruct images or videos at a relatively high frame rate, the reconstructed images (or image frames) from the reconstructed events may be used in a task-specific loss function that emphasizes on overall quality of the images rather than emphasizing on prediction accuracies of individual events. Likewise, in some operational scenarios in which the events are used to recognize objects depicted in the events at a relatively low latency, object recognition and classification accuracy may be measured with a task-specific loss function to emphasize the overall quality of the object recognition rather than emphasizing on prediction accuracies of individual events.

[0049] In a model application or inference phase, positional data representing different positions may be used as input to the neural network operating with optimized operational parameter values to generate or infer predicted polarities for these positions.

[0050] These training / fitting and application / inference phases may be repeatedly executed with different sequences of events and / or with different event camera data sets covering different time period (e.g., different hours, days, weeks, etc.). 3. POSITIONAL ENCODING

[0051] Downstream or subsequent media or image rendering operations based at least in part on outputs from a neural field or MLP network used to process event camera data. Positional encoding and / or sinusoid activation function (SIREN) may be implemented to reduce or prevent a loss of high frequency details in these media or image rendering operations and thereby improve scene representation by the neural field or MLP network.

[0052] In positional encoding, inputs to the neural field or MLP network are mapped to a relatively high dimensional space that include dimensions representing relatively high spatial frequency components in coordinate or spatial locational / positional data. This can be used to correct biases of neural networks towards learning relatively low frequency functions or components represented in the event camera data and to enable the neural field or MLP network to represent or learn relatively high (e.g., coordinate, spatial locational / positional, etc.) frequency variation in color and geometry / coordinate / location / position.

[0053] Hence, the performance of a neural field or MLP network may be significantly improved by mapping position coordinates / 0 from a one-dimensional real value domain R to 2L- dimensional real value domain ^12, where L represents the total number of frequencies, for example generated by the positional encoding (or coordinate mapping) 3 implemented with sinusoid activation functions (SIRENs). The positional encoding (or coordinate mapping) 3 acting on coordinate / 0 can be represented as follows: 3^ / 0^ = 4sin^2789 / 0^ cos^2789 / 0^ sin^27<9 / 0^ cos^27<9 / 0^ … sin^27> <9 / 0^ cos^27> <9 / 0^ ?(5)where {@A, @^, … , @2B^^ are integers. In an example setting, @C = D.4. EVENT DATA

[0054] Event cameras are asynchronous (physical) sensors that have a different representation of a scene as compared with other types of image sensors such as CCD. The event cameras can generate event camera data with a relatively high sensitivity to low light as well as light changes in a physical environment or visual scene, a relatively high temporal resolution, a relatively low latency (in the order of microseconds), a relatively (e.g., very, etc.) high dynamic range (e.g., a range of 140dB versus a range of 60dB of a camera with a different type of image sensor, etc.), and a relatively low power consumption.

[0055] A non-event camera may acquire full images at an image acquisition or refresh rate specified by or in reference to an external clock (e.g., 30fps, etc.). In contrast, an event camera or an event sensor therein such as a dynamic vision sensor (DVS) respond to luminance or brightness changes in a physical environment or visual scene instantly, asynchronously and independently for every pixel or sensor element in the event camera or sensor. A per-pixel reference log intensity (e.g., of received photons or light, etc.) may be maintained for each pixel or sensor element in the event camera or sensor. A pixel or sensor element fires or outputs a binary output / event if the magnitude of the log intensity changes beyond a maximum (log) intensity change threshold for positive or negative change. Mathematically, this event firing process can be expressed as follows: 1, H^%^, E^^ > JK(6)whereH^%^, E^^ = log^I^%^, E^^^ − log ^I^%^, E^ − WE^^^ (7)where I^%^, E^^ represents the intensity (e.g., of photons or light, etc.) at the pixel location %^=^^^, X^^, and time E^. WE^ represents the time elapsed since the last event at the same pixellocation or sensor element. I^%^, E^ − WE^^ represents the intensity (e.g., of photons or light,etc.) at the pixel location %^ = ^^^ , X^^. Also, JK represents the maximum (log) intensity changethreshold for the positive value change and JNrepresent the maximum (log) intensity changethreshold for the magnitude of the negative value change. Here, ^^%^, E^^ outputs or representsthe polarity of the event data or (log) intensity value change at the given pixel location %^andthe given time point E^ , which may be of one of the values 1, -1 and 0. If ^^%^, E^^ = 0, then noevents are generated. Regardless of whether an event is generated for a pixel position or sensor element at a given time point, the polarity of the pixel position or sensor element at the giventime point may be denoted as Y^ for simplicity. Events with Y^ = 0 for pixel positions or sensorelements may not be stored or explicitly represented in the event camera data or a stream thereof and may be implicitly assumed in the absence of positive or negative events for these pixel positions or sensor elements. Here k represents an event index, for example, used to identify or order events belonging to the same sequence of events. Different indexes k may refer to multiple events occurred at the same time or at different times.

[0056] An event ^^^E^– if fired – at a pixel location or sensor element may be defined as a4-parameter tuple ^^^E^ = ^^^, X^, E^, Y^^ containing coordinate information about its pixellocation ^^^, X^^, the exact trigger time E^, and the 1-bit polarity Y^. In some variants of eventcameras / sensors, instead of merely denoting whether a pixel at which an event is fired has increased / decreased in brightness or luminance, the event also contains information about or specifies how much intensity has changed. Other pixel locations are inferred to have the same intensity values as before and are thus not represented in the event camera data or the stream thereof.

[0057] Note that the event camera data or stream may be composed of sparse points or pixels as it may only explicitly represent – or contain information – these sparse points or pixels at each of which a sufficiently large intensity change as compared with a maximum intensity change threshold is detected or sensed by a corresponding sensor element.

[0058] The event camera data may be generated by an event camera and represented ordefined by a sequence of events (denoted as Z[E\ , EC]) between an initial time (point) E\ and afinal time (point) EC as Z[E\ , EC] = 〈^^E^|E,A time point E within the interval betweenE\and ECmay also be referred to as a time instance. Similarly, the sequence of events may also be referred to as an event stream.

[0059] Due to the advantages of event camera / sensors, event camera data generated by these cameras / sensors and modeled / transformed with neural fields or MLP networks under techniques as described herein can be used in various applications including but not limited to video gaming, AR / VR / MR, object and / or image rendering, applications with real-time interactions such as in robotics, applications in which relatively low power, low latency and robust performance under varying lighting conditions are needed, surveillance, object segmentation, object detection, object tracking, 2D / 3D sensing of physical environments, gesture recognition, object recognition, optical flow estimation, HDR image reconstruction, robotics related applications such as SLAM (Simultaneous Location and Mapping), image reconstruction, image deblurring (under relatively extreme lighting conditions and / or with relatively fast motions), astronomical research, and so forth. 5. NEURAL FIELD CODING FRAMEWORK

[0060] A neural field coding framework may be implemented to model or transform event camera data using neural fields. The framework includes an encoder side architecture as illustrated in FIG.1C and implemented by an upstream device that generates optimized values for operational parameters of a neural field or MLP network using event camera data as input, compresses or encodes the optimized values into a coded bitstream, and transmits or otherwise delivers the coded bitstream to downstream recipient device(s). The framework further includes a decoder side architecture as illustrated in FIG.1D and implemented by a downstream recipient device that receives the coded bitstream directly or indirectly from the upstream device, decodes or decompresses the optimized values of the operational parameters for the neural field or MLP network from the coded bitstream, applies the optimized values of the operational parameters to generate outputs from the neural field or MLP network. These outputs may be used in a widevariety of applications including generating or rendering images, object segmentation, object detection, object tracking, AR / VR / MR, real time or non-real time applications, etc.

[0061] As illustrated in FIG.1C, the encoder side architecture receives event camera data as input and uses a neural field or MLP network to represent the event camera data after training or fitting the neural field or MLP network with the event camera data.

[0062] In some operational scenarios, the event camera data used to train or fit the neural field or MLP network may be represented or specified as a sequence of (aggregated) event camera images at a sequence of time points covering a time interval or duration, such as a video or a video clip.

[0063] The sequence of event camera images may include event camera images of a single view or originating from an event camera of a single view at or for different time points in the sequence of time points. The sequence of event camera images may include a single image of the single view at a corresponding time point in the sequence of time points.

[0064] Additionally, optionally or alternatively, the sequence of event camera images may include event camera images of multiple views or originating from multiple event cameras (or camera elements) of multiple views at or for different time points in the sequence of time points. The sequence of event camera images may include multiple images of the multiple views at a corresponding time point in the sequence of time points.

[0065] In some operational scenarios, the event camera data used to train or fit the neural field or MLP network may be represented or specified as a sequence of (raw) event camera data at a sequence of time points covering a time interval or duration. For the purpose of better modeling the event camera data, raw non-event data not necessarily explicitly represented in the sequence of raw event camera data may be generated or provided as input to train or fit the neural field or MLP network.

[0066] The neural field or MLP network once trained or fitted with the input event camera data provides a representation (or neural field encoding) of the event camera data by way ofoptimized values generated for weights, biases and / or other operational parameters of the neural field or MLP network. These optimized values of the operational parameters of the neural field or MLP network may be encoded or compressed (with neural network compression) into a coded bitstream (or neural network bitstream), which may be transmitted or otherwise delivered from an upstream device implementing the encoder side architecture to a downstream recipient device implementing the decoder side architecture.

[0067] As illustrated in FIG.1D, the decoder side architecture receives the coded bitstream (or neural network stream) as input, and decompress or decodes the bitstream (with neural network decompression) to retrieve or generate the optimized values of the operational parameters (or neural field coefficients) such as the weights, biases, etc. A designated user may specify user input including but not limited to those pertaining to the event camera data such as timestamps, views and so on. Based at least in part on the user input, the decoder side architecture performs neural field decoding to generate a reconstructed version of the event camera data (or reconstructed event data) that is the same as, or relatively closely approximate, the input event camera data on the encoder side architecture. The neural field at the decoder may be implemented with an MLP which can be used to reconstruct the event camera data. 6. MODELING AGGREGATED EVENTS WITH NEURAL FIELD

[0068] As mentioned earlier, an event ^^^E^ as represented or included in event camera datamay be defined as a 4-parameter tuple ^^^E^ = ^^^, X^ , E^, Y^^. To help preserve high spatialfrequency event or image features represented in the event camera data, positional encoding 3may be applied to coordinates ^^, X^ in events represented or included in the event camera data.

[0069] Given these coordinates and a given timestamp bE, the neural field may be implemented with a MLP network (denoted as "^#c) trained or fitted with the events in the event camera data to predict or output polarities for coordinates and the same given timestamp bE in a reconstruction domain (or a reconstructed event camera data domain). A predictedpolarity (of an event) for coordinate ^^, X^ at the timestamp bE may be given or represented asfollows: Ŷ = "^#c$^3^^^, 3^X^, bE^ (8)

[0070] Under some approach, an MLP network that acts as a regressor (e.g., predicting all possible values in a value range, predicting all possible color values in a color, etc.) may be used to implement a neural field. In comparison, under techniques as described herein, an MLP network or neural field is modified or implemented to act as a classifier (e.g., two or more classification categories, polarity in luminance or brightness changes, etc.). The MLP network or neural field can be used to predict the polarity of an event represented in the event camera data.

[0071] In some operational scenarios the neural field or MLP network is implemented as a multi-class classifier. This may be achieved by replacing the sigmoid activation function in the neural field block or MLP layer by a Softmax function. The Softmax function can be represented as follows: efg^hi^^Y+\ = ∑m^ln< efg^hl^^ (9)where o\and pq is the number of classes.

[0072] The Softmax function takes an pq dimensional vector of classes from the previouslayers and transforms each component of the vector into a real number in a value range 40, 1?,such that all components of the vector add up to one (1).

[0073] The MLP network "^#′ as modified or implemented – for example, using the Softmax activation function instead of the sigmoid function as activation function in an (e.g., each, etc.) MLP layer therein – as a classifier takes input % represented in an input event camera data domain and outputs classification categories Ŷ in a reconstructed event camera data domain. Specific representations of the input % depend on specific use cases. The output classification categories Ŷ from the MLP network "^#′ may be represented as follows: Ŷ = "^#c$^%^ (10)

[0074] Given ground truth polarities ^&'generated or represented in the input event camera data, the formal problem formulation to obtain optimized values of operational parameters such as weights, biases, etc., of the (modified) MLP network may be expressed as follows: ^∗ = *+, m c &'$in (^ "^# $^%^, ^ ^ (11)where (^, ^ represents the loss function to be minimized for the purpose of obtaining theoptimized values of the operational parameters. As the MLP is to be used as a classifier, instead of using the MSE loss function for a regressor, a cross-entropy loss function may be used in expression (11) and represented as follows: (= − ∑Nq\t^ Y0\. log^Y+\^ (12)where Y+\is the output of the Softmax function for the i-th component of the vector received from the previous (MLP) layers; Y0\is the mapped target polarity for the i-th component; and pq is the number of classes or components in the vector. For simplicity in event camera datamodeling, polarities Y from the raw event camera data are mapped to Y0 such that Y = +1corresponds to Y0 = 0 and Y = −1 corresponds to Y0 = 1.7. RECONSTRUCTING AGGREGATED EVENTS

[0075] Techniques as described herein may be used to implement deep learning algorithms that can be applied to aggregated events in synchronous frames. These aggregated events may be generated from asynchronous events included or represented in raw event camera data produced with one or more event cameras.

[0076] Given the event camera data as an event dataset, as mentioned earlier, an eventsequence for the entire dataset may be given or specified as: Z[E\ , EC] = 〈^^E^|E ∈ 4E\ , EC?〉. Toaggregate the events in the event sequence or dataset,event pixel location, which refers to a pixel or pixel location at which at least one event occurs during the time interval T, the latest or last event is selected or picked from all the event(s) occurred during the time interval T at the event pixel location. The selected event may be referred to as anaggregated event denoted as Zw&&_^y^N'[E\, EC], which may be given or represented as follows:Zw&&z{zm|[E\ , EC] = 〈^^%^, E^̂|Ê ∈ }E\, EC~, Ê = max^E^^ ∀^%^^〉 (13)whereof the selected event at the latest timeinterval Ê and event pixel location %^the latter of which corresponds to the pixel coordinate^^^, X^^.

[0077] Pixel locations that are not event pixel locations or that do not belong to any event pixel location in the aggregated event set in expression (13) above may be assigned the 3rd polarity – a value of 0.

[0078] Assume that an event / image frame that includes all pixel locations at which sensor elements are to detect or fire events of polarities has a width of ^ and height of ^. A set of non- event data for pixel locations in the event / image frame at which sensor elements do not detect or fire events of polarities may be determined or identified, as follows: ^Zw&&z{zm|[E\, EC] = ^0|^^q^, Xq^^ ∉ ^^^, X^^, ^q^ ∈ 41, ^?, Xq^ ∈ 41, ^? ^ (14)where interval; thepolarity is non- lie withinthe boundary of the event / image frame and do not correspond or belong to any event pixel location within that given time interval. The given time interval may be represented by a reference timestamp of a reference time point in the given time interval.

[0079] The final or combined dataset – which contains event or polarity information for all pixel locations of the event / image frame – used as input to train the neural field or MLP network on the aggregated events / contents may be represented or given as follows: Zw&&[E\, EC] = Zw&&z{zm|[E\, EC] ∪ ^Zw&&z{zm|[E\, EC] (15)

[0080] can be used to model aggregated events using last occurred events captured in event camera data for corresponding event pixel locations. It should be noted, however, that in other operational scenarios a neural field can also be used to model event camera data that has been aggregated using other techniques such as Time Surface (TS), Event Voxel, and so on. Under the Time Surface techniques, an event stream or events therein may be converted into a spatiotemporal point cloud. A time surface may be generated for a (e.g., each, etc.) event in the point cloud by applying an exponential decay kernel (or function) to a pixel neighborhood of a specific pixel location to which the event correspond or is located. Hence, a time surface may be used to account for the event as well as a dynamic spatiotemporal context around the event that providesinformation about the history of the activity in the neighborhood. Time surfaces generated from the events provides a Time Surface representation for the events and can be used as inputs to a neural field as described herein to generate predicted events. Under the Event Voxel techniques, an event stream may be converted into a fixed-size tensor representation in which events may be encoded in a spatiotemporal voxel grid. A time duration ∆T = tkk N −1 − t 0 spanned or covered by the events may be discretized into B temporal bins. Every event in the events can distribute its polarity to its two closest spatiotemporal voxels in the spatiotemporal voxel grid. The voxels with the events distributed therein provides an event voxel representation for the events and can be used as inputs to a neural field as described herein to generate predicted events. Example event aggregations with Time Surface and Event Voxel are described in Yiheng Xie et al., “Neural Fields in Visual Computing and Beyond,” Eurographics / CGF State-of-the-Art Report (2022); Lagorce et. al., “HOTS: A hierarchy of event based time surfaces for pattern recognition,” IEEE TPAMI (2017); and Rebecq et. al., “Events-to-video: Bringing modern computer vision to event cameras,” CVPR (2019), all of which are incorporated herein by reference in entirety.

[0081] As the first use-case, events in event camera data may be aggregated with a given time interval of 50ms. An (e.g., entire, etc.) event sequence in the event camera data may be partitioned into event sub-sequences of time intervals of 50ms. These event sub-sequences each of which covers a corresponding time interval of 50ms – in an overall time duration covered bythe event sequence in the event camera data – may be respectively denoted as Z^E^, E^A^,Z^E^^, E^AA^, and so on. Raw or pre-aggregated events in these event sub-sequences of timeintervals of 50ms may be respectively aggregated into aggregated events denoted asZw&&^E^, E^A^, Zw&&^E^^, E^AA^, and so on. Each aggregated event Zw&&[E\ , EC] may be identified orindexed with a corresponding timestamp bE (e.g., ^D^ / 50, etc.) for a respective time interval (e.g., between E\and EC, etc.) of 50ms.

[0082] FIG.2A illustrates an example image containing aggregated events generated from aggregating (raw or pre-aggregated) events in (input) event camera data generated by an event camera over 50ms. Each of dark or gray values or pixels represents a “-1” polarity (indicating a decrease in brightness or luminance) or a “+1” polarity (indicating an increase in brightness or luminance), whereas each of white values or pixels represents a “0” polarity (indicating no change in brightness or luminance).

[0083] As an example of modeling aggregated events generated from the same event camera data using a neural field or a MLP network as illustrated in FIG.3A, ten event frames – indexedby respective timestamps bE = 41, 10? – each of which includes aggregated events generated byaggregating (raw or pre-aggregated) events in the event camera data for a respective time interval of 50ms in a sequence of (e.g., consecutive, sequential, etc.) time intervals of 50ms. The aggregated events in the ten event frames may be used to train or fit the neural field or MLP network and generate optimized values for operational parameters such as weights, biases, etc., used in the neural field or MLP network.

[0084] Correspondingly, ten reconstructed event frames – also indexed by respectivetimestamps bE = 41, 10? – each of which includes predicted aggregated events generated orpredicted by the neural field or MLP network operating with the optimized values for the operational parameters for a respective time interval of 50ms in a sequence of (e.g., consecutive, sequential, etc.) time intervals of 50ms.

[0085] Accuracy – of predictions by the neural field or MLP network – can be computed as follows: ^oo^+*oX = Number of correct predictionsTotal number of predictions (16) which can also be represented as ^oo^+*oX = TP + TN(17) where TPFN for False Negative.

[0086] As illustrated in FIG.3A, in the MLP network used to implement the neural field, the solid black arrows represent affine transformations with ReLU activation function, and the broken arrow represents an affine transformations, for example with a non-ReLU function. The affine transformations can be described as linear transformations with added bias. The numbers in each block describe fully connected layers. The output from the MLP network consists of three (3) values respectively corresponding to three (3) classification categories: +1, -1 and 0. The input to the MLP network is represented by a 41-dimensional vector – for each coordinate^^, X^, positional encoding 3 with ^ = 10 is applied resulting in 40 values and an additionalvalue indicating a timestamp (bE) for the vector.

[0087] FIG.2B illustrates an example reconstructed event frame with prediction accuracy computed based on comparing predicted aggregated events generated from the neural field or MLP network with aggregated events generated or aggregated from the event camera data for the same time interval of 50ms (with timestamp dt = 1). Similarly, other reconstructed event frames for other time intervals of 50ms (with other timestamps) may be generated by predicted aggregated events from the neural field or MLP network with respective computed prediction accuracies.

[0088] Additionally, optionally or alternatively, time intervals of different values other than 50ms may be used to aggregate events in the event camera data into event frames for the purpose of validating general applicability of data modeling operations as described herein. For example, the events in the event camera data may be aggregated using a time interval of 10ms, 1ms, etc. A slight drop in accuracy (e.g., ~97%, etc.) may be observed for relatively small time intervals such as time intervals of 1ms, as compared with relatively large time intervals such as time intervals of 10ms or 50ms. This may be due to class imbalance in that pixels without any polarity (belonging to non-event pixel locations) outnumber pixels with polarity (increase or decrease in brightness or luminance). This class imbalance may make the neural field modelingmore challenging. An example solution to tackle the class imbalance problem will be described later in further detail.

[0089] In some operational scenarios, post-training (e.g., static, dynamic, linear, non-linear, etc.) quantization may be applied to the operational parameters or the optimized values thereof in connection with the neural field or MLP network. For example, the weights, biases and / or activation function parameters may be internally represented in relatively high precision numeric values such as floating points, etc. These operational parameter values can be quantized into relatively low precision numeric values such as integer values. The quantized values may be signaled or included in a coded bitstream from an upstream device to a downstream recipient device to enable the downstream device to decode the quantized values for the operational parameters of the neural field or MLP network and use the neural field or MLP network operating with quantized values for the optional parameters to make predictions of aggregated events for the reconstructed event camera data domain.

[0090] By way of example but not limitation, relatively high precision operational parameter values of the neural field trained with the aggregated events – an example event frame for some or all of these aggregated events is illustrated in FIG.2A – as generated from the input event camera data may have a size of 0.97MB. After applying quantization, relatively low precision operational parameter values – or quantized operational parameter values – of the neural field or MLP network have a decreased size of 0.27MB.

[0091] FIG.2C illustrates an example reconstructed event frame with prediction accuracy computed based on comparing predicated aggregated events generated from the neural field or MLP network with aggregated events generated or aggregated from the event camera data for the same time interval of 50ms (with timestamp dt = 1) using quantized operational parameters.

[0092] Compared to the neural field or MLP network operating with relatively high precision (pre-quantized) operational parameters model, the neural field or MLP network operating with the quantized operational parameters results in a drop in (prediction) accuracy.

[0093] As previously noted, in some operational scenarios, event camera data may include multiple sequences of (polarity, sensor element fired) events generated by a system of multiple event cameras for multiple views. These multiple event cameras may be spatially located and / or oriented differently from one another in a physical environment or visual scene.

[0094] Under techniques as described herein, such event camera data can be modeled with a neural field or MLP network. An event ^^^E^ in the event camera data may be defined orrepresented as a 4-parameter tuple ^^^E^ = ^^^, X^ , E^, Y^^. Positional encoding 3 may beapplied to coordinates ^^, X^ of events in the event camera data. Given the coordinates ^^, X^ ofa pixel location, a timestamp bE, and a view ^, the neural field or MLP network may be used to predict a polarity of the pixel location in a reconstructed event data domain, as follows: Ŷ = "^#c$^3^^^, 3^X^, bE, ^^ (18)where, Ŷ represents the predicted polarity of the pixel location having the co-ordinates ^^, X^ atthe timestamp bE and view ^.

[0095] In some operational scenarios, the event camera data may include events for two views: left (or left-eye) view and right (or right-eye) view. These views ^ may be respectively denoted as 0 and 1.

[0096] Each event sequence for each view may be divided or partitioned into time intervals of a specific length such as 50ms. Events within each time interval may be aggregated (e.g., taking the last occurring event or polarity for an event pixel location, etc.) into aggregated events in an event / image frame for each view independently. These aggregated events in the event / image frame may be used as event training data. Non-events or non-event locations in the event / image frame may be identified or determined – all pixel locations other than the event pixel locations – for each view independently. Polarities of the non-event location may be assigned a specific value such as zero (a) and included in overall training data with the aggregated events for the event pixel locations.

[0097] In a model (or neural field) training phase implemented with the upstream device, the training data may be used to train or fit the neural field or MLP network to generate optimized values for operational parameters such as weights, biases, activation function parameters, etc., for the neural field or MLP network.

[0098] In the model application (or inference) phase, the optimized operational parameters or a quantized version thereof may be used by a downstream recipient device to generate reconstructed aggregated events for various pixel locations in a reconstructed image / event frame. 8. MODELING RAW EVENTS

[0099] Techniques as described herein may be used to implement deep learning algorithms that can be applied to raw event data or (sensor-generated / raw) events represented therein – without using synchronous frames to aggregate these events. In some operational scenarios, a (sensor-generated / raw) event in the raw event data contains information about only two (2) polarities: positive polarity for brightness / luminance increase, and negative polarity for brightness / luminance decrease. As noted before, these two polarities may be denoted by the values +1 and −1. [000100] The problem formulation for modeling raw events using a neural field may be similar to the problem formulation for modeling aggregated events as discussed earlier. For example, an event ^^^E^ in the (input) event camera data can be specified or defined as a 4-parameter tuple^^^E^ = ^^^, X^, E^, Y^^, where k represents an (event) index or identifier used to distinguishamong all the events included in the event camera data. Positional encoding 3 can be applied tocoordinates ^^, X^. Given specific coordinates of a pixel location and specific time (ortimestamp) E, the neural field or MLP network can be trained or fitted to predict a corresponding polarity of an event at the pixel location, as follows: Ŷ = "^#c$^3^^^, 3^X^, E^ (19)where Ŷ represents the predicted polarity of the event for the coordinate ^^, X^ at the timestampE. [000101] The design of the neural field or MLP network modeling the raw events may also be similar to that used to model aggregated events. Since a raw event stream – or a sequence of rawevents in the event camera data – denoted as Z[E\, EC] = 〈^^E^|E ∈ 4E\, EC?〉 consists of events ofonly two polarities (e.g., +1 and -1, etc.) to either an increase or decrease inbrightness or luminance, the neural field or MLP network may be adapted such that it can output only two classification categories or only two polarities, as illustrated in FIG.3B. [000102] This neural field or MLP network can be used to model the raw event stream (orsequence of raw events) Z[E\, EC] = 〈^^E^|E ∈ 4E\, EC?〉, but may not model the polarity (e.g., 0 ora value other than +11 polarities, etc.) of non-events or pixel locations at which no event occurs at a given time (or timestamp) t. [000103] As an example of using neural field to model raw event data, assume that event camera data includes a raw event stream such as a sequence of raw events covering an (event) time duration of interval of 10ms and that a temporal resolution of timestamps of the events ismicrosecond. The raw event stream may be denoted as Z^EA, E^AAAA^. Timestamp values t of theevents may be normalized – into a value range of [0, 1] – as follows: EN^^^ = 'B 'lim'l^!B 'lim (20)where10000(microseconds). Hence, the raw event stream may be alternatively denoted as Z[EN^^^8 , EN^^^<]using normalized timestamp value, where EN^^^8 = 0 and EN^^^< = 1. For the purpose ofillustration only, assuming the total number of events in the time duration or interval between the normalized timestamp values EN^^^8and EN^^^<is 185337. The total number of events in theraw event stream or a corresponding dataset may be given as ^Z[EN^^^8 , EN^^^<]^ = 185337.[000104] The problem to predict (reconstructed) eventsdomain usingthe neural field corresponding to the raw event stream Z[EN^^^8, EN^^^<] may be formulated asfollows: Ŷ = "^#cc$^3^^^, 3^X^, EN^^^^ (21)where Ŷ represents the predicted polarity of the event for the pixel location at the coordinates^^, X^ at a given normalized timestamp EN^^^.[000105] In a model training phase such as implemented or executed by an upstream device, the raw event stream or the events therein may be used as training data to train or fit the neural field or MLP network. Optimized values for operational parameters such as weights, biases, activation function parameters, etc., used in the neural field or MLP network may be obtained by minimizing or back propagating prediction errors as measured with a loss function selected for the neural field or MLP network. The optimized operational parameter values (e.g., unquantized, quantized, etc.) may be conveyed or delivered directly or indirectly by the upstream device to a downstream recipient device. [000106] In a model application (or inference) phase such as implemented or executed by the downstream device, the neural field or MLP network operating with the optimized operational parameter values may be used to predict or generate reconstructed events for given coordinates at given time (or timestamp) in a time duration or interval corresponding to that covered by the raw event stream. [000107] Accuracy of predictions by the neural field or MLP network – can be computed as follows: ^oo^+*oX = Number of correct predictionsTotal number of predictions (22) [000108] In some operational scenarios, an accuracy of 100% may be achieved by the neural field or MLP network operating with non-quantized or pre-quantized optimized operational parameter values on predicting or generating the reconstructed raw events as compared with the (input) raw events in the (input) event camera data. [000109] In some operational scenarios, post-training quantization (e.g., static quantization routine in Pytorch, etc.) is applied to generate or obtain a quantized model or quantized optimized operational parameter values for the neural field or MLP network, the predictionaccuracy drops down slightly to 99%, which still demonstrates that a neural field can be effectively used to model raw event data. [000110] Since the neural field or MLP network has not modeled – or has not been trained or fitted with – polarity values at non-event (pixel) locations, the neural field or MLP network is not configured to predict or generate reconstructed polarity values at these non-event locations. In other words, the model has not learned the third polarity or the third classification category other than the two polarities (or two corresponding classification categories) of the raw events in the event camera data and is thus unable to predict the third polarity or the third classification category. Rather, the model can be trained or used to predict the two polarities such as +1 and –1 for the event (pixel) locations. Therefore, the model can predict a specific polarity given a spatiotemporal location (e.g., (event) pixel location coordinates, event timestamp, etc.) of an event. However, the prediction from the model for non-event (pixel) location remains ambiguous. 9. MODELING RAW EVENT DATA WITH NON-EVENTS [000111] A spatiotemporal location (e.g., pixel location coordinates, timestamp, etc.) at which no event is present in the event camera data may be referred to as a non-event. To solve the limitations of the model trained or fitted for predicting the raw events only, the neural field or MLP network may be adapted or enhanced to model or predict the raw events as well as non- events. [000112] The problem to predict (reconstructed) raw events corresponding to the raw eventstream Z[EN^^^8 , EN^^^<] as well as non-events – which are not explicitly represented in the rawevent stream [EN^^^8 , EN^^^<] – in a reconstructed event domain using the neural field may beformulated as follows: Ŷ = "^#c$^3^^^, 3^X^, EN^^^^ (23)[000113] The neural field may be implemented with an MLP network such as shown in FIG. 3A that can predict three polarities. Training data used as input to the MLP network may beenhanced or augmented to include both the raw event data including the raw events included or explicitly represented in the event camera data and (additional) non-event data including the non-events not included or explicitly represented in the event camera data. [000114] The non-event data may be created or generated as follows: p^^^E^ = [^q^, Xq^, EN^^^^ , 0] (24)where the locations.[000115]or constructed as follows: ^Z[EN^^^8, EN^^^<] = 〈p^^E^|E ∈ 4EN^^^8, EN^^^<?〉 (25)[000116] networkfor modeling both raw event stream and non-events may be given as follows: Z[EN^^^8, EN^^^<] ∪ ^Z[EN^^^8, EN^^^<] (26)[000117] inthe raw events at every (non-event) pixel location at which no event is present in the raw camera data and to include these non-events as a part of training data for the neural field, such training data (set) would become very large. Consequently, a much larger neural field would be needed to model the training data including the raw events and non-events. [000118] Instead, selected non-event pixel locations may be sampled from all (non-event) pixels or pixel locations in an image (of a given time or timestamp) that includes all raw events at the given time or timestamp along with the all (non-event) pixels or pixel locations. In some operational scenarios, the neural field can be trained with training data that include non-events at the selected non-event pixel locations and used to support interpolation for the purpose of predicting polarities of other non-events at other (non-selected) non-event pixel locations. [000119] In some operational scenarios, selected non-event pixel locations may be sampled as pixel locations regularly placed from across the image in a grid-wise spatial arrangement. More specifically, the selected non-event pixel may be composed of or selected from p1evenly spacedpoints (or pixel locations) forming a grid (p evenly distributed samples across the width and p evenly distributed samples across the height) across the image, as follows:4^q^, Xq^? = ^^Sℎ,+Rb^^, ^^ (27)where ^= 1: p: ^, W is the width of the image;^ = 1: p: ^, H is the height of the image;and meshgrid() represents a (e.g., Matlab, etc.) function to create a grid of selected non-eventpixels or pixel locations 4^q^, Xq^? based on two vectors or dimensions such as X and Y.[000120] The initially created non-event stream may include a subset of non-event pixel locations identified or selected in the image that coincide with some of the event pixel locations.Hence, the non-event stream ^Z[EN^^^8 , EN^^^<] may be updated or modified by removing thosenon-event coordinates for every timestamp EN^^^^if these coordinates correspond to an event pixel location, as follows: ^^^ = {^^^^ − {[^^^ ∩ ^^]^ ∀ EN^^^^ (28)where ^^^pixellocations for each time stamp EN^^^^. [000121] In some operational scenarios, a neural field trained or fitted with the overall training data including the raw events in the event camera data and selected non-events for all timestamps (or timepoint / time instances) represented in the raw events can achieve a relatively high prediction accuracy such as approximately 99.5% accuracy for all reconstructed events corresponding to the raw events. Furthermore, the neural field can achieve a relatively high prediction accuracy for all reconstructed events corresponding to the selected non-events arranged or placed on the grid across the image for every timestamp. For example, the neural field can overall predict the raw events as well as the selected non-events on the grid locations with approximately 99% accuracy. [000122] Since the grid locations selected as the non-event pixel locations may be fixed for each timestamp, the neural field does not receive any training outside the fixed locations. Hence,the neural field may have a tendency to just predict the polarity at these fixed selected non-event locations and does not have sufficient knowledge or understanding from training on how to predict or interpolate polarities for non-selected non-event locations across the grid locations. [000123] In some operational scenarios, instead of using a regular grid of pixels for the third polarity indicating non-events, a random grid of pixels for the third polarity is used in an imagefor each timestamp EN^^^^. More specifically, selected non-event pixel ^^q^ , Xq^^ may becomposed of or selected from p1points (or pixel locations) forming a grid (p randomly distributed samples across the width and p randomly distributed samples across the height)across the image such that ^q^ ∈ ^ and ^q^ ∈ 41, ^? and Xq^ ∈ ^ and Xq^ ∈ 41, ^?, where H is theheight of the Image and W is the width of the image. As noted previously, the initial non-eventstream ^Z[EN^^^8 , EN^^^<] formed by sampled points may be updated or modified by removingnon-event coordinates for every timestamp EN^^^^if the coordinates correspond to an event pixel location.[000124] For the purpose of illustration only, p = 100 non-event pixel locations are randomlysampled – or individually or differently randomly sampled for each timestamp represented in the raw events in the event camera data – across each of X and Y dimensions of an image at each such timestamp. The neural field may be trained or fitted using this overall training data including the raw events as well as the non-events corresponding on randomly sampled grids of all the timestamps represented in the raw events as excluded through any raw events with coinciding coordinates on any sampled grid points. [000125] FIG.2D illustrates example reconstructed or predicted events and non-events from the neural field for timestamp = 0. In this image, the dark pixel stands for -1 polarity (implying decrease in brightness or luminance) or +1 polarity (implying increase in brightness or luminance) and the gray pixel indicates 0 polarity (indicating no change in brightness or luminance).[000126] Across all the timestamps, the neural network may achieve a relatively high prediction accuracy such as approximately 80% at the event pixel locations as well as improvement for the non-event pixel locations as compared with regularly sampled grids. [000127] In some operational scenarios, it may be observed that the neural field – e.g., with random sampled grid points, etc. – still tends to aggregate data across different timestamps and hence fails to sufficiently appreciate or learn differences among different data portions across these different timestamps. SIREN activation function or an NeRV architecture – examples of which are described in Chen et. al., “NeRV: Neural Representation for Videos,” NeurIPS (2021), all of which are incorporated herein by reference in entirety – may be used or implemented to improve the neural field for better learning and predictions across different timestamps. [000128] To further understand or assess efficacy of random sampling of pixels, one may consider or evaluate only a relatively small subset in all timestamps such as only three (3)different timestamps EN^^^^ = {0, 0.5, 1^ and determine whether the neural field can modelboth the events and non-events across those timestamps. [000129] To make the evaluation more tractable, instead of including non-event pixels for every timestamp, only a relatively small subset – e.g., 40000 samples – in all non-events of an image at a given timestamp is considered. For example, if the size of the image is 640x480 pixels and if there are 20 events in the image, then there are 307180 non-events in that image for that timestamp. Hence, 40000 samples may represent only a small subset among 307180 non- events in that image for that timestamp.[000130] The total number of non-events for three (3) timestamps EN^^^^ = {0, 0.5, 1^ may begiven as: |pZ| ≈ 40000 ∗ 3. This number is approximate since non-event pixel locations that liein the same place as or have coordinates coinciding the event pixel locations are excluded from the overall training data. [000131] Prediction accuracies for these timestamps exhibit similar trends: while modeling or predicting events at the event pixel locations well (e.g., relatively accurately, etc.), the neuralfield may not model or predict non-events at non-event locations well. As noted earlier, the neural field exhibits a tendency to aggregate data (e.g., between or among adjacent timestamps, etc.). These trends are also observed with considering or evaluating more timestamps such aseleven (11) timestamps EN^^^^ = {0,0.1,0.2,0.3,0.4, 0.5, 0.6,0.7,0.8,0.9,1^ and / or considering orevaluating different subsets such as 1500 samples among all non-events of an image at a giventimestamp and / or different total number of non-events for all these timestamps |pZ| ≈ 15000 ∗11. [000132] In addition, to further understand or assess efficacy of random sampling of pixels, one may consider or evaluate one pixel from an image and use a neural field to model its polarity across all the timestamps represented in the raw events in the event camera data. More specifically, all the events and non-events to be modeled by the neural network given aparticular pixel location ^^, X^ for all timestamps between an initial time pO+^A and a final timepO+^^can be represented as follows: Z[EN^^^8, EN^^^<] ∪ ^Z[EN^^^8, EN^^^<] ∶ ^^, X^ = [*0, ¥¦] (29)where [*0,[000133] For example, an event stream including the events can be denoted as Z^EA, E^AAAA^ for10000 timestamps between the initial time pO+^A= 0 and the final time pO+^^= 10000 giventhe particular pixel ^^, X^ = ^373, 225^. For the purpose of illustration only, there are only 24events across 10000 timestamps in the event stream Z^EA, E^AAAA^: ^Z[EN^^^8 , EN^^^<]^ = 24.[000134] FIG.2E illustrates example modeling of (timeneural field with predicted events. As can be seen, the neural field may not be able to model these varying events. This problem can be caused by a class imbalance problem between the classes or polarities of +1 and -1 represented in the raw events and the class or priority assigned to the non- events, as the total number of the non-events far outnumber the total number of the raw eventsgiven ^Z[EN^^^8 , EN^^^<]^ = 24 *pb ^pZ[EN^^^8 , EN^^^<]^ = 9976 at the pixel location ^^, X^ =[000135] To address this issue, more class balance may be created using a number of different approaches. In a first example, the events may be repeated 100 times, as follows: Z[EN^^^8 , EN^^^<] ∪ Z§[¨EN¨^¨^^¨¨8,¨ EN¨^^̈^ ¨<¨]¨ ∪¨ …©¨ .∪¨ Z¨[¨EN¨^¨^¨^¨8 ,¨ EN¨^¨^^¨<ª] ∪ ^Z[EN^^^8, EN^^^<]«« '\^^¬, As aresult of this approach to create more class balance, an improvement may be observed in the prediction of raw events generated by the neural field as compared with predictions by a neural field trained with training data with a relatively pronounced class imbalance problem. [000137] In a second example, the events may be repeated 200 times, as follows: Z[EN^^^8 , EN^^^<] ∪ Z§[¨EN¨^¨^^¨8¨,¨ EN¨^¨^^¨<¨]¨ ∪¨ …©¨ .∪¨ Z¨[¨E¨N^¨^^¨8¨, ËN ¨^¨^^¨<ª] ∪ ^Z[EN^^^8 , EN^^^<]^«« '\^^¬[000138] This results in ^Z[EN^^^8, EN^^^<]^ = 4800 *pb ^pZ[EN^^^8 , EN^^^<]^ = 9976, whichcreates more class balance than the first example. As a result of this approach to create more class balance, an even larger improvement may be observed in the prediction of raw events generated by the neural field, which can predict most of the events along with an improved prediction for the non-events. [000139] In a third example, the events may be repeated 500 times, as follows: Z[EN^^^8 , EN^^^<] ∪ Z§[¨EN¨^¨^^¨8¨ ,¨ EN¨^¨^^¨<¨]¨ ∪¨ …©¨ .∪¨ Z¨[¨EN¨^¨^¨^8¨, ËN ¨^¨^^¨<ª]9976. Animprovement may also be observed in the prediction of raw events generated by the neural field, which can predict most of the events along with an improved prediction for the non-events. [000141] Hence, it is possible to model changing polarities of one pixel across all timestamps. This can be translated to state that the changing polarities of all pixels of the image can be modeled using a neural field. However, there may still be a tendency of the neural field to modelaggregated polarity data (and to fail to capture some of the subtle variations over different timestamps). This aggregation problem may be mitigated by parameter tuning and / or by loss function selection. 10. COMPRESSING RAW EVENT CAMERA DATA USING NEURAL FIELDS [000142] Optimized values of operational parameters may be used to define or specify a neural field of a selected structure used to model polarities of raw events in event camera data or aggregated events generated from the raw events. These optimized values of operational parameters along with the selected structure logically provide a representation of the neural field as well as a representation of the event camera data used to generate training data to train or fit the neural field. In some operational parameters, the representation of the neural field or the event camera data may be compressed. The compressed representation may be explicitly or implicitly coded into a coded bitstream to be delivered from an upstream device to a downstream recipient device. [000143] For example, the neural field may be implemented with an MLP network. The selected structure of the neural field may be a selected MLP network structure among one or more different candidate MLP network structures. The selected neural field or MLP network structure may be determined or adopted based at least in part on evaluating and comparing the different candidate neural field or MLP structures. The evaluation and comparison can be performed in part or in whole to balance the need for relatively high accuracy in prediction and the need for relatively effective data compression. [000144] For the purpose of illustration only, different candidate neural field or MLP network structures of FIG.3B through FIG.3F may be evaluated or compared for compression performance. Each of these neural field or MLP network structures may be implemented by a neural field or MLP network to model the same raw events in the same event camera data. Theseraw events may be included in a raw event stream Z[EN^^^8 , EN^^^<] with only two possiblepolarities. The candidate neural field or MLP networks of through FIG.3F havedifferent data sizes for operational parameter values associated with different total numbers of neurons used in (e.g., input, hidden, output, etc.) layers in the candidate neural fields or MLP networks and provide different accuracies in reconstructing the events as illustrated in TABLE 1 below. TABLE 1 Neural Model Size Accuracy (%) Quantized model size Accuracy (%) 1702KB 100 464KB 100 [0, evaluated) MLP network as illustrated in FIG.3C has a data size 1702KB and provides a prediction accuracy of 100%. A (candidate or evaluated) second MLP network as illustrated in FIG.3D has a data size 1375KB and provides a prediction accuracy of 100%. A (candidate or evaluated) third MLP network as illustrated in FIG.3B has a data size 961KB and provides a prediction accuracy of 100%. A (candidate or evaluated) fourth MLP network as illustrated in FIG.3E has a data size 770KB and provides a prediction accuracy of 99%. The (candidate or evaluated) fifth MLP network as illustrated in FIG.3F has a data size 217KB and provides a prediction accuracy of 92%. [000146] The operational parameter values of these different MLPs may be quantized respectively. [000147] As illustrated in (the two rightmost columns of) TABLE 1, the first (candidate or evaluated) MLP network as illustrated in FIG.3C operating with quantized operational parameter values has a data size 464KB and provides a prediction accuracy of 100%. The (candidate or evaluated) second MLP network as illustrated in FIG.3D operating with quantized operational parameter values has a data size 378KB and provides a prediction accuracy of 100%. The (candidate or evaluated) third MLP network as illustrated in FIG.3B operating withquantized operational parameter values has a data size 277KB and provides a prediction accuracy of 99%. The (candidate or evaluated) fourth MLP network as illustrated in FIG.3E operating with quantized operational parameter values has a data size 226KB and provides a prediction accuracy of 97%. The (candidate or evaluated) fifth MLP network as illustrated in FIG.3F operating with quantized operational parameter values has a data size 75KB and provides a prediction accuracy of 83%. [000148] In some operational scenarios, in addition to including (optimized or trained) operational parameter values of the neural field or MLP network in a coded bitstream as described herein, coordinates of the event pixel locations at which the raw events occur may be compressed or included in the coded bitstream. [000149] Similar evaluation or comparison or selection may be made with neural fields or MLPs used to model aggregated events generated from aggregating raw events in event camera data over time intervals of a fixed length such as 50ms, 10ms, etc. [000150] In an example, a quantized neural field modeling aggregated events with 50ms time intervals with a data size of 277KB provides a 99% prediction accuracy. The data size of the quantized neural field may be further reduced, for example by saving or encoding into a lossless PNG format, resulting in a further compressed data size of around 80KB. [000151] For example, a quantized neural field modeling aggregated events with 50ms time intervals with a data size of 277KB provides a 99% prediction accuracy. The data size of the quantized neural field may be further reduced, for example by saving or encoding into a lossless PNG format, resulting in a further compressed data size of around 80KB. [000152] The quantization of neural fields demonstrate that relatively small models (e.g., with relatively small total number of neurons and / or relatively small data sizes for operational parameter values, etc.) can be used to provide a representation of modeled raw events or modeled aggregated events without incurring a significant drop in accuracy. Hence, neural fields can be used as an approach to compress event camera data.[000153] For the purpose of illustration only, it has been described that an MLP network may be used to implement a neural field and predict polarities and / or non-events. It should be noted, however, that more than one MLP network may be used to implement a neural field as described herein. In an example, the neural field may include – or may be implemented with – two different MLP networks each of which is specifically trained to predict one type of (positive or negative polarities given pixel locations. In another example, the neural field may include – or may be implemented with – three different MLP networks each of two of which is specifically trained to predict one type of (positive or negative polarities given pixel locations and one of which is specifically trained to predict non-events given pixel locations. Additionally, optionally or alternatively, other artificial neural networks including but not limited to MLP networks may be used to implement at least in part the neural field as described herein. For example, pre- processing or post-processing (e.g., CNN, MLP, non-MLP, transformer, etc.) neural networks may be implemented with the neural field. A post-processing neural network may be used to receive outputs from MLP networks that predict events and / or non-events given pixel locations and generate further processed outputs including but not limited to predicted events and / or non- events with a relatively high accuracy. 12. EXAMPLE PROCESS FLOWS [000154] FIG.4A illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., an upstream device, one or more event data codecs, an encoding device / module, a transcoding device / module, a media device / module, etc.) may perform this process flow. [000155] In block 402, an upstream device receives event camera data containing a sequence of raw events. Each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor. Each sensor element among the plurality of sensor elements of the event image sensor corresponds to a respective pixel position among a plurality of pixel locations.[000156] In block 404, the upstream device generates neural field training data from the sequence of raw events in the event camera data. The neural field training data includes a plurality of training instances. Each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance. [000157] In block 406, in a training phase, the upstream device generates optimized values for operational parameters of a neural field by training the neural field with the neural field training data based at least in part on a specific loss function selected for the neural field. The neural field is trained to output respective predicted classification categories for pixel locations in the plurality of pixel locations at a plurality of time points. [000158] In block 408, the upstream device encodes neural field operational parameter values in a coded bitstream. The neural field operational parameter values are generated from the optimized values for the operational parameters of the neural field. The coded bitstream causes a recipient device to use the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data. [000159] In an embodiment, each of the ground truth classification category and the predicted classification categories belongs to a common set of classification categories; the common set of classification categories represents a set of different luminance change indicators. [000160] In an embodiment, the common set of classification categories represents a set of two classification categories consisting of a first luminance change indicator for a positive luminance change over a positive luminance change threshold, and a second luminance change indicator for a negative luminance change over a negative luminance change threshold. [000161] In an embodiment, the common set of classification categories represents a set of three classification categories consisting of a first luminance change indicator for a positive luminance change over a positive luminance change threshold, a second luminance changeindicator for a negative luminance change over a negative luminance change threshold, and a third luminance change indicator for no luminance change over the positive or negative luminance change threshold. [000162] In an embodiment, events represented in the plurality of training instances are one of: raw events or aggregated events. [000163] In an embodiment, each of the aggregated events is generated from aggregating raw events at a specific event pixel location for a specific time interval; the specific time interval is set based at least in part on user input. [000164] In an embodiment, the neural field operational parameter values encoded in the coded bitstream are generated by quantizing the optimized values for the operational parameters of the neural field. [000165] In an embodiment, the neural field is implemented with a multi-layer perceptron (MLP) network; the MLP network is trained using positional data in the training data as input and polarities in the training data as ground truth. [000166] In an embodiment, the positional data is generated by applying positional encoding to coordinates of pixel locations represented in the training data. [000167] In an embodiment, the MLP network includes an output layer that uses activation functions that are of a different function type from other activation functions used in other layers of the MLP network. [000168] In an embodiment, a neuron in each of all subsequent layers of the MLP network is fully connected with all neurons in an immediately preceding layer in the MLP network. [000169] FIG.4B illustrates an example process flow according to an embodiment. In some embodiments, one or more computing devices or components (e.g., a downstream device, one or more event data codecs, a decoding device / module, a transcoding device / module, a media device / module, etc.) may perform this process flow.[000170] In block 452, a downstream device decodes, from a coded bitstream generated by an upstream device, neural field operational parameter values. The neural field operational parameter values are generated from optimized values for operational parameters of a neural field. [000171] The optimized values for the operational parameters of the neural field were generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field. The neural field is trained to output predicted classification categories. [000172] The neural field training data includes a plurality of training instances. Each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance. The neural field training data is generated from a sequence of raw events in event camera data. [000173] Each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor. Each sensor element among the plurality of sensor elements of the event image sensor corresponds to a respective pixel position among a plurality of pixel locations. [000174] In block 454, the downstream device uses the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data. [000175] In an embodiment, the reconstructed events are used to generate visual information relating to a physical environment. [000176] In an embodiment, a computing device such as a display device, a mobile device, a set-top box, a multimedia device, etc., is configured to perform any of the foregoing methods. In an embodiment, an apparatus comprises a processor and is configured to perform any of the foregoing methods. In an embodiment, a non-transitory computer readable storage medium,storing software instructions, which when executed by one or more processors cause performance of any of the foregoing methods. [000177] In an embodiment, a computing device comprising one or more processors and one or more storage media storing a set of instructions which, when executed by the one or more processors, cause performance of any of the foregoing methods. [000178] Note that, although separate embodiments are discussed herein, any combination of embodiments and / or partial embodiments discussed herein may be combined to form further embodiments. 11. IMPLEMENTATION MECHANISMS – HARDWARE OVERVIEW [000179] Embodiments of the present invention may be implemented with a computer system, systems configured in electronic circuitry and components, an integrated circuit (IC) device such as a microcontroller, a field programmable gate array (FPGA), or another configurable or programmable logic device (PLD), a discrete time or digital signal processor (DSP), an application specific IC (ASIC), and / or apparatus that includes one or more of such systems, devices or components. The computer and / or IC may perform, control, or execute instructions relating to the adaptive perceptual quantization of images with enhanced dynamic range, such as those described herein. The computer and / or IC may compute any of a variety of parameters or values that relate to the adaptive perceptual quantization processes described herein. The image and video embodiments may be implemented in hardware, software, firmware and various combinations thereof. [000180] Certain implementations of the inventio comprise computer processors which execute software instructions which cause the processors to perform a method of the disclosure. For example, one or more processors in a display, an encoder, a set top box, a transcoder or the like may implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processors. Embodiments of the invention may also be provided in the form of a programproduct. The program product may comprise any non-transitory medium which carries a set of computer-readable signals comprising instructions which, when executed by a data processor, cause the data processor to execute a method of an embodiment of the invention. Program products according to embodiments of the invention may be in any of a wide variety of forms. The program product may comprise, for example, physical media such as magnetic data storage media including floppy diskettes, hard disk drives, optical data storage media including CD ROMs, DVDs, electronic data storage media including ROMs, flash RAM, or the like. The computer-readable signals on the program product may optionally be compressed or encrypted. [000181] Where a component (e.g. a software module, processor, assembly, device, circuit, etc.) is referred to above, unless otherwise indicated, reference to that component (including a reference to a "means") should be interpreted as including as equivalents of that component any component which performs the function of the described component (e.g., that is functionally equivalent), including components which are not structurally equivalent to the disclosed structure which performs the function in the illustrated example embodiments of the invention. [000182] According to one embodiment, the techniques described herein are implemented by one or more special-purpose computing devices. The special-purpose computing devices may be hard-wired to perform the techniques, or may include digital electronic devices such as one or more application-specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs) that are persistently programmed to perform the techniques, or may include one or more general purpose hardware processors programmed to perform the techniques pursuant to program instructions in firmware, memory, other storage, or a combination. Such special- purpose computing devices may also combine custom hard-wired logic, ASICs, or FPGAs with custom programming to accomplish the techniques. The special-purpose computing devices may be desktop computer systems, portable computer systems, handheld devices, networking devices or any other device that incorporates hard-wired and / or program logic to implement the techniques.[000183] For example, FIG.5 is a block diagram that illustrates a computer system 500 upon which an embodiment of the invention may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for communicating information, and a hardware processor 504 coupled with bus 502 for processing information. Hardware processor 504 may be, for example, a general purpose microprocessor. [000184] Computer system 500 also includes a main memory 506, such as a random access memory (RAM) or other dynamic storage device, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 also may be used for storing temporary variables or other intermediate information during execution of instructions to be executed by processor 504. Such instructions, when stored in non-transitory storage media accessible to processor 504, render computer system 500 into a special-purpose machine that is customized to perform the operations specified in the instructions. [000185] Computer system 500 further includes a read only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions for processor 504. A storage device 510, such as a magnetic disk or optical disk, is provided and coupled to bus 502 for storing information and instructions. [000186] Computer system 500 may be coupled via bus 502 to a display 512, such as a liquid crystal display, for displaying information to a computer user. An input device 514, including alphanumeric and other keys, is coupled to bus 502 for communicating information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, a trackball, or cursor direction keys for communicating direction information and command selections to processor 504 and for controlling cursor movement on display 512. This input device typically has two degrees of freedom in two axes, a first axis (e.g., x) and a second axis (e.g., y), that allows the device to specify positions in a plane. [000187] Computer system 500 may implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logicwhich in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques as described herein are performed by computer system 500 in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the sequences of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In other embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. [000188] The term “storage media” as used herein refers to any non-transitory media that store data and / or instructions that cause a machine to operation in a specific fashion. Such storage media may comprise non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of storage media include, for example, a floppy disk, a flexible disk, hard disk, solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge. [000189] Storage media is distinct from but may be used in conjunction with transmission media. Transmission media participates in transferring information between storage media. For example, transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio-wave and infra-red data communications. [000190] Various forms of media may be involved in carrying one or more sequences of one or more instructions to processor 504 for execution. For example, the instructions may initially be carried on a magnetic disk or solid state drive of a remote computer. The remote computer can load the instructions into its dynamic memory and send the instructions over a telephone lineusing a modem. A modem local to computer system 500 can receive the data on the telephone line and use an infra-red transmitter to convert the data to an infra-red signal. An infra-red detector can receive the data carried in the infra-red signal and appropriate circuitry can place the data on bus 502. Bus 502 carries the data to main memory 506, from which processor 504 retrieves and executes the instructions. The instructions received by main memory 506 may optionally be stored on storage device 510 either before or after execution by processor 504. [000191] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides a two-way data communication coupling to a network link 520 that is connected to a local network 522. For example, communication interface 518 may be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, communication interface 518 may be a local area network (LAN) card to provide a data communication connection to a compatible LAN. Wireless links may also be implemented. In any such implementation, communication interface 518 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information. [000192] Network link 520 typically provides data communication through one or more networks to other data devices. For example, network link 520 may provide a connection through local network 522 to a host computer 524 or to data equipment operated by an Internet Service Provider (ISP) 526. ISP 526 in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet” 528. Local network 522 and Internet 528 both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link 520 and through communication interface 518, which carry the digital data to and from computer system 500, are example forms of transmission media.[000193] Computer system 500 can send messages and receive data, including program code, through the network(s), network link 520 and communication interface 518. In the Internet example, a server 530 might transmit a requested code for an application program through Internet 528, ISP 526, local network 522 and communication interface 518. [000194] The received code may be executed by processor 504 as it is received, and / or stored in storage device 510, or other non-volatile storage for later execution. 13. EQUIVALENTS, EXTENSIONS, ALTERNATIVES AND MISCELLANEOUS [000195] In the foregoing specification, embodiments of the invention have been described with reference to numerous specific details that may vary from implementation to implementation. Thus, the sole and exclusive indicator of what is claimed embodiments of the invention, and is intended by the applicants to be claimed embodiments of the invention, is the set of claims that issue from this application, in the specific form in which such claims issue, including any subsequent correction. Any definitions expressly set forth herein for terms contained in such claims shall govern the meaning of such terms as used in the claims. Hence, no limitation, element, property, feature, advantage or attribute that is not expressly recited in a claim should limit the scope of such claim in any way. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense. Enumerated Exemplary Embodiments [000196] The invention may be embodied in any of the forms described herein, including, but not limited to the following Enumerated Example Embodiments (EEEs) which describe structure, features, and functionality of some portions of embodiments of the present invention. [000197] EEE1. A method, comprising: receiving event camera data containing a sequence of raw events, wherein each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor, wherein each sensor element among the plurality of sensor elements of the event image sensor corresponds toa respective pixel position among a plurality of pixel locations; generating neural field training data from the sequence of raw events in the event camera data, wherein the neural field training data includes a plurality of training instances, wherein each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance; in a training phase, generating optimized values for operational parameters of a neural field by training the neural field with the neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output respective predicted classification categories for pixel locations in the plurality of pixel locations at a plurality of time points; encoding neural field operational parameter values in a coded bitstream, wherein the neural field operational parameter values are generated from the optimized values for the operational parameters of the neural field, wherein the coded bitstream causes a recipient device to use the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data. [000198] EEE2. The method as recited in EEE1, wherein the neural field is implemented with a single multi-layer perceptron (MLP) network in which output neurons of the MLP network predict, for an input pixel location, one of (a) a positive or negative event polarity when the input pixel location represents an event pixel location, or (b) a positive or negative event polarity at the input pixel location when the input pixel location represents the event pixel location and a non-event at the input pixel location when the input pixel location does not represent an event pixel location. [000199] EEE3. The method as recited in EEE1 or EEE2, wherein the neural field is implemented at least with a first multi-layer perceptron (MLP) network and a second MLP network separate from the first MLP network; wherein the first MLP network predicts positive polarities for input pixel locations; wherein the second MLP network predicts negative polarities for the input pixel locations.[000200] EEE4. The method as recited in EEE3, wherein the neural field is further implemented with a third MLP network separate from the first and second MLP networks; wherein the third MLP network predicts non-events for input pixel locations. [000201] EEE5. The method as recited in EEE4, wherein the neural field includes a post- processing MLP network receiving outputs from the first, second and third MLP networks as input and generating predictions of event polarities for given pixel locations. [000202] EEE6. The method as recited in EEE3, wherein the neural field includes a post- processing MLP network receiving outputs from the first and second MLP networks as input and generating predictions of event polarities for given pixel locations. [000203] EEE7. The method as recited in EEE1-EEE6, wherein each of the ground truth classification category and the predicted classification categories belongs to a common set of classification categories; wherein the common set of classification categories represents a set of different luminance change indicators. [000204] EEE8. The method as recited in any of EEE1-EEE7, wherein the specific loss function is selected specifically to support a downstream task to be performed with the reconstructed events; wherein the downstream task relates to one of: an image generation task to be performed by a downstream device, an object recognition task to be performed by the downstream device, or another task to be performed by the downstream device with the reconstructed events. [000205] EEE9. The method as recited in EEE8, wherein the common set of classification categories represents a set of two classification categories consisting of a first luminance change indicator for a positive luminance change over a positive luminance change threshold, and a second luminance change indicator for a negative luminance change over a negative luminance change threshold. [000206] EEE10. The method as recited in EEE8, wherein the common set of classification categories represents a set of three classification categories consisting of a first luminancechange indicator for a positive luminance change over a positive luminance change threshold, a second luminance change indicator for a negative luminance change over a negative luminance change threshold, and a third luminance change indicator for no luminance change over the positive or negative luminance change threshold. [000207] EEE11. The method as recited in any of EEE1-EEE10, wherein events represented in the plurality of training instances are one of: raw events, or aggregated events, time surfaces, event voxels, or input data items in other event representations. [000208] EEE12. The method as recited in EEE11, wherein each of the aggregated events is generated from aggregating raw events at a specific event pixel location for a specific time interval; wherein the specific time interval is set based at least in part on user input. [000209] EEE13. The method as recited in EEE11, wherein each of the time surfaces is generated from applying an exponential decay kernel to a pixel neighborhood around a specific event pixel location to which a respective event in the events corresponds. [000210] EEE14. The method as recited in EEE11, wherein the event voxels are generated from distributing each of the events to nearest voxels in an event voxel grid formed by a point cloud representation of the events. [000211] EEE15. The method as recited in any of EEE1-EEE14, wherein the neural field operational parameter values encoded in the coded bitstream are generated by quantizing the optimized values for the operational parameters of the neural field. [000212] EEE16. The method as recited in any of EEE1-EEE15, wherein the neural field is implemented with a multi-layer perceptron (MLP) network; wherein the MLP network is trained using positional data in the training data as input and polarities in the training data as ground truth. [000213] EEE17. The method as recited in any of EEE1-EEE16, wherein the neural field is further implemented with one or more of: convolutional neural networks, transformer neuralnetworks, artificial neural networks (ANNs) implementing neural representation for videos, or ANNs other than MLP networks. [000214] EEE18. The method as recited in EEE16, wherein the positional data is generated by applying positional encoding to coordinates of pixel locations represented in the training data. [000215] EEE19. The method as recited in EEE16-EEE18, wherein the MLP network includes an output layer that uses activation functions that are of a different function type from other activation functions used in other layers of the MLP network. [000216] EEE20. The method as recited in any of EEE16-EEE19, wherein a neuron in each of all subsequent layers of the MLP network is fully connected with all neurons in an immediately preceding layer in the MLP network. [000217] EEE21. A method, comprising: decoding, from a coded bitstream generated by an upstream device, neural field operational parameter values, wherein the neural field operational parameter values are generated from optimized values for operational parameters of a neural field; wherein the optimized values for the operational parameters of the neural field were generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output predicted classification categories; wherein the neural field training data includes a plurality of training instances, wherein each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event represented in the training instance, wherein the neural field training data is generated from a sequence of raw events in event camera data; wherein each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor, wherein each sensor element among the plurality of sensor elements of the event image sensor corresponds to a respective pixel position among a plurality of pixel locations; using the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented inthe training data. [000218] EEE22. The method as recited in EEE21, wherein the reconstructed events are used to generate visual information relating to a physical environment. [000219] EEE23. The method as recited in EEE21 or EEE22, wherein the plurality of events approximated by the plurality of reconstructed events constitute a subset of all events in the training data; wherein all the events used to train the neural field by the upstream device correspond to first pixel locations and first time points at a first spatiotemporal resolution; wherein the plurality of reconstructed events corresponds to second pixel locations and second time points at a second spatiotemporal resolution different from the first spatiotemporal resolution; wherein the plurality of reconstructed events is used as input to an image generator to generate a sequence of images at one or more of: a relatively high frame rate or a relatively high dynamic range. [000220] EEE24. An apparatus performing any of the methods as recited in EEE1-EEE23. [000221] EEE25. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any of the methods as recited in EEE1-EEE23.

Claims

CLAIMS What is claimed is:

1. A method, comprising: receiving event camera data containing a sequence of raw events at a plurality of time points, wherein each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor, wherein each sensor element among the plurality of sensor elements of the event image sensor corresponds to a respective pixel position among a plurality of pixel locations; generating neural field training data from the sequence of raw events in the event camera data, wherein the neural field training data includes a plurality of training instances, wherein each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event at a given time point represented in the training instance; in a training phase, generating optimized values for operational parameters of a neural field by training the neural field with the neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output respective predicted classification categories for pixel locations in the plurality of pixel locations at the plurality of time points, wherein each of the ground truth classification category and the predicted classification categories belongs to a common set of classification categories; wherein the common set of classification categories represents a set of different luminance change indicators; encoding neural field operational parameter values in a coded bitstream, wherein the neural field operational parameter values are generated from the optimized values for the operational parameters of the neural field, wherein the coded bitstream causes a recipient device to use the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data.

2. The method as recited in Claim 1, wherein the neural field is implemented with a singlemulti-layer perceptron (MLP) network in which output neurons of the MLP network predict, for an input pixel location, one of (a) a positive or negative event polarity when the input pixel location represents an event pixel location, or (b) a positive or negative event polarity at the input pixel location when the input pixel location represents the event pixel location and a non- event at the input pixel location when the input pixel location does not represent an event pixel location.

3. The method as recited in Claim 1, wherein the neural field is implemented at least with a first multi-layer perceptron (MLP) network and a second MLP network separate from the first MLP network; wherein the first MLP network predicts positive polarities for input pixel locations; wherein the second MLP network predicts negative polarities for the input pixel locations.

4. The method as recited in Claim 3, wherein the neural field is further implemented with a third MLP network separate from the first and second MLP networks; wherein the third MLP network predicts non-events for input pixel locations.

5. The method as recited in Claim 4, wherein the neural field includes a post-processing MLP network receiving outputs from the first, second and third MLP networks as input and generating predictions of event polarities for given pixel locations.

6. The method as recited in Claim 3, wherein the neural field includes a post-processing MLP network receiving outputs from the first and second MLP networks as input and generating predictions of event polarities for given pixel locations.

7. The method as recited in any of Claims 1-6, wherein the specific loss function is selected specifically to support a downstream task to be performed with the reconstructed events; wherein the downstream task relates to one of: an image generation task to be performed by adownstream device, an object recognition task to be performed by the downstream device, or another task to be performed by the downstream device with the reconstructed events.

8. The method as recited in any of Claims 1-7, wherein the common set of classification categories represents a set of two classification categories consisting of a first luminance change indicator for a positive luminance change over a positive luminance change threshold, and a second luminance change indicator for a negative luminance change over a negative luminance change threshold.

9. The method as recited in any of Claims 1-7, wherein the common set of classification categories represents a set of three classification categories consisting of a first luminance change indicator for a positive luminance change over a positive luminance change threshold, a second luminance change indicator for a negative luminance change over a negative luminance change threshold, and a third luminance change indicator for no luminance change over the positive or negative luminance change threshold.

10. The method as recited in any of Claims 1-9, wherein events represented in the plurality of training instances are one of: raw events, aggregated events, time surfaces, event voxels, or input data items in other event representations.

11. The method as recited in Claim 10, wherein each of the aggregated events is generated from aggregating raw events at a specific event pixel location for a specific time interval; wherein the specific time interval is set based at least in part on user input.

12. The method as recited in Claim 10, wherein each of the time surfaces is generated from applying an exponential decay kernel to a pixel neighborhood around a specific event pixel location to which a respective event in the events corresponds.

13. The method as recited in Claim 10, wherein the event voxels are generated from distributing each of the events to nearest voxels in an event voxel grid formed by a point cloud representation of the events.

14. The method as recited in any of Claims 1-13, wherein the neural field operational parameter values encoded in the coded bitstream are generated by quantizing the optimized values for the operational parameters of the neural field.

15. The method as recited in any of Claims 1-14, wherein the neural field is implemented with a multi-layer perceptron (MLP) network; wherein the MLP network is trained using positional data in the training data as input and polarities in the training data as ground truth.

16. The method as recited in any of Claims 1-15, wherein the neural field is further implemented with one or more of: convolutional neural networks, transformer neural networks, artificial neural networks (ANNs) implementing neural representation for videos, or ANNs other than MLP networks.

17. The method as recited in Claim 15, wherein the positional data is generated by applying positional encoding to coordinates of pixel locations represented in the training data.

18. The method as recited in any of Claims 15-17, wherein the MLP network includes an output layer that uses activation functions that are of a different function type from other activation functions used in other layers of the MLP network.

19. The method as recited in any of Claims 15-18, wherein a neuron in each of all subsequent layers of the MLP network is fully connected with all neurons in an immediately preceding layer in the MLP network.

20. A method, comprising: decoding, from a coded bitstream generated by an upstream device, neural field operational parameter values, wherein the neural field operational parameter values are generated from optimized values for operational parameters of a neural field; wherein the optimized values for the operational parameters of the neural field were generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output predicted classification categories; wherein the neural field training data includes a plurality of training instances, wherein each training instance in the plurality of training instances of the neural field training data includes coordinates of an event pixel location and a ground truth classification category of an event at a given time point represented in the training instance, wherein the neural field training data is generated from a sequence of raw events at a plurality of time points in event camera data; wherein each raw event in the sequence of raw events is generated in response to a luminance change by a specific sensor element among a plurality of sensor elements of an event image sensor, wherein each sensor element among the plurality of sensor elements of the event image sensor corresponds to a respective pixel position among a plurality of pixel locations; wherein each of the ground truth classification category and the predicted classification categories belongs to a common set of classification categories; wherein the common set of classification categories represents a set of different luminance change indicators; using the neural field operating with the neural field operational parameter values to generate a plurality of reconstructed events approximating a plurality of events represented in the training data.

21. The method as recited in Claim 20, wherein the reconstructed events are used to generate visual information relating to a physical environment.

22. The method as recited in Claim 20 or 21, wherein the plurality of events approximated by the plurality of reconstructed events constitute a subset of all events in the training data; wherein all the events used to train the neural field by the upstream device correspond to first pixel locations and first time points at a first spatiotemporal resolution; wherein the plurality of reconstructed events corresponds to second pixel locations and second time points at a second spatiotemporal resolution different from the first spatiotemporal resolution; wherein the plurality of reconstructed events is used as input to an image generator to generate a sequence of images at one or more of: a relatively high frame rate or a relatively high dynamic range.

23. An apparatus performing any of the methods as recited in Claims 1-22.

24. A non-transitory computer readable medium, storing software instructions, which when executed by one or more processors cause performance of the steps of any of the methods as recited in Claims 1-22.

Citation Information

Patent Citations

  • Event Camera Based Navigation Control

    US20220197312A1

Cited By

  • Three-channel event determination method and device based on action dynamic characteristics, equipment, medium and product

    CN121585922A

  • Event camera motion small target detection method and system based on topology constraint

    CN121962585A