Encoding event camera data using neural fields
By encoding event camera data using neural fields and utilizing MLP networks to achieve efficient encoding and decoding of event data, the problem of low efficiency in event camera data processing in existing technologies is solved, and the accuracy and efficiency of data reconstruction and recognition are improved.
Patent Information
- Application Number
- CN202480080327.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-24
- Filing Date
- 2024-12-16
- Publication Date
- 2026-07-24
AI Technical Summary
Existing technologies struggle to efficiently process event camera data, especially for accurately identifying and reconstructing objects in physical environments or visual scenes under real-time and low-latency conditions.
Neural fields are used to encode event camera data, and multilayer perceptron (MLP) networks are used to model and compress the event camera data. The coded bitstream is generated by training the operating parameters of the neural fields, thus achieving efficient encoding and decoding of event data.
It improves the encoding efficiency and reconstruction accuracy of event camera data, enabling high-quality reconstruction and recognition of physical environments and visual scenes under low latency and low power consumption conditions.
Smart Images

Figure CN122460075A_ABST
Abstract
Description
Cross-references to related applications
[0001] This application claims priority to U.S. Provisional Application No. 63 / 611,975, filed December 19, 2023, and European Patent Application No. 24159531.3, filed February 24, 2024, each of which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure generally relates to visual computing, and more specifically to encoding event camera data using neural fields. Background Technology
[0003] Event cameras with imaging sensors that respond to local brightness changes can be used to generate event camera data from a physical environment or visual scene. For example, event camera data can be generated from an array or spatial distribution of pixels or sensor elements of an image sensor within the event camera, without the use of a shutter. Each pixel or sensor element in the image sensor can operate independently and asynchronously with other pixels or sensor elements in the same image sensor. Event camera data can include a stream of pixel-specific events occurring asynchronously at multiple points in time covering a certain time interval or duration. Each such event indicates (in the physical environment or visual scene) the polarity (increase or decrease) or instantaneous measurement of brightness or luminance at a specific spatial location or orientation corresponding to a specific pixel at a specific point in time represented in the event camera data.
[0004] Techniques are being developed for processing event camera data generated using different types of event cameras. For example, temporal smoothing can be applied to event camera data to construct images of physical environments or visual scenes with relatively wide or relatively dark brightness or luminance ranges. Event camera data can also be processed or filtered to relatively efficiently and responsively identify stationary or moving objects or pedestrians in physical environments or visual scenes.
[0005] The methods described in this section are permissible but not necessarily methods that have been previously conceived or employed. Therefore, unless otherwise instructed, no method described in this section should be considered prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise instructed, any issues concerning one or more methods should not be considered to be in any prior art based on this section. Attached Figure Description
[0006] The present disclosure is illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals refer to similar elements, and in the accompanying drawings: Figure 1AThe illustration shows an example multilayer perceptron (MLP) network that can be used to implement neural fields; Figure 1B The diagram illustrates an example of the basic building blocks of an MLP layer in an MLP network; Figure 1C and Figure 1D The diagram illustrates the structure of an example neural field encoder and decoder; Figure 2A The illustration shows an example image containing an aggregation event; Figures 2B to 2D The illustration shows an example of reconstructed event frames and prediction accuracy; Figure 2E The illustration shows an example of modeling the original event and predicting the event using a neural field; Figures 3A to 3F The illustration shows an example MLP network in which the MLP layers have a different number of neurons; Figure 4A and Figure 4B The example process flow is illustrated; and Figure 5 An example hardware platform on which a computer or computing device as described herein can be implemented is illustrated. Detailed Implementation
[0007] This document describes example embodiments involving the encoding or modeling of event camera data using neural fields. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a thorough understanding of this disclosure. However, it will be apparent that this disclosure can be practiced without these specific details. In other instances, well-known structures and devices have not been described exhaustively in order to avoid unnecessarily obscuring, obscuring, or confusing this disclosure.
[0008] This document describes example embodiments based on the following summary: 1. General Overview 2. Neural Fields and MLP Networks 3. Location coding 4. Event Data 5. Neural Field Coding Framework 6. Modeling aggregation events using neural fields 7. Reconstruct the aggregated events. 8. Model the original event. 9. Modeling raw event data and non-event data. 10. Compress raw event camera data using neural fields. 11. Example Process Flow 12. Implementation Mechanism – Hardware Overview 13. Equivalents, extensions, substitutes, and others 1. General Overview This overview provides a basic description of some aspects of exemplary embodiments of the invention. It should be noted that this overview is not a broad or exhaustive summary of all aspects of the exemplary embodiments. Furthermore, it should be understood that this overview is not intended to identify any particularly important aspects or elements of the exemplary embodiments, nor is it intended to be construed as particularly depicting any scope of the exemplary embodiments, nor is it a general description of this disclosure. This overview only introduces some concepts relevant to the exemplary embodiments in a condensed and simplified format and should be understood as merely a conceptual prelude to a more detailed description of the following exemplary embodiments.
[0009] Neural fields can be used as coordinate-based neural networks to parameterize the physical properties of a scene or object in space and time. Neural fields have applications in visual computing problems such as 3D image reconstruction and image synthesis. In the neural field framework, field quantities are generated by sampling coordinates and then fed into the neural network. The field quantities can represent samples in the target (or desired) reconstruction domain.
[0010] For example, neural fields, such as neural radiation fields (NeRF), can be used to provide 3D scene representations. NeRF can be generated from or located at a specific coordinate (e.g., a specific spatial position (x, y, z)) and / or a specific viewing direction (φ, z). The neural field takes a specific input at a given coordinate and uses that input to generate a specific output, such as the volume density and / or view-dependent emission at that coordinate. Then, using the output volume density and the corresponding RGB color of the view-dependent emission, a 2D projected image of the 3D scene is constructed through volume rendering operations. Therefore, the neural field can be used as a locally implicit image function to represent an image as a set of latent codes, thereby predicting the RGB value of a given pixel at a given coordinate or spatial location (x, y) and / or predicting 2D features.
[0011] The techniques described herein can be implemented to provide a novel framework that uses neural fields (e.g., as implicit functions, etc.) to model or transform one or more sequences of event camera data generated from one or more event cameras at one or more viewpoints in relation to a physical environment or visual scene. Each of the event camera data sequences may include sensing data collected from the sensor elements of the corresponding event camera in the one or more event cameras.
[0012] In some approaches, neural networks can act as filters to accept or use data captured by one-dimensional or two-dimensional (1D / 2D) sensors as input. The neural network can then generate the filtered result as output.
[0013] In contrast, the neural field described in this paper accepts or employs coordinate or spatial location information associated with the signal (e.g., event camera data sequences generated by an event camera) as input. By training with data captured by sensors (such as polarity in the event sequence) as real data, the neural field memorizes the signal or event sequence itself along with the coordinate or spatial location information. The neural field can then be used to predict or generate reconstructed event sequences that approximate the event sequences used as input to the neural field and the real data.
[0014] A neural field can be referred to as an implicit neural function, which contains hidden layers that operate with weights, biases, operating parameters, and activation functions. A trained neural field, or a fitted implicit neural function represented therein, maps an input sequence of events to a reconstructed sequence of events.
[0015] Under the techniques described herein, neural fields can be used to implement (e.g., backbone, non-backbone, etc.) multimedia encoders or codecs that encode and / or represent event camera data using trained or fitted operating parameters employed in the neural field. These operating parameters can be passed from an upstream device to a downstream receiving device. The downstream device can use the trained or fitted operating parameters to operate the corresponding neural field to predict or generate reconstructed events that approximate those represented in or derived from the event camera data with relatively high accuracy. Neural field multimedia codecs, as described herein, can help achieve relatively high rate-distortion (RD) performance and relatively high compression performance during encoding operations.
[0016] In some example operational scenarios, neural fields, as described in this paper, can be implemented using multilayer perceptron (MLP) networks. MLP networks can be trained or fitted relatively efficiently through machine learning-based training to describe, encode, and / or represent multimedia data (such as the trained or fitted operational parameters of the MLP network or its layers), which can then be used to predict or generate reconstructed event camera data corresponding to event camera data collected using one or more event cameras.
[0017] Compared to deep learning architectures that take raw multimedia data as input and perform filtering-like operations, MLP networks can take coordinate or spatial location information related to or embedded with event camera data as input and memorize or retain coordinate or spatial location information in the output (reconstructed event) generated by the MLP network.
[0018] In some approaches, MLP stands for regressor or implicit continuous function, which outputs predicted output values (e.g., continuous brightness or RGB values) covering all continuous values within a range or space of values (e.g., all byte values, all integer values, etc.).
[0019] In contrast, MLP networks, as described in this paper, operate or function as classifiers (e.g., for predicting the positive or negative polarity of brightness variations), rather than as regressors. While other layers of an MLP network can output continuous values, the last layer (or output layer) of an MLP network uses an activation function that outputs a finite number (e.g., 2, 3, etc.) of discrete values, representing appropriate subsets of all continuous values within a range or space of values, or selected values from all byte values, all integer values, etc.
[0020] In some operational scenarios, neural fields or MLP networks, as described herein, can be used to model aggregated event camera data sequences generated by aggregating raw or pre-aggregated event camera data sequences collected from event cameras. Aggregated events can be generated or aggregated from raw events in the event camera data, at least in part, based on user input specifying the time interval for aggregation (e.g., based on timestamps, etc.). Additionally, optionally, or alternatively, neural fields can model multiple event camera data sequences based on multiple viewpoints of multiple event cameras or camera elements relative to the physical environment or visual scene.
[0021] In some operational scenarios, neural fields or MLP networks, as described in this paper, can be used to model one or more raw or pre-aggregated sequences of event camera data collected from a physical environment or visual scene by one or more event cameras. Additionally, optionally or alternatively, neural fields can model non-event data corresponding to coordinates or spatial locations in the raw event camera data that are not represented by the original events.
[0022] A neural field (along with trained or fitted operating parameters, including but not limited to weights, biases, activation function parameters, etc.) can be viewed as an implicit function that logically approximates or provides a representation of aggregated or raw event camera data used as input to train or fit the neural field or operating parameters (e.g., to generate their optimized values, etc.). Under the techniques described herein, compression of aggregated or raw event camera data can be implemented or accomplished by compressing the neural field or by compressing the operating parameters that define or specify the neural field.
[0023] In some operational scenarios, event location data used to train neural fields can be encoded in a bitstream along with optimized values of operational parameters.
[0024] Alternatively or alternatively, while the optimized values of the operating parameters can be represented in relatively high-precision numerical formats such as floating-point or double-precision, they can be quantized to relatively low-precision numerical formats (compared to high-precision formats) such as integers (e.g., linear quantization, nonlinear quantization, etc.). The quantized optimized values of the neural field's operating parameters can be encoded in a bitstream instead of the pre-quantized optimized values used during training.
[0025] The example embodiments described herein relate to encoding event camera data using neural field operating parameters. Event camera data having a raw event sequence is received. Each raw event in the raw event sequence is generated by a specific sensor element of a plurality of sensor elements of an event image sensor in response to a change in brightness. Each sensor element of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions. Neural field training data is generated from the raw event sequence in the event camera data. The neural field training data includes a plurality of training instances. Each training instance of the neural field training data includes the event pixel position coordinates of the event represented in the training instance and the ground truth classification category. During the training phase, optimized values of the operating parameters of the neural field are generated by training the neural field using the neural field training data, at least in part based on a specific loss function chosen for the neural field. The neural field is trained to output corresponding predicted classification categories for pixel positions among a plurality of pixel positions at a plurality of time points. The neural field operating parameter values are encoded in an encoded bitstream. The neural field operating parameter values are generated from the optimized values of the neural field operating parameters. The encoded bitstream enables the receiving device to use a neural field operated on with neural field operation parameter values to generate multiple reconstructed events that approximate multiple events represented in the training data.
[0026] The example embodiments described herein relate to decoding event camera data using neural field operating parameters. The neural field operating parameter values are decoded from an encoded bitstream generated by an upstream device. The neural field operating parameter values are generated from optimized values of the neural field's operating parameters. These optimized values are generated by the upstream device training the neural field using neural field training data, at least in part based on a specific loss function chosen for the neural field. The neural field is trained to output predicted classification categories. The neural field training data includes multiple training instances. Each training instance of the neural field training data includes the event pixel location coordinates of the event represented in the training instance and the ground truth classification category. The neural field training data is generated from a raw event sequence in the event camera data. Each raw event in the raw event sequence is generated by a specific sensor element of a plurality of sensor elements of an event image sensor in response to a change in brightness. Each sensor element of the event image sensor corresponds to a corresponding pixel location among a plurality of pixel locations. The neural field, operated with the neural field operating parameter values, is used to generate multiple reconstructed events that approximate the multiple events represented in the training data.
[0027] In some example embodiments, mechanisms as described herein form part of a media processing system, including but not limited to any of the following: handheld devices, game consoles, televisions, laptop computers, netbook computers, tablet computers, desktop computers, computer workstations, computer kiosks, or various other types of computing devices and media processing units.
[0028] Various modifications to the embodiments and general principles, as well as the features described herein, will be readily apparent to those skilled in the art. Therefore, this disclosure is not intended to be limited to the illustrated embodiments, but rather to be accorded the maximum scope consistent with the principles and features described herein.
[0029] 2. Neural Fields and MLP Networks Figure 1A An example multilayer perceptron (MLP) network that can be used to implement neural fields as described herein is illustrated. An MLP network may include MLP layers, such as an input layer, an output layer, and one or more hidden layers between the input and output layers. These MLP layers are fully connected—each neuron in a subsequent layer of an MLP layer is connected to all neurons in the immediately preceding layer (if any) and thus receives the output generated from all neurons in the immediately preceding layer.
[0030] An activation function is used in each neuron of a hidden layer in an MLP network. The activation function of a neuron in a given layer of an MLP network takes the outputs of all neurons in the immediately preceding layer (relative to the given layer) as input and generates an output.
[0031] The MLP network described in this paper does not act as a regressor or implicit function (which generates continuous output values covering all continuous values within a specific range or value space), but rather as a classifier (which generates a sparse or appropriate subset of values within a specific range or value space (e.g., all byte values, all integer values, etc., as opposed to all continuous values)). In some operational scenarios, some or all of the input or hidden layers preceding the output layer in an MLP network can use activation functions that generate continuous output values. In contrast, the last layer (or output layer) in an MLP network as described in this paper uses an activation function that outputs discrete, sparse values (not covering all continuous values within a range or value space, or covering selected values from all byte values, all integer values, etc.).
[0032] like Figure 1A As illustrated, the input layer of an MLP network consists of neurons (represented by circles) that receive inputs from the neural field or MLP network (e.g., input0, input1, etc.) and generate intermediate outputs that are received as inputs by one or more hidden layers or neurons. The output layer of an MLP network consists of neurons (represented by circles) that receive outputs from neurons in the immediately preceding layer (relative to the output layer) as inputs and generate predictions (or MLP predictions) as outputs from the neural field or MLP network (e.g., output0).
[0033] As described in this paper, different layers of an MLP network can have different total numbers of neurons. Deep or multi-layer MLP networks, which include multiple MLP layers, can be constructed by sequentially stacking or arranging multiple building blocks of multiple MLP layers within them.
[0034] Figure 1B The illustration shows an example basic building block of neurons in an MLP layer of an MLP network. Each of the other neurons in the same MLP layer and / or in other MLP layers may have the same or similar structure.
[0035] As shown, neurons in an MLP layer receive multiple inputs (fully connected) and generate outputs, which in turn are generated as outputs from all neurons in the immediately preceding layer. A neuron may include a linear operation followed by a rectified linear unit (ReLU) activation function.
[0036] Linear operations can receive input x 1 、x 2 、x 3 、……、x m, where m represents an integer not less than one (1), which represents the total number of neurons in the immediately preceding layer. Linear operations multiply these inputs by the corresponding weights (or weight factors). w 1 、w 2 、 w 3 、……、w m (They can be collectively referred to as) This generates a weighted sum. Linear operations further add a bias (or bias factor) to the weighted sum. To generate or output the output value of a linear operation. x ReLU activation function f(x) (For example, max(0, x) (etc.) receives the output value from the linear operation. x As input, and based on the input received from the previous linear operation x Generate output value .
[0037] More generally, the operating parameters (such as those used in forming MLP networks or neural fields) The first MLP layer The weights and biases at each MLP layer (e.g., each neuron in that layer), where, (representing an integer not less than one (1)) is denoted as Weight ( )and bias ( ), which can be collectively referred to as For simplicity, the neuron indexes have been omitted, but it should be noted that different weights and biases can be used for different neurons in the same MLP layer.
[0038] During the model training phase, these operational parameters can be trained or fitted to optimize or minimize the prediction error, such as that measured or calculated by the loss function chosen for the MLP (or neural) network implementing the neural field. (From the...) The output values of an MLP layer or its output neurons can be generated or adjusted by activation functions such as the Rectified Linear Unit (ReLU). It should be noted that for the output layer, different activation functions (such as the Softmax function) can be used to generate discrete, sparse values representing the output of the entire MLP network as a classifier. The ReLU activation function can be given as follows: (1) In some operational scenarios, the final output value of the MLP network (or neural field) can be represented or restricted to a specific range of values (e.g., [0,1]). A sigmoid layer can then be added as the final layer of the MLP network, which squeezes each real number from the immediately preceding layer (input) of the MLP network (e.g., the output from the last ReLU activation, etc.) into a specific range between 0 and 1. The activation function of the sigmoid layer can be given as the sigmoid function as follows: (2) Therefore, neural fields can be implemented using an end-to-end MLP network, which consists of... It consists of several MLP layers, each with its own operating parameters. Weight and bias or This MLP network uses input x and the following output : (3) Given real data signals (e.g., based on samples from event camera data, etc.) The formal problem statement used to fit or train an MLP network or neural field (e.g., during the training phase) can be defined or specified as optimizing operational parameters. The value is minimized by the following loss function : (4) During the model training or fitting phase, location data in the training data (e.g., pixels or pixel locations, sensor element placement, etc.) can be used as input to an MLP or neural field to generate polarity predictions for different locations. A (single) loss function can be used to measure the prediction error between the predicted polarity in the training data and the true data polarity. Given that the activation functions in the MLP network or neural field are continuous and differentiable, gradients or derivatives can be computed at least in part based on the prediction error, and these gradients or derivatives can be backpropagated to iteratively adjust the weights and biases until the prediction error measured by the loss function is minimized (e.g., below a prediction error threshold) and / or until all training instances of the training data have been processed. Various candidate loss functions can be considered as the loss function for training the MLP network or neural field. In some operational scenarios, the loss function can be specifically selected from candidate loss functions, or used to measure the accuracy of predicting event polarity and / or non-events by referencing the event polarity generated by the sensor and / or the non-events implied by the events, or to compute the error of predicting event polarity and / or non-events. In some operational scenarios, task-specific (e.g., application-specific, task-oriented, application-specific, etc.) loss functions can be selected from candidate loss functions or from downstream applications or tasks of a specific type used to support the reconstruction of events. Different loss functions of different types can be used as task-specific loss functions for different downstream applications or tasks. Examples of downstream applications or tasks described herein could include image / video generation, object recognition, etc. By way of illustration, and not limitation, in some operational scenarios where events are used to generate or reconstruct images or videos at a relatively high frame rate, the reconstructed image (or image frame) from the reconstructed event can be used with a task-specific loss function that prioritizes the overall quality of the image rather than the prediction accuracy of individual events. Similarly, in some operational scenarios where events are used to identify objects depicted in an event with relatively low latency, the accuracy of object recognition and classification can be measured using a task-specific loss function that prioritizes the overall quality of object recognition rather than the prediction accuracy of individual events.
[0039] In the model application or inference phase, location data representing different locations can be used as input to a neural network that operates to optimize the values of its operating parameters in order to generate or infer the predicted polarity of these locations.
[0040] These training / fitting and application / inference phases can be repeatedly executed using different event sequences and / or different event camera datasets covering different time periods (e.g., different hours, days, weeks, etc.).
[0041] 3. Location coding At least in part, this is based on downstream or subsequent media or image rendering operations derived from the output of a neural field or MLP network used to process event camera data. Position encoding and / or a sinusoidal activation function (SIREN) can be implemented to reduce or prevent the loss of high-frequency details in these media or image rendering operations, thereby improving the neural field or MLP network's representation of the scene.
[0042] In location coding, the input to a neural field or MLP network is mapped to a relatively high-dimensional space that includes dimensions representing relatively high spatial frequency components in coordinates or spatial positioning / location data. This can be used to correct the bias that neural networks tend to learn relatively low-frequency functions or components represented in event camera data, and to enable the neural field or MLP network to represent or learn relatively high (e.g., coordinates, spatial positioning / location, etc.) frequency variations in color and geometry / coordinates / position / location.
[0043] Therefore, by using position coordinates From the one-dimensional real value domain R mapping To the 2L-dimensional real domain This can significantly improve the performance of neural field or MLP networks, among which, L This can be represented, for example, by implementing position encoding (or coordinate mapping) using the Sinusoidal Activation Function (SIREN). The total number of frequencies generated. Acting on coordinates. Location encoding (or coordinate mapping) It can be represented as follows: (5) in, It is an integer. In the example setup, .
[0044] 4. Event Data Event cameras are asynchronous (physical) sensors that represent scenes differently from other types of image sensors, such as CCDs. Event cameras can generate event camera data with the following characteristics: relatively high sensitivity to low light and changes in light in the physical environment or visual scene; relatively high temporal resolution; relatively low latency (microseconds); relatively (e.g., very high) dynamic range (e.g., a range of 140 dB, compared to the 60 dB range of cameras with different types of image sensors); and relatively low power consumption.
[0045] Non-event cameras can acquire complete images at an image acquisition or refresh rate specified or referenced by an external clock (e.g., 30 fps). In contrast, event cameras, or event sensors within them (such as dynamic vision sensors (DVS)), respond instantly, asynchronously, and independently to changes in brightness or luminance in the physical environment or visual scene. A per-pixel reference logarithmic intensity (e.g., for received photons or light) can be maintained for each pixel or sensor element in the event camera or sensor. If the magnitude of the logarithmic intensity change exceeds a maximum (logarithmic) intensity change threshold for positive or negative changes, the pixel or sensor element excites or outputs a binary output / event. Mathematically, this event excitation process can be expressed as follows: (6) in, (7) in, Indicates pixel position and time The intensity at a point (e.g., the intensity of a photon or light). It indicates the time elapsed since the last event at the same pixel location or sensor element. Indicates pixel position The intensity (e.g., the intensity of photons or light). Furthermore, This represents the threshold for the maximum (logarithmic) intensity change of a positive value, and The threshold representing the maximum (logarithmic) intensity change of negative values. Here, Output or represent the position of a given pixel and a given time point The polarity of the event data or (logarithmic) intensity value change, which can be one of the values 1, -1, and 0. If If no event is generated, then no event is generated. Regardless of whether an event is generated for a pixel position or sensor element at a given time point, for simplicity, the polarity of the pixel position or sensor element at a given time point can be represented as... Events related to pixel location or sensor element ( This can be either not stored or explicitly represented in the event camera data or its stream, and can be implicitly assumed to be an event even if there is no positive or negative event at these pixel locations or sensor elements. Here, k This represents an event index, for example, used to identify or sort events belonging to the same event sequence. Different indexes k can refer to multiple events that occur at the same time or at different times.
[0046] Events at pixel locations or sensor elements (If excited) can be defined as a 4-parameter tuple This contains coordinate information about its pixel location. Precise trigger time and 1 polarity In some variants of event cameras / sensors, the event also includes information about or specifies how much the intensity has changed, rather than simply indicating whether the brightness or luminance of the pixel that was triggered by the event has increased / decreased. It is inferred that other pixel locations have the same intensity value as before, and therefore are not represented in the event camera data or its stream.
[0047] Note that event camera data or streams can consist of sparse points or pixels, as it can explicitly represent (or contain information) only these sparse points or pixels, at each of which the corresponding sensor element detects or senses an intensity change that is sufficiently large compared to the maximum intensity change threshold.
[0048] Event camera data can be generated by the event camera and is based on an initial time (point). With the final time (point) The sequence of events between (represented as) ) represents or is defined as . and Time points within the interval between It can also be called a time instance. Similarly, an event sequence can also be called an event stream.
[0049] Due to the advantages of event cameras / sensors, event camera data generated by these cameras / sensors and modeled / transformed using neural fields or MLP networks with the techniques described herein can be used for a variety of applications, including but not limited to video games, AR / VR / MR, object and / or image rendering, applications for real-time interaction such as robotics, applications requiring robust performance under relatively low power, low latency, and varying lighting conditions, surveillance, object segmentation, object detection, object tracking, 2D / 3D sensing of physical environments, gesture recognition, object recognition, optical flow estimation, HDR image reconstruction, robotic applications such as SLAM (Simultaneous Localization and Mapping), image reconstruction, image deblurring (under relatively extreme lighting conditions and / or relatively fast motion), astronomical research, etc.
[0050] 5. Neural Field Coding Framework A neural field coding framework can be implemented to model or transform event camera data using neural fields. This framework includes, for example... Figure 1CThe illustrated encoder-side architecture, implemented by an upstream device, uses event camera data as input to generate optimized values for the operating parameters of a neural field or MLP network, compresses or encodes these optimized values into an encoded bitstream, and transmits or otherwise delivers the encoded bitstream to a downstream receiving device. This framework further includes, as shown... Figure 1D The illustrated decoder-side architecture, implemented by a downstream receiver device, receives an encoded bitstream directly or indirectly from an upstream device, decodes or decompresses optimized values of the operating parameters of a neural field or MLP network from the encoded bitstream, and applies these optimized values to generate outputs from the neural field or MLP network. These outputs can be used for a wide variety of applications, including image generation or rendering, object segmentation, object detection, object tracking, AR / VR / MR, real-time or non-real-time applications, etc.
[0051] like Figure 1C As illustrated, the encoder-side architecture receives event camera data as input and uses the neural field or MLP network to represent the event camera data after training or fitting the neural field or MLP network with the event camera data.
[0052] In some operational scenarios, event camera data used to train or fit neural fields or MLP networks can be represented or specified as (aggregated) event camera image sequences at time points covering a certain time interval or duration, such as videos or video clips.
[0053] An event camera image sequence may include event camera images from a single viewpoint, or event camera images from a single viewpoint at different points in a time sequence or for different points in time. An event camera image sequence may also include a single image from a single viewpoint at a corresponding point in a time sequence.
[0054] Additionally, optionally, or alternatively, an event camera image sequence may include event camera images from multiple viewpoints, or event camera images from multiple event cameras (or camera elements) originating from multiple viewpoints at different times in a time-point sequence or for different times. An event camera image sequence may include multiple images from multiple viewpoints at corresponding times in a time-point sequence.
[0055] In some operational scenarios, event camera data used to train or fit neural fields or MLP networks can be represented or specified as (raw) event camera data sequences at time points covering a certain time interval or duration. To better model event camera data, raw non-event data, which may not be explicitly represented in the original event camera data sequences, can be generated or provided as input to train or fit neural fields or MLP networks.
[0056] After training or fitting with input event camera data, a neural field or MLP network provides a representation (or neural field encoding) of the event camera data through optimized values generated for the weights, biases, and / or other operating parameters of the neural field or MLP network. These optimized values of the operating parameters of the neural field or MLP network can be encoded or compressed (using neural network compression) into an encoded bitstream (or neural network bitstream), which can be transmitted from an upstream device implementing an encoder-side architecture or otherwise passed to a downstream receiving device implementing a decoder-side architecture.
[0057] like Figure 1D As illustrated, the decoder-side architecture receives an encoded bitstream (or neural network stream) as input and decompresses or decodes the bitstream (using neural network decompression) to obtain or generate optimized values for operating parameters (or neural field coefficients) (such as weights, biases, etc.). A specified user can specify user input, including but not limited to user input related to event camera data (such as timestamps, viewpoints, etc.). Based at least in part on the user input, the decoder-side architecture performs neural field decoding to generate a reconstructed version of the event camera data (or reconstructed event data) that is identical to or closely approximated by the input event camera data on the encoder-side architecture. The neural field at the decoder can be implemented using an MLP that can be used to reconstruct the event camera data.
[0058] 6. Modeling aggregation events using neural fields As mentioned earlier, events represented or included in event camera data It can be defined as a 4-parameter tuple To help preserve high spatial frequency events or image features represented in event camera data, location encoding can be used. Event coordinates used to represent or include in event camera data .
[0059] Given these coordinates and a given timestamp Neural fields can utilize MLP networks (represented as...) This is implemented by training or fitting the network with events from event camera data to predict or output coordinates in the reconstructed domain (or the reconstructed event camera data domain) and the same given timestamp. The polarity of the timestamp. coordinates of The predictive polarity of the event can be given or represented as follows: (8) In some approaches, MLP networks acting as regressors (e.g., predicting all possible values within a range, predicting all possible color values, etc.) can be used to implement neural fields. In contrast, in techniques as described herein, MLP networks or neural fields are modified or implemented to act as classifiers (e.g., two or more classification categories, the polarity of brightness or luminance variations, etc.). MLP networks or neural fields can be used to predict the polarity of events represented in event camera data.
[0060] In some operational scenarios, neural fields or MLP networks are implemented as multi-class classifiers. This can be achieved by replacing the sigmoid activation function in the neural field block or MLP layer with a softmax function. The softmax function can be represented as follows: (9) in, It is a class from a previous (e.g., MLP, etc.) layer. The predicted polarity value, and It is the number of classes.
[0061] The Softmax function takes data from previous layers. A dimensional vector is formed, and each component of the vector is transformed into a range of values. The real number in the vector is such that all components of the vector sum to one (1).
[0062] MLP networks modified or implemented as classifiers (For example, using the Softmax activation function instead of the sigmoid function as the activation function in this MLP network (e.g., each of the MLP layers)) The input is represented in the input event camera data domain. Output classification categories in the reconstructed event camera data domain. .enter The specific representation depends on the specific use case. (From MLP network) Output classification categories It can be represented as follows: (10) Given the polarity of the real data generated or represented in the input event camera data. The formal problem of obtaining the optimization values of the operating parameters (such as weights, biases, etc.) of a (modified) MLP network can be expressed as follows: (11) in, This represents the loss function to be minimized in order to obtain the optimized values of the operating parameters. Since the MLP is to be used as a classifier rather than the MSE loss function as a regressor, the cross-entropy loss function can be used in expression (11) and expressed as follows: (12) in, The Softmax function is applied to the vector received from the previous (MLP) layer. i The output of each component; yes i-th The target polarity of each component; and It represents the number of classes or components in the vector. To simplify event camera data modeling, the polarity from the raw event camera data... Mapped to , making Corresponding to and Corresponding to .
[0063] 7. Reconstruct the aggregated events. The techniques described herein can be used to implement deep learning algorithms that can be applied to aggregated events in synchronized frames. These aggregated events can be generated from asynchronous events included or represented in raw event camera data produced using one or more event cameras.
[0064] Given event camera data as an event dataset, as mentioned earlier, the event sequence of the entire dataset can be given or specified as follows: In order to time intervals The events in the aggregated event sequence or dataset, for each event pixel position (which refers to the time interval) T (pixels or pixel locations where at least one event occurs during the time interval) T Select the newest or last event from all events occurring at the event pixel position during the period. The selected events can be called aggregate events, denoted as... It can be given or represented as follows: (13) in, Indicates the latest time interval and event pixel position The polarity of the selected event (e.g., +1 or -1, etc.), where the event pixel position corresponds to pixel coordinates. .
[0065] Pixel positions that are not event pixel positions or do not belong to any event pixel position in the aggregated event set in the above expression (13) can be assigned a third polarity - value 0.
[0066] Assuming the event / image frame includes all pixel locations where the sensor element wants to detect or trigger a polarity event, it has a width and height The following is a non-event dataset that can determine or identify pixel locations in an event / image frame where sensor elements did not detect or trigger polarity events: (14) in, This represents the set of all event pixel positions within the represented time interval; for non-event pixel positions... The polarity is always 0. Non-event pixel locations are located within the boundaries of the event / image frame and do not correspond to or belong to any event pixel location within that given time interval. A given time interval can be represented by a reference timestamp of a reference time point within that given time interval.
[0067] The final or combined dataset (containing event or polarity information for all pixel locations of an event / image frame) used as input to train a neural field or MLP network on aggregated events / content can be represented or given as follows: (15) For illustrative purposes only, it has been described that neural fields can be used to model aggregated events using the last event captured in event camera data for the corresponding event pixel location. However, it should be noted that in other operational scenarios, neural fields can also be used to model event camera data that has already been aggregated using other techniques such as temporal surfaces (TS), event voxels, etc. Under the temporal surface technique, an event stream, or the events within it, can be converted into a spatiotemporal point cloud. A temporal surface can be generated for events in the point cloud (e.g., each, etc.) by applying an exponentially decaying kernel (or function) to the pixel neighborhood of the specific pixel location corresponding to or in which the event is located. Therefore, the temporal surface can be used to interpret the event and the dynamic spatiotemporal context surrounding the event, which provides information about the history of activity in the neighborhood. The temporal surface generated from the event provides a temporal surface representation of the event and can be used as input to a neural field as described herein to generate predicted events. Under the event voxel technique, an event stream can be converted into a fixed-size tensor representation, where events can be encoded in a spatiotemporal voxel grid. The duration spanned or covered by the event. T = t k N 1 t k0 can be discretized into B time chambers. Each of these events can have its polarity distributed to its two nearest spatiotemporal voxels in a spatiotemporal voxel grid. The voxels containing the events provide the event voxel representation of the events and can be used as input to a neural field as described in this paper to generate predicted events. Example event aggregations utilizing time surfaces and event voxels are described in the following literature: Yiheng Xie et al., “ Neural Fields in Visual Computing and Beyond [Neural Fields in Visual Computing and Other Fields], Eurographics / CGF State-of-the-Art Report (2022); Lagorce et al., HOTS: A hierarchy of events based time surfaces for pattern recognition [HOTS: Event-based temporal surface hierarchy for pattern recognition], IEEE TPAMI (2017); and Rebecq et al., " Events-to-video: Bringing modern computer vision to event cameras [Event to Video: Applying Modern Computer Vision to Event Cameras], CVPR (2019), all of which are incorporated herein by reference in their full text.
[0068] As a first use case, events in event camera data can be aggregated at given time intervals of 50 ms. An event sequence in the event camera data (e.g., the entire sequence) can be partitioned into event subsequences with time intervals of 50 ms. Each of these event subsequences, covering a corresponding time interval of 50 ms (within the total duration covered by the event sequences in the event camera data), can be represented as follows: , Etc. The original or pre-aggregated events in these event subsequences with a time interval of 50 ms can be aggregated into aggregated events, denoted as... , Etc. A corresponding time interval of 50 ms can be used (e.g., between...). and Corresponding timestamps between (e.g.) (For example, (etc.) to process each aggregate event To identify or index.
[0069] Figure 2AThe illustration shows an example image containing aggregated events, generated by aggregating (raw or pre-aggregated) events from (input) event camera data produced by the event camera within 50ms. Each dark or gray value or pixel represents a "-1" polarity (indicating a decrease in brightness or luminance) or a "+1" polarity (indicating an increase in brightness or luminance), while each white value or pixel represents a "0" polarity (indicating no change in brightness or luminance).
[0070] As used as Figure 3A The illustration shows an example of a neural field or MLP network modeling aggregated events generated from camera data of the same event, consisting of ten event frames (each with a corresponding timestamp). Each in the index includes an aggregated event generated by aggregating (raw or pre-aggregated) events from event camera data within corresponding 50 ms time intervals in a 50 ms time interval sequence (e.g., consecutive, sequential, etc.). The aggregated events in the ten event frames can be used to train or fit a neural field or MLP network and generate optimized values for operating parameters (such as weights, biases, etc.) used in the neural field or MLP network.
[0071] Correspondingly, ten reconstruction event frames (also marked with corresponding timestamps) Each of the indexes includes predicted aggregate events generated or predicted by a neural field or MLP network operating with optimized values of operating parameters within corresponding 50 ms time intervals in a 50 ms time interval sequence (e.g., continuous, sequential, etc.).
[0072] The accuracy of predictions made by neural fields or MLP networks can be calculated as follows: (16) It can also be expressed as (17) TP represents true positive; TN represents true negative; FP represents false positive; and FN represents false negative.
[0073] like Figure 3A As illustrated, in the MLP network used to implement the neural field, solid black arrows represent affine transformations and ReLU activation functions, while dashed arrows represent, for example, affine transformations and non-ReLU functions. An affine transformation can be described as a linear transformation with added bias. The numbers in each block describe fully connected layers. The output from the MLP network consists of three (3) values corresponding to three (3) classification categories: +1, -1, and 0. The input to the MLP network is represented by a 41-dimensional vector, with each coordinate... ,application Position encoding This yields 40 values and a timestamp for the indicator vector. The added value of ).
[0074] Figure 2B The illustration shows an example reconstructed event frame and prediction accuracy, which is based on the aggregated prediction events generated from a neural field or MLP network and the same time interval of 50 ms (where the timestamp is...). dt = 1 The calculation is performed by comparing aggregated events generated or aggregated from event camera data within the specified time interval. Similarly, additional reconstructed event frames within other time intervals (with other timestamps) can be generated from predicted aggregated events from neural fields or MLP networks with corresponding computational prediction accuracy.
[0075] Alternatively, or alternatively, events in the event camera data can be aggregated into event frames using time intervals other than 50 ms to verify the general applicability of the data modeling operations described herein. For example, time intervals of 10 ms, 1 ms, etc., can be used to aggregate events in the event camera data. Compared to relatively large time intervals such as 10 ms or 50 ms, a slight decrease in accuracy (e.g., approximately 97%) can be observed with relatively small time intervals such as 1 ms. This is likely due to class imbalance, where the number of pixels without any polarity (pixels belonging to non-event pixel locations) far exceeds the number of pixels with polarity (pixels whose brightness or luminance increases or decreases). This class imbalance can make neural field modeling more challenging. Example solutions to address the class imbalance problem will be described in further detail later.
[0076] In some operational scenarios, post-training quantization (e.g., static, dynamic, linear, nonlinear, etc.) can be combined with neural fields or MLP networks to apply to operational parameters or their optimized values. For example, weights, biases, and / or activation function parameters can be internally represented as relatively high-precision numerical values, such as floating-point values. These operational parameter values can be quantized to relatively low-precision numerical values, such as integer values. Quantized values can be transmitted as signals from upstream devices to downstream receiving devices or included in the coded bitstream, enabling downstream devices to decode the quantized values of the operational parameters of the neural field or MLP network and use the neural field or MLP network operated with the quantized values of optional parameters to predict aggregated events in the reconstructed event camera data domain.
[0077] By using examples rather than limitations, such as generating aggregated events from input event camera data ( Figure 2AThe diagram shows some or all of the example event frames from these aggregated events. The size of the relatively high-precision operating parameter values of the trained neural field can be 0.97 MB. After applying quantization, the size of the relatively low-precision operating parameter values or quantized operating parameter values of the neural field or MLP network is reduced by 0.27 MB.
[0078] Figure 2C The illustration shows an example reconstructed event frame and prediction accuracy, which is based on the aggregated prediction events generated from a neural field or MLP network and the same time interval of 50 ms (where the timestamp is...). dt = 1 The calculation is performed by comparing aggregated events generated or aggregated from event camera data using quantization operation parameters.
[0079] Compared to neural field or MLP networks that operate with relatively high-precision (pre-quantized) operating parameter models, neural field or MLP networks that operate with quantized operating parameters result in a decrease in (prediction) accuracy.
[0080] As mentioned earlier, in some operational scenarios, event camera data can include multiple (polarity, sensor element excitation) event sequences generated by a system of multiple event cameras for multiple viewpoints. These multiple event cameras can be spatially located and / or oriented differently from each other in the physical environment or visual scene.
[0081] Using the techniques described in this paper, such event camera data can be modeled using neural fields or MLP networks. Events in the event camera data It can be defined or represented as a 4-parameter tuple It can be used to analyze event coordinates in event camera data. Application location coding Given the coordinates of a pixel position timestamp and perspective Neural fields or MLP networks can be used to predict the polarity of pixel locations in the reconstructed event data domain, as shown below: (18) in, Indicates the timestamp and perspective The location has coordinates The predicted polarity of the pixel position.
[0082] In some operational scenarios, event camera data can include events from two perspectives: a left (or left-eye) perspective and a right (or right-eye) perspective. These perspectives They can be represented as 0 and 1 respectively.
[0083] Each event sequence for each viewpoint can be divided or partitioned into time intervals of a specific length (e.g., 50 ms). Events within each time interval can be independently aggregated for each viewpoint (e.g., by the last occurrence of the event pixel location or its polarity) into aggregated events in the event / image frame. These aggregated events in the event / image frame can be used as event training data. Non-event or non-event locations (all pixel locations other than the event pixel location) in the event / image frame can be independently identified or determined for each viewpoint. The polarity of non-event locations can be assigned a specific value, such as zero (a), and included in the overall training data along with the aggregated events for the event pixel location.
[0084] During the training phase of the model (or neural field) implemented using upstream equipment, training data can be used to train or fit the neural field or MLP network to generate optimized values for the operating parameters of the neural field or MLP network (such as weights, biases, activation function parameters, etc.).
[0085] During the model application (or inference) phase, downstream receiving devices can use optimized operating parameters or their quantized versions to generate reconstructed aggregate events for individual pixel locations in the reconstructed image / event frame.
[0086] 8. Model the original event. The techniques described herein can be used to implement deep learning algorithms that can be applied to raw event data or the (sensor-generated / raw) events represented therein—without the need to aggregate these events using synchronization frames. In some operational scenarios, the (sensor-generated / raw) events in the raw event data contain only information about two (2) polarities: a positive polarity of increasing brightness / luminance and a negative polarity of decreasing brightness / luminance. As previously stated, these two polarities can be derived from values and express.
[0087] Problem formulations for modeling raw events using neural fields can be similar to those for modeling aggregated events, as discussed earlier. For example, events in (input) event camera data. It can be specified or defined as a 4-parameter tuple ,in, k This represents the (event) index or identifier used to distinguish all events included in the event camera data. Location encoding. Can be applied to coordinates Given the specific coordinates of a pixel location and a specific time (or timestamp). The neural field or MLP network can be trained or fitted to predict the corresponding polarity of the event at this pixel location, as follows: (19) in, Indicates the timestamp coordinates The predictive polarity of events at a given location.
[0088] The design of neural fields or MLP networks for modeling raw events can be similar to that used for modeling aggregated events. Since it is represented as... The raw event stream (or raw event sequence in the event camera data) consists only of events corresponding to two polarities (e.g., +1 and -1) of brightness or light intensity. Therefore, neural fields or MLP networks can be adapted to output only two classification categories or only two polarities, such as... Figure 3B As shown.
[0089] This neural field or MLP network can be used to process raw event streams (or raw event sequences). Modeling can be performed, but it may not be possible to model at a given time (or timestamp). t Model the polarity of non-event or pixel locations where no event has occurred (e.g., 0 or values other than polarities +1 and -1).
[0090] As an example of using neural fields to model raw event data, assume the event camera data includes a raw event stream, such as a raw event sequence covering the duration of events at 10 ms intervals, and assume the time resolution of the event timestamps is in microseconds. The raw event stream can be represented as... The timestamp value of the event. t It can be normalized to the value range [0, 1] as follows: (20) in, It is the minimum value, such as 0, and This is the maximum value, such as 10,000 (microseconds). Therefore, the original event stream can be alternatively represented using normalized timestamp values. ,in, and For illustrative purposes only, we assume normalized timestamp values. and The total number of events within the duration or interval between events is 185,337. The total number of events in the original event stream or the corresponding dataset can be given as... .
[0091] Use with the original event stream The problem of using corresponding neural fields to predict (reconstructed) events in the reconstructed event domain can be formulated as follows: (twenty one) in, Indicates a given normalized timestamp In coordinates The predicted polarity of the event at the pixel location.
[0092] In model training phases, such as those implemented or executed by upstream devices, the raw event stream or events within it can be used as training data to train or fit a neural field or MLP network. Optimized values for the operational parameters used in the neural field or MLP network (e.g., weights, biases, activation function parameters, etc.) can be obtained by minimizing prediction errors, such as those measured using a loss function chosen for the neural field or MLP network, or by backpropagating these prediction errors. Optimized operational parameter values (e.g., unquantized, quantized, etc.) can be transmitted or passed directly or indirectly from the upstream device to the downstream receiving device.
[0093] In the model application (or inference) phase, for example, implemented or executed by downstream devices, neural fields or MLP networks that operate to optimize operating parameter values can be used to predict or generate reconstructed events at a given coordinate at a given time (or timestamp) within a duration or interval corresponding to the duration or interval covered by the original event stream.
[0094] The accuracy of predictions made by neural fields or MLP networks can be calculated as follows: (twenty two) In some operational scenarios, neural fields or MLP networks operating with unquantized or prequantized optimized operational parameter values can achieve 100% accuracy in predicting or generating reconstructed original events compared to using the original events from the (input) event camera data.
[0095] In some operational scenarios, applying post-training quantization (e.g., static quantization routines in PyTorch) to generate or obtain quantized models or quantized optimized operational parameter values for neural fields or MLP networks slightly reduces prediction accuracy to 99%, which still demonstrates that neural fields can be effectively used to model raw event data.
[0096] Because the neural field or MLP network has not yet modeled (or has not been trained or fitted with) the polarity values at non-event (pixel) locations, it is not configured to predict or generate reconstructed polarity values at these locations. In other words, the model has not learned a third polarity or third classification class in addition to the two polarities (or two corresponding classification classes) of the original events in the event camera data, and therefore cannot predict a third polarity or third classification class. Instead, the model can be trained or used to predict the two polarities of event (pixel) locations, such as +1 and -1. Therefore, the model can predict a specific polarity given the spatiotemporal location of the event (e.g., (event) pixel location coordinates, event timestamp, etc.). However, the predictions from the model for non-event (pixel) locations remain ambiguous.
[0097] 9. Modeling raw event data and non-event data. Spatiotemporal locations of events that do not exist in event camera data (e.g., pixel location coordinates, timestamps, etc.) can be referred to as non-events. To address the limitations of models trained or fitted only to predict raw events, neural fields or MLP networks can be adapted or augmented to model or predict both raw events and non-events.
[0098] Using neural fields to predict the reconstructed event domain and its relation to the original event flow. Corresponding (reconstructed) original events and events not in the original event stream The problem of non-events explicitly represented in the middle can be stated as follows: (twenty three) Neural fields can utilize, for example Figure 3A The MLP network shown can predict three polarities to be implemented. The training data used as input to the MLP network can be augmented or expanded to include both raw event data and (additional) non-event data. The raw event data includes raw events included or explicitly represented in the event camera data, and the (additional) non-event data includes non-events not included or explicitly represented in the event camera data.
[0099] Non-event data can be created or generated as follows: (twenty four) For non-event pixels, the polarity is always 0, and The pixel position of the event.
[0100] Therefore, non-event data or its corresponding stream can be represented or constructed as follows: (25) The final or overall training data used to train or fit a neural field or MLP network to model both the original event stream and non-events can be given as follows: (26) While it's possible to create non-events at each (non-event) pixel location for each timestamp represented in the original event, where the event doesn't exist in the original camera data, and include these non-events as part of the neural field's training data, such a training data set would become extremely large. Therefore, a much larger neural field would be needed to model the training data, which includes both original events and non-events.
[0101] Conversely, a selected non-event pixel location can be sampled from all (non-event) pixels or pixel locations in an image (at a given time or timestamp) that includes all raw events at a given time or timestamp, as well as all (non-event) pixels or pixel locations. In some operational scenarios, the neural field can be trained using training data that includes non-events at the selected non-event pixel locations and used to support interpolation in order to predict the polarity of other non-events at other (unselected) non-event pixel locations.
[0102] In some operational scenarios, the selected non-event pixel locations can be sampled as pixel locations placed in a grid-like spatial arrangement on the image. More specifically, the selected non-event pixels can be formed by a grid across the image ( Samples with uniformly distributed span width and (samples evenly distributed across the height) Composed of or selected from uniformly spaced points (or pixel positions), as follows: (27) in, W is the width of the image; H is the height of the image; and meshgrid() This indicates that it is used based on two vectors or dimensions (e.g.) X and Y To create a selection of non-event pixels or pixel locations. Functions that construct a grid (e.g., in Matlab).
[0103] The initially created non-event stream can include a subset of non-event pixel locations in the image that coincide with some event pixel locations among the event pixel locations, identified or selected. Therefore, for each timestamp... If those non-event coordinates correspond to event pixel positions, then the non-event stream can be updated or modified by removing these coordinates. ,as follows: (28) For each timestamp , Is the non-event pixel position and It is the pixel position of the event.
[0104] In some operational scenarios, neural fields trained or fitted using the overall training data (including the original events in the event camera data and selected non-events representing all timestamps (or time points / time instances) in the original events) can achieve relatively high prediction accuracy; for example, the accuracy is approximately 99.5% for all reconstructed events corresponding to the original events. Furthermore, for each timestamp, the neural field can achieve relatively high prediction accuracy for all reconstructed events corresponding to selected non-events arranged or placed on a grid across the image. For example, the overall prediction accuracy of the neural field for the original events and selected non-events at the grid locations can be approximately 99%.
[0105] Since the grid locations selected as non-event pixel positions can be fixed for each timestamp, the neural field is not trained outside of these fixed locations. Therefore, the neural field may tend to predict only the polarity at these fixed, selected non-event locations, and lack sufficient knowledge or understanding regarding how to predict or interpolate the polarity of unselected non-event locations across grid locations during training.
[0106] In some operational scenarios, instead of using a rule-based pixel grid to indicate the third polarity of non-events, for each timestamp... A random pixel grid is used for the third polarity. More specifically, the selected non-event pixels... It can be formed by grids across the image ( A sample with a width of random distribution and (a sample randomly distributed across heights) A number of points (or pixel positions) are composed of or selected from them, such that and and and Where H is the height of the image and W is the width of the image. As mentioned earlier, for each timestamp If non-event coordinates correspond to event pixel locations, then the initial non-event stream formed by the sampling points can be updated or modified by removing these coordinates. .
[0107] For illustrative purposes only, for each timestamp represented in the raw events in the event camera data, at each such timestamp, each pair across the X and Y dimensions of the image... Random sampling is performed on non-event pixel locations (either individually or in different ways). This overall training data, which includes the original events and the corresponding non-events on a randomly sampled grid representing all timestamps in the original events, is used to train or fit the neural field.
[0108] Figure 2D The illustration shows an example of reconstructed or predicted events and non-events from a neural field at timestamp = 0. In this image, dark pixels represent polarity -1 (suggesting a decrease in brightness or luminance) or +1 polarity (suggesting an increase in brightness or luminance), and gray pixels indicate polarity 0 (suggesting no change in brightness or luminance).
[0109] Across all timestamps, neural networks can achieve relatively high prediction accuracy (e.g., around 80%) at event pixel locations compared to regular sampling grids, and also improve prediction accuracy for non-event pixel locations.
[0110] In some operational scenarios, it can be observed that neural fields (e.g., in the case of randomly sampled grid points) still tend to aggregate data across different time stamps, and therefore cannot fully understand or learn the differences between different data portions across these different time stamps. The SIREN activation function or NeRV architecture (examples of which are in Chen et al.'s paper) NeRV: Neural Representation for Videos NeRV : [Video Neural Representation], described in NeurIPS (2021), and all these examples (which are cited in their entirety hereinafter incorporated herein by reference) can be used or implemented to improve neural fields for better learning and prediction across different time stamps.
[0111] To further understand or evaluate the effectiveness of randomly sampled pixels, only a relatively small subset of all timestamps (e.g., only three (3) distinct timestamps) can be considered or evaluated. And determine whether the neural field can model both events and non-events across these timestamps.
[0112] To make the evaluation more tractable, we consider only a relatively small subset (e.g., 40,000 samples) of all non-events in the image at a given timestamp, instead of including all non-event pixels for every timestamp. For example, if the image size is 640 × 480 pixels, and if there are 20 events in the image, then there are 307,180 non-events in the image for that timestamp. Therefore, 40,000 samples can represent only a small subset of the 307,180 non-events in the image for that timestamp.
[0113] For three (3) timestamps The total number of non-events can be given as follows: This number is approximate because non-event pixel locations that are in the same position as the event pixel location or whose coordinates coincide with the event pixel location are excluded from the overall training data.
[0114] The prediction accuracy of these timestamps exhibits a similar tendency: while modeling or predicting events at event pixel locations well (e.g., relatively accurately, etc.), neural fields may not model or predict non-events at non-event locations well. As previously mentioned, neural fields tend to aggregate data (e.g., aggregating data between or within adjacent timestamps, etc.). When considering or evaluating more timestamps (e.g., eleven (11) timestamps), And / or consider or evaluate different subsets (e.g., 1500 samples of images across all non-events at a given timestamp) and / or the different total numbers of non-events across all those timestamps. These tendencies can also be observed in certain situations.
[0115] Furthermore, to further understand or evaluate the power of randomly sampled pixels, a pixel from the image can be considered or evaluated, and a neural field can be used to model its polarity across all timestamps representing the original event in the event camera data. More specifically, given an initial time... With final time Specific pixel positions of all timestamps between All events and non-events to be modeled by a neural network can be represented as follows: (29) in, It is a fixed pixel position.
[0116] For example, given a specific pixel At the initial time = 0 and final time When there are 10,000 timestamps between 10,000, the event stream including the events can be represented as follows: For illustrative purposes only, the event flow... Only 24 events exist across 10,000 timestamps: .
[0117] Figure 2EThe illustration shows an example of modeling (time-varying) raw events and predicting events using a neural field. It can be seen that the neural field may fail to model these varying events. This problem may be caused by a class imbalance between the +1 and -1 classes or polarities represented in the raw events and the classes or priorities assigned to non-events, considering the pixel location... Place That is, the total number of non-events far exceeds the total number of original events.
[0118] To solve this problem, many different methods can be used to create higher-order class balances. In the first example, the event can be repeated 100 times, as follows:
[0119] (30) This leads to Because of this method for creating a higher degree of class balance, improvements can be observed in predicting the original events generated by the neural field compared to predictions made using training data with relatively obvious class imbalance problems.
[0120] In the second example, the event can be repeated 200 times, as follows:
[0121] (31) This leads to The class balance it creates is of a higher degree than that of the first example. Due to this method of creating a higher degree of class balance, even greater improvements can be observed in predicting the original events generated by the neural field, which can predict most events and is able to improve the prediction of non-events.
[0122] In the third example, the event can be repeated 500 times, as follows:
[0123] (32) This leads to An improvement can also be observed in predicting the original events generated by the neural field, which can predict most events and improve the prediction of non-events.
[0124] Therefore, the polarity variation of a pixel can be modeled across all time stamps. This can be translated to show that a neural field can be used to model the polarity variation of all pixels in an image. However, neural fields may still tend to model aggregated polarity data (and fail to capture some of the subtle variations within different time stamps). This aggregation problem can be mitigated through parameter tuning and / or through loss function selection.
[0125] 10. Compress raw event camera data using neural fields. Optimized values for the operating parameters can be used to define or specify a neural field of a selected structure, which is used to model the polarity of raw events or aggregated events generated from raw events in the event camera data. These optimized values for the operating parameters, along with the selected structure, logically provide a representation of the neural field and a representation of the event camera data used to generate training data for training or fitting the neural field. In some operating parameters, the representation of the neural field or event camera data can be compressed. The compressed representation can be explicitly or implicitly encoded into the encoded bitstream to be transmitted from the upstream device to the downstream receiving device.
[0126] For example, neural fields can be implemented using MLP networks. The selected structure of the neural field can be one of one or more different candidate MLP network structures. The selected neural field or MLP network structure can be determined or adopted at least in part based on evaluating and comparing different candidate neural field or MLP structures. The evaluation and comparison can be performed partially or entirely to balance the need for relatively high prediction accuracy with the need for relatively efficient data compression.
[0127] For illustrative purposes only, compression performance can be considered. Figures 3B to 3F Different candidate neural field or MLP network structures are evaluated or compared. Each of these neural field or MLP network structures can be implemented by a neural field or MLP network to model the same raw events in the same event camera data. These raw events can be included in a raw event stream with only two possible polarities. middle. Figures 3B to 3F Candidate neural fields or MLP networks have different operating parameter values and data sizes (which are related to the different total number of neurons used in the candidate neural field or MLP network (e.g., input, hidden, output, etc.) layers) and provide different accuracies when reconstructing events, as shown in Table 1 below.
[0128] Table 1
[0129] As illustrated in Table 1 (the leftmost two columns), Figure 3CThe first (candidate or evaluated) MLP network illustrated has a data size of 1702 KB and provides 100% prediction accuracy. Figure 3D The illustrated second MLP network (candidate or evaluation) has a data size of 1375 KB and provides 100% prediction accuracy. Figure 3B The illustrated third MLP network (candidate or evaluation) has a data size of 961 KB and provides 100% prediction accuracy. Figure 3E The illustrated fourth MLP network (candidate or evaluation) has a data size of 770 KB and provides a prediction accuracy of 99%. Figure 3F The illustrated fifth MLP network (candidate or evaluation) has a data size of 217 KB and provides a prediction accuracy of 92%.
[0130] The operating parameter values of these different MLPs can be quantified separately.
[0131] As illustrated in Table 1 (the two rightmost columns), Figure 3C The illustrated first (candidate or evaluated) MLP network, operated with quantized operation parameter values, has a data size of 464 KB and provides 100% prediction accuracy. Figure 3D The illustrated second MLP network (candidate or evaluation) operated with quantized operation parameter values has a data size of 378 KB and provides 100% prediction accuracy. Figure 3B The illustrated third MLP network (candidate or evaluation) operated with quantized operation parameter values has a data size of 277 KB and provides a prediction accuracy of 99%. Figure 3E The illustrated fourth MLP network (candidate or evaluation) operated with quantized operation parameter values has a data size of 226 KB and provides a prediction accuracy of 97%. Figure 3F The illustrated fifth MLP network (candidate or evaluation) operated with quantized operation parameter values has a data size of 75 KB and provides a prediction accuracy of 83%.
[0132] In some operational scenarios, in addition to including the (optimized or trained) operational parameter values of the neural field or MLP network in the encoded bitstream as described in this paper, the coordinates of the event pixel location where the original event occurred can also be compressed or included in the encoded bitstream.
[0133] Similar evaluations, comparisons, or selections can be made for neural fields or MLPs used to model aggregated events generated by aggregating raw events in event camera data over time intervals of fixed length (e.g., 50 ms, 10 ms, etc.).
[0134] In the example, a quantized neural field modeling events aggregated at 50 ms time intervals, with a data size of 277 KB, provides a prediction accuracy of 99%. The data size of the quantized neural field can be further reduced, for example, by saving or encoding it into a lossless PNG format, resulting in a further compressed data size of approximately 80 KB.
[0135] For example, a quantized neural field modeling events aggregated at 50 ms time intervals with a data size of 277 KB provides a prediction accuracy of 99%. The data size of the quantized neural field can be further reduced, for example, by saving or encoding it into a lossless PNG format, resulting in a further compressed data size of approximately 80 KB.
[0136] The quantization of neural fields demonstrates that relatively small models (e.g., a relatively small total number of neurons and / or a relatively small data size for operating parameter values) can provide representations of modeled raw events or modeled aggregate events without a significant drop in accuracy. Therefore, neural fields can be used as a method for compressing event camera data.
[0137] For illustrative purposes only, it has been described that MLP networks can be used to implement neural fields and predict polarity and / or non-events. However, it should be noted that more than one MLP network can be used to implement the neural fields as described herein. In the examples, the neural field may include (or can be implemented using) two different MLP networks, each of which is specifically trained to predict one type of positive or negative polarity at a given pixel location. In another example, the neural field may include (or can be implemented using) three different MLP networks, two of which are specifically trained to predict one type of positive or negative polarity at a given pixel location, and one of which is specifically trained to predict non-events at a given pixel location. Additionally, alternatively, or alternatively, other artificial neural networks (including, but not limited to, MLP networks) can be used to implement the neural fields as described herein at least in part. For example, neural fields can be combined to implement preprocessing or post-processing neural networks (e.g., CNNs, MLPs, non-MLPs, Transformers, etc.). Post-processing neural networks can be used to receive the output from an MLP network that predicts events and / or non-events at a given pixel location and generate outputs for further processing (including but not limited to predicted events and / or non-events with relatively high accuracy).
[0138] 12. Example Process Flow Figure 4AAn example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., upstream devices, one or more event data codecs, encoding devices / modules, transcoding devices / modules, media devices / modules, etc.) may perform this process flow.
[0139] In box 402, the upstream device receives event camera data containing a raw event sequence. Each raw event in the raw event sequence is generated by a specific sensor element of a plurality of sensor elements of the event image sensor in response to a change in brightness. Each sensor element of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions.
[0140] In box 404, the upstream device generates neural field training data from the raw event sequence in the event camera data. The neural field training data includes multiple training instances. Each training instance in the neural field training data includes the event pixel location coordinates and the ground truth data classification category of the event represented in the training instance.
[0141] In box 406, during the training phase, the upstream device generates optimized values for the operating parameters of the neural field by training the neural field using the neural field training data, at least in part based on a specific loss function chosen for the neural field. The neural field is trained to output corresponding predicted classification categories for pixel locations at multiple time points.
[0142] In box 408, the upstream device encodes neural field operation parameter values in an encoded bitstream. These neural field operation parameter values are generated from optimized values of the neural field's operation parameters. The encoded bitstream enables the receiving device to use the neural field operated on with these neural field operation parameter values to generate multiple reconstructed events that approximate the multiple events represented in the training data.
[0143] In this embodiment, each of the real data classification category and the predicted classification category belongs to a common classification category set; this common classification category set represents a set of different brightness change indicators.
[0144] In an embodiment, the common classification category set represents a set of two classification categories, which consist of the following: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, and a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold.
[0145] In an embodiment, the public classification category set represents a set of three classification categories, which consist of the following: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold, and a third brightness change indicator for brightness changes that do not exceed either the positive or negative brightness change threshold.
[0146] In the embodiments, the events represented in the multiple training instances are one of the following: original events or aggregated events.
[0147] In one embodiment, each of the aggregated events is generated by aggregating the original events at a specific event pixel location within a specific time interval; this specific time interval is set at least in part based on user input.
[0148] In this embodiment, the neural field operation parameter values encoded in the encoded bitstream are generated by quantizing the optimized values of the neural field operation parameters.
[0149] In this embodiment, the neural field is implemented using a multilayer perceptron (MLP) network; the MLP network is trained using positional data from the training data as input and polarity from the training data as real data.
[0150] In this embodiment, the location data is generated by applying location encoding to the coordinates of the pixel locations represented in the training data.
[0151] In an embodiment, the MLP network includes an output layer that uses activation functions of a different type than other activation functions used in other layers of the MLP network.
[0152] In this embodiment, neurons in each of the subsequent layers of the MLP network are fully connected to all neurons in the immediately preceding layer of the MLP network.
[0153] Figure 4B An example process flow according to an embodiment is illustrated. In some embodiments, one or more computing devices or components (e.g., downstream devices, one or more event data codecs, decoding devices / modules, transcoding devices / modules, media devices / modules, etc.) may perform this process flow.
[0154] In box 452, the downstream device decodes neural field operation parameter values from the encoded bitstream generated by the upstream device. These neural field operation parameter values are generated from optimized values of the neural field's operation parameters.
[0155] The optimized values of the neural field's operating parameters are generated by upstream equipment through training the neural field using training data, at least in part based on a specific loss function chosen for the neural field. The neural field is trained to output predicted classification categories.
[0156] The neural field training data comprises multiple training instances. Each training instance in the neural field training data includes the event pixel location coordinates and the ground truth data classification category of the event represented in the training instance. The neural field training data is generated from raw event sequences in event camera data.
[0157] Each raw event in the raw event sequence is generated by a specific sensor element among multiple sensor elements of the event image sensor in response to a change in brightness. Each sensor element of the event image sensor corresponds to a specific pixel position among multiple pixel positions.
[0158] In box 454, the downstream device uses a neural field operated on with neural field operation parameter values to generate multiple reconstructed events that approximate multiple events represented in the training data.
[0159] In this embodiment, the reconstruction event is used to generate visual information related to the physical environment.
[0160] In embodiments, computing devices such as display devices, mobile devices, set-top boxes, and multimedia devices are configured to perform any of the methods described above. In embodiments, an apparatus includes a processor and is configured to perform any of the methods described above. In embodiments, a non-transitory computer-readable storage medium stores software instructions that, when executed by one or more processors, cause any of the methods described above to be performed.
[0161] In one embodiment, a computing device includes one or more processors and one or more storage media storing an instruction set that, when executed by the one or more processors, causes any of the methods described above to be performed.
[0162] Note that although individual embodiments are discussed herein, any combination of the embodiments and / or some of the embodiments discussed herein can be combined to form further embodiments.
[0163] 11. Implementation Mechanism – Hardware Overview Embodiments of the present invention may be implemented using computer systems, systems configured with electronic circuits and components, integrated circuit (IC) devices (such as microcontrollers, field-programmable gate arrays (FPGAs) or other configurable or programmable logic devices (PLDs), discrete-time or digital signal processors (DSPs), application-specific integrated circuits (ASICs)), and / or means including one or more of such systems, devices, or components. The computer and / or IC may execute, control, or implement instructions relating to adaptive perceptual quantization of images with enhanced dynamic range, as described herein. The computer and / or IC may calculate any of the various parameters or values relating to the adaptive perceptual quantization process described herein. Image and video embodiments may be implemented in hardware, software, firmware, and various combinations thereof.
[0164] Some embodiments of the present invention include a computer processor that executes software instructions that cause the processor to perform the methods of the present disclosure. For example, one or more processors, such as those in a display, encoder, set-top box, transcoder, etc., can implement methods related to adaptive perceptual quantization of HDR images as described above by executing software instructions in a program memory accessible to the processor. Embodiments of the present invention may also be provided in the form of a program product. The program product may include any non-transitory medium carrying a set of computer-readable signals, including instructions that, when executed by a data processor, cause the data processor to perform the methods of embodiments of the present invention. Program products according to embodiments of the present invention may take any of a variety of forms. The program product may include, for example, physical media, such as magnetic data storage media including floppy disks and hard disk drives, optical data storage media including CD-ROMs and DVDs, electronic data storage media including ROMs and flash RAMs, etc. The computer-readable signals on the program product may optionally be compressed or encrypted.
[0165] In the case of the components mentioned above (e.g., software modules, processors, components, devices, circuits, etc.), unless otherwise specified, references to these components (including references to “devices”) should be interpreted as including any component that performs the function of the described component as an equivalent of that component (e.g., functionally equivalent), including components that are structurally different from those that perform the functions in the illustrated exemplary embodiments of the invention.
[0166] According to one embodiment, the techniques described herein are implemented by one or more dedicated computing devices. The dedicated computing device may be hardwired to perform these techniques, or may include digital electronic devices persistently programmed to perform these techniques, such as one or more application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or may include one or more general-purpose hardware processors programmed to perform these techniques according to program instructions in firmware, memory, other storage devices, or combinations thereof. Such a dedicated computing device may also combine custom hardwired logic, ASICs, or FPGAs with custom programming to implement these techniques. The dedicated computing device may be a desktop computer system, a portable computer system, a handheld device, a networking device, or any other device incorporating hardwired and / or program logic to implement the techniques.
[0167] For example, Figure 5 This is a block diagram illustrating a computer system 500 on which embodiments of the present invention may be implemented. The computer system 500 includes a bus 502 or other communication mechanism for transmitting information, and a hardware processor 504 coupled to the bus 502 to process information. The hardware processor 504 may be, for example, a general-purpose microprocessor.
[0168] Computer system 500 also includes main memory 506, such as random access memory (RAM) or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When stored in a non-transitory storage medium accessible to processor 504, such instructions enable computer system 500 to become a dedicated machine defined to perform the operations specified in the instructions.
[0169] Computer system 500 further includes read-only memory (ROM) 508 or other static storage device coupled to bus 502 for storing static information and instructions of processor 504. Storage device 510 (such as a magnetic disk or optical disk) is provided and coupled to bus 502 for storing information and instructions.
[0170] Computer system 500 can be coupled to display 512, such as an LCD, via bus 502 for displaying information to a computer user. Input device 514, including alphanumeric keys and other keys, is coupled to bus 502 for transmitting information and command selections to processor 504. Another type of user input device is cursor control 516, such as a mouse, trackball, or cursor arrow keys, for transmitting directional information and command selections to processor 504 and for controlling cursor movement on display 512. Typically, this input device has two degrees of freedom on two axes (a first axis (e.g., x-axis) and a second axis (e.g., y-axis)), allowing the device to specify a position in a plane.
[0171] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic. These custom hardwired logics, one or more ASICs or FPGAs, firmware, and / or program logic, combined with the computer system, enable computer system 500 to be a dedicated machine or programmed to be a special-purpose machine. According to one embodiment, computer system 500 performs the techniques described herein in response to processor 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium (such as storage device 510). Execution of the sequence of instructions contained in main memory 506 causes processor 504 to perform the process steps described herein. In other embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0172] As used herein, the term "storage medium" refers to any non-transitory medium that stores data and / or instructions that enable a machine to operate in a particular manner. Such storage media can include non-volatile media and / or volatile media. Non-volatile media include, for example, optical discs or magnetic disks, such as storage device 510. Volatile media include dynamic memory, such as main memory 506. Common forms of storage media include, for example, floppy disks, floppy hard disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, flash EPROMs, NVRAMs, any other memory chips or memory cartridges.
[0173] Storage media differ from transmission media but can be used in conjunction with them. Transmission media participate in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and optical fibers, including conductors containing bus 502. Transmission media can also take the form of sound waves or light waves, such as those generated during radio wave and infrared data communication.
[0174] Various forms of media can involve loading one or more sequences of one or more instructions to processor 504 for execution. For example, instructions may initially be loaded onto a disk or solid-state drive of a remote computer. The remote computer may load the instructions into its dynamic memory and transmit them over a telephone line using a modem. A modem local to computer system 500 may receive data over the telephone line and convert the data into an infrared signal using an infrared transmitter. An infrared detector may receive the data carried in the infrared signal, and appropriate circuitry may place the data on bus 502. Bus 502 loads the data into main memory 506, from which processor 504 fetches and executes the instructions. Instructions received in main memory 506 may optionally be stored on storage device 510 before or after execution by processor 504.
[0175] Computer system 500 also includes a communication interface 518 coupled to bus 502. Communication interface 518 provides bidirectional data communication coupled to network link 520, which connects to local network 522. For example, communication interface 518 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing data communication connectivity to a corresponding type of telephone line. As another example, communication interface 518 may be a Local Area Network (LAN) card for providing data communication connectivity to a compatible LAN. A wireless link may also be implemented. In any such implementation, communication interface 518 transmits and receives electrical, electromagnetic, or optical signals carrying streams of digital data representing various types of information.
[0176] Network link 520 typically provides data communication to other data devices via one or more networks. For example, network link 520 may provide a connection via local network 522 to host computer 524 or to data devices operated by Internet Service Provider (ISP) 526. ISP 526, in turn, provides data communication services via a global packet data communication network now commonly referred to as the “Internet” 528. Both local network 522 and Internet 528 use electrical, electromagnetic, or optical signals that carry digital data streams. Signals through various networks, as well as signals on network link 520 and through communication interface 518 (which carries digital data to and from computer system 500), are example forms of transmission media.
[0177] Computer system 500 can send messages and receive data, including program code, via a network, network link 520, and communication interface 518. In the Internet example, server 530 can transmit application request code via the Internet 528, ISP 526, local network 522, and communication interface 518.
[0178] The received code may be executed by processor 504 upon receipt and / or stored in storage device 510 or other non-volatile storage for later execution.
[0179] 13. Equivalents, extensions, substitutes, and others In the foregoing description, embodiments of the invention have been described with reference to numerous specific details, which may vary depending on the implementation. Therefore, the sole and exclusive indication of the claimed embodiments of the invention, and of the applicant's opinion, is the set of claims published in specific form according to this application, wherein such claim publication includes any subsequent corrections. Any definitions expressly set forth herein with respect to terms contained in such claims shall govern the meaning of such terms as used in the claims. Therefore, any limitations, elements, properties, characteristics, advantages, or attributes not expressly referenced in the claims should not in any way limit the scope of such claims. Therefore, this specification and drawings should be viewed in an illustrative rather than restrictive sense.
[0180] Exemplary examples of enumeration This invention may be practiced in any of the forms described herein, including but not limited to the enumerated example embodiments (EEE) that describe some parts of the structure, features and functions of embodiments of the invention.
[0181] EEE1. A method comprising: Receive event camera data containing a raw event sequence, wherein each raw event in the raw event sequence is generated by a specific sensor element of a plurality of sensor elements of an event image sensor in response to a brightness change, wherein each of the plurality of sensor elements of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions; Neural field training data is generated from the original event sequence in the event camera data, wherein the neural field training data includes multiple training instances, and each training instance in the multiple training instances of the neural field training data includes the event pixel position coordinates and the real data classification category of the event represented in the training instance; During the training phase, optimized values of the operating parameters of the neural field are generated by training the neural field with training data based at least in part on a specific loss function chosen for the neural field, wherein the neural field is trained to output corresponding predicted classification categories for pixel positions among the plurality of pixel positions at multiple time points. The neural field operation parameter values are encoded in an encoded bitstream, wherein the neural field operation parameter values are generated from optimized values of the operation parameters of the neural field, and wherein the encoded bitstream enables a receiving device to use the neural field operated with the neural field operation parameter values to generate multiple reconstructed events that approximate multiple events represented in the training data.
[0182] EEE2. The method as described in EEE1, wherein the neural field is implemented using a single multilayer perceptron (MLP) network, wherein the output neurons of the MLP network predict one of the following for an input pixel location: (a) a positive or negative event polarity when the input pixel location represents an event pixel location; or (b) a positive or negative event polarity at the input pixel location when the input pixel location represents the event pixel location, and a non-event at the input pixel location when the input pixel location does not represent an event pixel location.
[0183] EEE3. The method as described in EEE1 or EEE2, wherein the neural field is implemented using at least a first multilayer perceptron (MLP) network and a second MLP network separate from the first MLP network; wherein the first MLP network predicts the positive polarity of the input pixel location; and wherein the second MLP network predicts the negative polarity of the input pixel location.
[0184] EEE4. The method as described in EEE3, wherein the neural field is further implemented using a third MLP network separate from the first MLP network and the second MLP network; wherein the third MLP network predicts non-events at the input pixel locations.
[0185] EEE5. The method as described in EEE4, wherein the neural field includes a post-processing MLP network that receives outputs from the first MLP network, the second MLP network, and the third MLP network as inputs and generates a prediction of event polarity for a given pixel location.
[0186] EEE6. The method as described in EEE3, wherein the neural field includes a post-processing MLP network that receives outputs from the first MLP network and the second MLP network as inputs and generates a prediction of event polarity for a given pixel location.
[0187] EEE7. The method as described in EEE1 to EEE6, wherein each of the real data classification category and the predicted classification category belongs to a common classification category set; wherein the common classification category set represents a set of different brightness change indicators.
[0188] EEE8. The method of any one of EEE1 to EEE7, wherein the particular loss function is specifically selected to support a downstream task to be performed using the reconstruction event; wherein the downstream task involves one of the following: an image generation task to be performed by a downstream device, an object recognition task to be performed by the downstream device, or another task to be performed by the downstream device using the reconstruction event.
[0189] EEE9. The method as described in EEE8, wherein the common classification category set represents a set of two classification categories, the two classification categories being composed of: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, and a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold.
[0190] EEE10. The method as described in EEE8, wherein the common classification category set represents a set of three classification categories, the three classification categories being composed of the following: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold, and a third brightness change indicator for brightness changes not exceeding the positive brightness change threshold or the negative brightness change threshold.
[0191] EEE11. The method of any one of EEE1 to EEE10, wherein the events represented in the plurality of training instances are one of the following: original events, or aggregated events, time surfaces, event voxels, or input data items in other event representations.
[0192] EEE12. The method as described in EEE11, wherein each of the aggregated events is generated by aggregating raw events at a specific event pixel location within a specific time interval; wherein the specific time interval is set at least in part based on user input.
[0193] EEE13. The method as described in EEE11, wherein each of the time surfaces is generated by applying an exponential decay kernel to the pixel neighborhood surrounding the specific event pixel location corresponding to the corresponding event in the event.
[0194] EEE14. The method as described in EEE11, wherein the event voxels are generated by assigning each of the events to the nearest voxel in an event voxel grid formed by the point cloud representation of the events.
[0195] EEE15. The method of any one of EEE1 to EEE14, wherein the neural field operation parameter values encoded in the encoded bitstream are generated by quantizing optimized values of the neural field operation parameters.
[0196] EEE16. The method of any one of EEE1 to EEE15, wherein the neural field is implemented using a multilayer perceptron (MLP) network; wherein the MLP network is trained using positional data in the training data as input and polarity in the training data as real data.
[0197] EEE17. The method of any one of EEE1 to EEE16, wherein the neural field is further implemented using one or more of the following: a convolutional neural network, a Transformer neural network, an artificial neural network (ANN) implementing video neural representations, or an ANN other than an MLP network.
[0198] EEE18. The method as described in EEE16, wherein the location data is generated by applying location encoding to the pixel location coordinates represented in the training data.
[0199] EEE19. The method as described in EEE16 to EEE18, wherein the MLP network includes an output layer using an activation function of a different type than other activation functions used in other layers of the MLP network.
[0200] EEE20. The method of any one of EEE16 to EEE19, wherein neurons in each of all subsequent layers of the MLP network are fully connected to all neurons in the immediately preceding layer of the MLP network.
[0201] EEE21. A method comprising: Decode neural field operation parameter values from the encoded bitstream generated by the upstream device, wherein the neural field operation parameter values are generated from optimized values of the neural field operation parameters; The optimized values of the operating parameters of the neural field are generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output predicted classification categories. The neural field training data includes multiple training instances, and each training instance in the neural field training data includes the event pixel location coordinates and the real data classification category of the event represented in the training instance. The neural field training data is generated from the original event sequence in the event camera data. Each original event in the original event sequence is generated by a specific sensor element among a plurality of sensor elements of the event image sensor in response to a change in brightness, wherein each of the plurality of sensor elements of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions; The neural field, operated with the neural field operation parameter values, is used to generate multiple reconstructed events that approximate multiple events represented in the training data.
[0202] EEE22. The method as described in EEE21, wherein the reconstruction event is used to generate visual information related to the physical environment.
[0203] EEE23. The method as described in EEE21 or EEE22, wherein the plurality of reconstructed events constitute a subset of all events in the training data; wherein all events used by the upstream device to train the neural field correspond to a first pixel position and a first time point at a first spatiotemporal resolution; wherein the plurality of reconstructed events correspond to a second pixel position and a second time point at a second spatiotemporal resolution different from the first spatiotemporal resolution; wherein the plurality of reconstructed events are used as input to an image generator to generate an image sequence with one or more of the following: a relatively high frame rate or a relatively high dynamic range.
[0204] EEE24. An apparatus that performs any one of the methods described in EEE1 to EEE23.
[0205] EEE25. A non-transitory computer-readable medium storing software instructions that, when executed by one or more processors, cause to perform the steps of any one of the methods described in EEE1 through EEE23.
Claims
1. A method comprising: Receive event camera data containing a sequence of raw events at multiple time points, wherein each raw event in the sequence of raw events is generated by a specific sensor element of a plurality of sensor elements of an event image sensor in response to a change in brightness, wherein each of the plurality of sensor elements of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions; Neural field training data is generated from the sequence of the original events in the event camera data, wherein the neural field training data includes multiple training instances, and each training instance of the multiple training instances of the neural field training data includes the event pixel position coordinates and the real data classification category of the event at a given time point represented in the training instance. During the training phase, optimized values for the operating parameters of the neural field are generated by training the neural field using training data, at least in part based on a specific loss function chosen for the neural field. The neural field is trained to output a corresponding predicted classification category for pixel positions at a plurality of time points, wherein each of the real data classification category and the predicted classification category belongs to a common classification category set; wherein the common classification category set represents a set of different brightness change indicators. The neural field operation parameter values are encoded in an encoded bitstream, wherein the neural field operation parameter values are generated from optimized values of the operation parameters of the neural field, and wherein the encoded bitstream enables a receiving device to use the neural field operated with the neural field operation parameter values to generate multiple reconstructed events that approximate multiple events represented in the training data.
2. The method of claim 1, wherein, The neural field is implemented using a single multilayer perceptron (MLP) network, wherein the output neurons of the MLP network predict one of the following for an input pixel location: (a) the positive or negative event polarity when the input pixel location represents an event pixel location; or (b) the positive or negative event polarity at the input pixel location when the input pixel location represents the event pixel location, and the non-event at the input pixel location when the input pixel location does not represent an event pixel location.
3. The method as described in claim 1, wherein, The neural field is implemented using at least a first multilayer perceptron (MLP) network and a second MLP network separate from the first MLP network; wherein the first MLP network predicts the positive polarity of the input pixel location; and wherein the second MLP network predicts the negative polarity of the input pixel location.
4. The method of claim 3, wherein, The neural field is further implemented using a third MLP network separate from the first and second MLP networks; wherein the third MLP network predicts non-events at the input pixel locations.
5. The method of claim 4, wherein, The neural field includes a post-processing MLP network that receives outputs from the first MLP network, the second MLP network, and the third MLP network as inputs and generates a prediction of event polarity for a given pixel location.
6. The method of claim 3, wherein, The neural field includes a post-processing MLP network that receives outputs from the first MLP network and the second MLP network as inputs and generates a prediction of event polarity for a given pixel location.
7. The method according to any one of claims 1 to 6, wherein, The specific loss function is specifically selected to support a downstream task to be performed using the reconstruction event; wherein the downstream task involves one of the following: an image generation task to be performed by the downstream device, an object recognition task to be performed by the downstream device, or another task to be performed by the downstream device using the reconstruction event.
8. The method according to any one of claims 1 to 7, wherein, The public category set represents a set of two categories, which consist of the following: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, and a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold.
9. The method according to any one of claims 1 to 7, wherein, The public category set represents a set of three categories, which consist of the following: a first brightness change indicator for positive brightness changes exceeding a positive brightness change threshold, a second brightness change indicator for negative brightness changes exceeding a negative brightness change threshold, and a third brightness change indicator for brightness changes that do not exceed either the positive or negative brightness change thresholds.
10. The method according to any one of claims 1 to 9, wherein, The events represented in the plurality of training instances are one of the following: original events, aggregated events, time surfaces, event voxels, or input data items in other event representations.
11. The method of claim 10, wherein, Each of the aggregated events is generated by aggregating the original events at a specific event pixel location within a specific time interval; wherein the specific time interval is set at least in part based on user input.
12. The method of claim 10, wherein, Each of the time surfaces is generated by applying an exponential decay kernel to the pixel neighborhood around the specific event pixel location corresponding to the event in the event.
13. The method of claim 10, wherein, The event voxels are generated by assigning each of the events to the nearest voxel in the event voxel grid formed by the point cloud representation of the events.
14. The method according to any one of claims 1 to 13, wherein, The neural field operation parameter values encoded in the encoded bitstream are generated by quantizing the optimized values of the neural field operation parameters.
15. The method according to any one of claims 1 to 14, wherein, The neural field is implemented using a multilayer perceptron (MLP) network; wherein the MLP network is trained using positional data from the training data as input and polarity from the training data as real data.
16. The method according to any one of claims 1 to 15, wherein, The neural field is further implemented using one or more of the following: convolutional neural network, Transformer neural network, artificial neural network (ANN) implementing video neural representation, or ANN other than MLP network.
17. The method of claim 15, wherein, The location data is generated by applying location encoding to the pixel location coordinates represented in the training data.
18. The method according to any one of claims 15 to 17, wherein, The MLP network includes an output layer that uses an activation function of a different type than other activation functions used in other layers of the MLP network.
19. The method according to any one of claims 15 to 18, wherein, Neurons in each of the subsequent layers of the MLP network are fully connected to all neurons in the immediately preceding layer of the MLP network.
20. A method comprising: Decode neural field operation parameter values from the encoded bitstream generated by the upstream device, wherein the neural field operation parameter values are generated from optimized values of the neural field operation parameters; The optimized values of the operating parameters of the neural field are generated by the upstream device by training the neural field with neural field training data based at least in part on a specific loss function selected for the neural field, wherein the neural field is trained to output predicted classification categories. The neural field training data includes multiple training instances, and each training instance in the neural field training data includes the event pixel location coordinates and the real data classification category of the event at a given time point represented in the training instance. The neural field training data is generated from the sequence of raw events at multiple time points in the event camera data. Each original event in the sequence of original events is generated by a specific sensor element among a plurality of sensor elements of the event image sensor in response to a brightness change, wherein each of the plurality of sensor elements of the event image sensor corresponds to a corresponding pixel position among a plurality of pixel positions; wherein each of the real data classification category and the predicted classification category belongs to a common classification category set; wherein the common classification category set represents a set of different brightness change indicators; The neural field, operated with the neural field operation parameter values, is used to generate multiple reconstructed events that approximate multiple events represented in the training data.
21. The method of claim 20, wherein, The reconstruction event is used to generate visual information related to the physical environment.
22. The method of claim 20 or 21, wherein, The plurality of reconstructed events that are close to each other constitute a subset of all events in the training data; wherein all events used by the upstream device to train the neural field correspond to a first pixel position and a first time point at a first spatiotemporal resolution; wherein the plurality of reconstructed events correspond to a second pixel position and a second time point at a second spatiotemporal resolution different from the first spatiotemporal resolution; wherein the plurality of reconstructed events are used as input to an image generator to generate an image sequence with one or more of the following: a relatively high frame rate or a relatively high dynamic range.
23. An apparatus that performs any one of the methods claimed in claims 1 to 22.
24. A non-transitory computer-readable medium storing software instructions that, when executed by one or more processors, cause the steps of any one of the methods described in claims 1 to 22 to be performed.