Event-based image processing

The two-stage processing system for EBS enhances performance in low-light conditions by applying compressive nonlinearity and filters, improving dynamic range and noise suppression for flexible event detection and edge/motion identification.

JP2026514353APending Publication Date: 2026-05-11CUVOS PTY LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CUVOS PTY LTD
Filing Date
2024-03-19
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing event-based image sensors (EBS) face limitations in low-light and complex lighting conditions, requiring high-performance computing for noise handling and compromised performance reliability, and lack flexibility in nonlinearity and parameter settings.

Method used

A two-stage processing system with compressive nonlinearity and feedback loops for event detection, incorporating high-pass and band-pass filters to enhance dynamic range and suppress noise, allowing configurable event detection and edge/motion identification.

Benefits of technology

Enhances EBS performance in challenging lighting conditions, enabling high-speed, flexible event detection and edge/motion identification using commercially available sensors, reducing noise impact and improving dynamic range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514353000001_ABST
    Figure 2026514353000001_ABST
Patent Text Reader

Abstract

An event-based image processing method is disclosed. The method includes: acquiring time-series input signal data characterizing information from an input source; applying a first processing stage, which comprises applying a compression nonlinearity to the input signal data; and applying a second processing stage to the output from the first processing stage. The second processing stage comprises time and / or spatial processing within a feedback loop, applying a high-pass or band-pass time filter and / or a high-pass or band-pass spatial filter to the output from the first processing stage, thereby suppressing recurring changes and improving the ratio of events to recurring changes in the output from the first processing stage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method and apparatus for event-based image processing.

Background Art

[0002] Event-based image sensors (EBS), also known as dynamic vision sensors (DVS), have become popular in recent years. Their operation is based on the biology of animal eyes and is developed following a neuromorphic approach. EBS is preferred in situations where lighting conditions are harsh, data throughput is important, and power is constrained.

[0003] The earliest event-based sensor (EBS) was described by T. Delbruck in the paper "A 128x128 120 dB 15μs Latency Asynchronous Temporal Contrast Vision Sensor", IEEE Journal of Solid-State Circuits, vol. 43, no. 2, pp. 566-576, February 2008, doi: 10.1109 / JSSC.2007.914337. The study included a simplified circuit for EBS reproduced here in FIG. 1 and an operating principle reproduced here in FIG. 2.

[0004] The circuit shown in FIG. 1 stores the change in voltage at each time step, and the stored voltage change corresponds to the logarithm of the change in the photodiode current across capacitor C1. The value of this change in photodiode current is amplified by C1 / C2 to generate a voltage change value V diff to generate. V diff is used to determine whether the logarithmic change in the photodiode current is large enough to generate an "event". An "ON" event is when V diffAn "OFF" event is generated when the value increases positively by a minimum amount, and an "OFF" event is generated when there is a sufficiently large negative change. The "threshold," which is the required magnitude of change, is set globally for both positive and negative changes.

[0005] When the logarithm of the photodiode current is taken, small changes in current are amplified more than large changes in current. This compressive nonlinearity is similar to that observed in biological sensory organs. However, the difference from EBS is that 1) EBS only reports whether the change is positive, negative, or no change occurred, rather than the magnitude of the change, and 2) logarithmic compressive nonlinearity is established. Figure 3 illustrates this nonlinearity. Therefore, there is no way to explore different types of nonlinearity. Furthermore, the parameters of the logarithmic response are fixed.

[0006] In reality, the widespread adoption of EBS technology has been slow due to several drawbacks in its real-world performance and operational limitations.

[0007] Currently, custom EBS solutions are available on the market, such as those manufactured by Prophesee (https: / / www.prophesee.ai / ) and IniVation (https: / / iniVation.com / ). However, currently available EBS solutions still require optimal lighting conditions to maintain the stated performance specifications. Performance reliability is compromised in low-light conditions with few photons, or in complex lighting conditions where the amount of photons may not be stable. Furthermore, processing is required to interpret the event data provided by EBS. The event data is asynchronous, meaning that when pixels generate an event, they are timestamped and sent to the event bus. In low-light or suboptimal lighting conditions where noise can easily overwhelm the "event" data, processing must be performed by very fast hardware, or the event detection threshold must be set very high, in order to enable event generation. Setting the threshold high reduces the high dynamic range benefit of EBS. On the other hand, using high-speed post-processing hardware requires high-performance computing solutions such as GPUs or FPGAs.

[0008] There are several research papers on EBS embodiments using software or digital hardware such as FPGAs. Most of these papers focus on emulating the analog circuit performance of EBS, rather than the biological function of M-type retinal ganglion cells. Essentially, these models attempt to emulate something that is already an abstraction of the biological function. Software simulations or digital emulators are used as substitutes for actual EBS solutions, not as improvements to existing EBS systems.

[0009] Where prior art is referenced in this specification, it should be understood that such references do not constitute an acknowledgment that such prior art forms part of the common general knowledge in the art in Australia or any other country. [Overview of the project] [Problems that the invention aims to solve]

[0010] This specification discloses embodiments of alternative solutions to currently available EBS systems. [Means for solving the problem]

[0011] In a first embodiment, an image processing method is disclosed. The method comprises: acquiring time-series input signal data characterizing information from an input source; applying a first processing step, which includes applying compression nonlinearity (or alternative nonlinear signal processing that increases the information content of the input frame, e.g., delentropy) to the input signal data; and applying a second processing step to the output from the first processing step. The second processing step comprises spatial and / or temporal processing within a feedback loop, which applies a high-pass or band-pass spatial and / or temporal filter to the output from the first processing step, thereby suppressing recurring (and slow) changes and improving the ratio of events to recurring changes in the output from the first processing step. In some embodiments, the method comprises detecting events in the output of the second processing step. An event may be defined as a change in pixel intensity over time that satisfies a specific slew rate threshold.

[0012] In some forms, detecting an event involves applying at least one threshold to the pixel values ​​of the output of a second processing stage.

[0013] In some forms, the threshold applied includes multiple thresholds, each applied to a specific pixel or set of pixels.

[0014] In some forms, detecting an event involves setting an event rate for the detected event.

[0015] In some forms, event detection involves setting two different event rates for each of the two different detected events.

[0016] In some forms, the event rate is set by setting a counter that is the number of clock cycles during which a detected event is expected to persist.

[0017] In some forms, a time-high-pass or band-pass filter is an n-th order filter, where n corresponds to the number of time steps incorporated into the time processing.

[0018] In some forms, n is equal to 1.

[0019] In some forms, n is greater than 1.

[0020] In some forms, time-high-pass or band-pass filters are infinite impulse-response filters.

[0021] In some forms, the gain of a time-high-pass or band-pass filter is variable based on at least one or more of the characteristics of the input source and the characteristics of the environment in which the input source is acquiring the input data.

[0022] In some forms, the second processing stage further comprises spatial processing configured to detect edges. Examples of edge detection algorithms include, for example, Sobel filtering and Canny edge detection.

[0023] In some configurations, spatial processing is configured to occur before or after temporal processing.

[0024] In one form, the spatial processing applies a spatial high-pass filter implemented using convolution.

[0025] In one form, the first processing stage includes a time feedback loop, in which current and previous samples from the input signal data and at least one previous sample of the output from the first processing stage are used to obtain the current sample of the output from the first processing stage.

[0026] In one form, applying the compressive non-linearity to the input data includes applying a gain to the input signal data, the gain being variable based on the magnitude of the input signal data.

[0027] In one form, the first processing stage includes a plurality of processing modules configured to advantageously process input data having different characteristics.

[0028] In one form, the processing module includes a first processing module for advantageously processing input data of a lower magnitude and a second processing module for advantageously processing input data of a higher magnitude.

[0029] In one form, each processing module of the first processing stage includes a divisive low-pass filter.

[0030] In one form, the input source includes an image sensor, the time-series input data is a series of frames, each frame includes a plurality of pixels, and the first and second processing stages are applied pixel by pixel.

[0031] In one form, the image sensor is an electro-optical sensor, a camera, or an infrared image sensor, or another sensor including one or more pixels.

[0032] In a second aspect, a signal processing device comprising a plurality of processing modules is disclosed herein. These modules include a first processing module configured to receive and process a time-series input signal from an input source, the first processing module being configured to apply a compression nonlinearity to the received input signal. These modules also include a second processing module configured to receive the output from the first processing module, the second processing module being configured to perform spatial and / or time processing in a feedback loop and to apply a high-pass or band-pass spatial and / or time filter to the output from the first processing stage, thereby suppressing recurring changes and improving the ratio of events to recurring changes in the output from the first processing stage.

[0033] In some configurations, the device includes an event detection module configured to perform event detection on the output from a second processing module.

[0034] In some configurations, the second processing module is further configured to apply a spatial filtering process to detect edges in the output of the first processing module.

[0035] In some configurations, the first processing module, the second processing module, or both are implemented in hardware.

[0036] In some configurations, the input signal contains signal data for a set of pixels, and the processing module processes the input signal pixel by pixel.

[0037] Next, embodiments will be described as mere examples with reference to the attached drawings. [Brief explanation of the drawing]

[0038] [Figure 1] This is a schematic diagram of a prior art circuit for implementing an event-based sensor. [Figure 2]Figure 1 is a schematic diagram illustrating the operating principle of the event-based sensor circuit shown. [Figure 3] This is a schematic diagram illustrating the nonlinearity of logarithmic compression. [Figure 4] This is a schematic diagram of an event-based detection system according to one embodiment of the present invention. [Figure 5] This is a schematic diagram of how the sliding kernel works. [Figure 6-1] This figure shows a set of sliding kernels configured to detect vertical and horizontal edges. [Figure 6-2] This figure shows another set of sliding kernels configured to detect vertical and horizontal edges within a set of pixels. [Figure 7(a)] This diagram schematically illustrates the processing in the first processing stage of an event-based detection system according to one embodiment of the present invention. [Figure 7(b)] This diagram schematically illustrates the processing in the first processing stage of an event-based detection system according to one embodiment of the present invention. [Figure 8-1] This is a frame-based output "before the event" showing a moving hand. [Figure 8-2] This is a frame-based output from "before the event," showing the camera in motion. [Figure 8-3] This is the frame-based output "before the event" when there is no movement. [Figure 9] This is a conceptual diagram of pixel values ​​provided to the pipeline structure and used to generate filtering values ​​for local pixels. [Figure 10] This is an example of an image rendered from event output obtained by processing image data captured during drone flight. [Modes for carrying out the invention]

[0039] The following detailed description refers to the accompanying drawings, which form part of the detailed description. The exemplary embodiments described in the detailed description and shown in the drawings are not intended to limit. Other embodiments may be used and other modifications may be made without departing from the spirit or scope of the subject matter presented. It will be readily apparent that the aspects of the disclosure generally described herein and shown in the drawings can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are contemplated in the disclosure.

[0040] This specification discloses embodiments of event-based sensing systems. Embodiments of the system, including algorithms, hardware implementations, and software extensions, are discussed. As described herein, the event-based sensing systems are configured to apply novel processing to detect “motion events,” or “events” that indicate motion. Advantageously, the system can be configured to detect “events” based on pixel-level image data rather than requiring image data from a set of pixels. An event may also be defined as a change in pixel intensity over time that satisfies a specific slew rate threshold.

[0041] The advantage of this system is that it can operate with input data from commercially available (COTS) image sensor technology and, by applying digital signal processing (DSP), enhance, detect edges and motion, and then generate “events.” Embodiments of the described system can incorporate EBS-compatible Addressed Event Representation (AER) output, which, along with signal processing, is implemented on a field-programmable gate array (FPGA) or application-specific integrated circuit (ASIC). AER is an output protocol commonly used in EBS. The AER protocol addresses the pixel where an event occurred by its x, y coordinates and the time of the event. While directly related to the hardware, it is essentially a file protocol that can be used by EBS algorithms to process the data. The disclosed solution means that any image sensor can be transformed into a high-performance EBS camera. By including image enhancement that is low latency and delivers performance equivalent to or exceeding custom EBS in complex and challenging lighting conditions, the described system provides an easily integrated and usable alternative to currently available EBS solutions.

[0042] Embodiments of the event-based sensing system 100 disclosed herein include a first processing stage 102 and a second processing stage 104, as conceptually shown in Figure 4. Input to the first stage is provided by a sensor 106, which may be external to the system 100 or built into the system, as shown. The first stage 102 of the system 100 implements parametric compression nonlinearity, increasing the gain of low input signals and saturating the gain of high input signals, thereby extending the dynamic range of each pixel output. The second stage 104 of the system is configured to detect motion and edges in the image.

[0043] In some embodiments, the pixel output from the first stage 102 may also be used to provide an enhanced image 112 with a better dynamic range and is compared to the original input from the sensor 106.

[0044] The first stage of the system requires the use of memory to track pixel changes, thereby providing feedback for pixel enhancement. The purpose of this stage is to extend the dynamic range beyond what is available from the image sensor, as will be discussed later. Local pixel changes, global changes, or both may be tracked. A “local” pixel refers to the individual pixel being processed. A “global” pixel refers to pixels across the entire frame or pixels in a specific region. An example of a function used in the first stage 102 is described later with reference to Figure 7.

[0045] As shown in Figure 4, the first stage 102 may be implemented using an N-pipeline approach. In the N-pipeline approach applied to time-series image data, pixel data from different parts of the frame is processed by each pipeline. Pixels from the image are provided pixel by pixel to each pipeline in the first stage 102, enabling tracking of changes occurring at the local pixel level. The pipelines may be further configured so that different subsets of pixels provide input to each pipeline, thereby enabling tracking of changes "globally". The N-pipeline approach is discussed in more detail in the applicant's Australian Provisional Patent Application No. 2022902462, which is incorporated herein by reference. Processing in each pipeline involves time processing in that pixel data from frames acquired at different sample times (e.g., the current frame and the previous frame) is used as input for processing. Thus, processing accesses memory to access data from previous time steps. This improves the system's dynamic range and allows processing of image frames acquired under a wider range of lighting conditions.

[0046] A second stage 104 of the event-based detection system applies processing to the output from the first stage 102 to find edges and motion in the image. Edge and motion detection may be performed via temporal processing, spatial processing, or both. In one embodiment, the second stage 102 does this by applying a high-pass or band-pass temporal filter 108 and a spatial high-pass filter 110. The spatial high-pass filter 110 may be applied after the high-pass temporal filter 108. Both filters 108 and 110 look for edges in the data (in pixels and in the resulting image). Therefore, the order in which the filters are applied may be reversed without affecting the result. Also, since both filters 108 and 110 function to look for edges, either one or both (i.e., spatial and / or temporal filters) may be used, but it is preferable that at least the temporal filter 108 is included.

[0047] In some embodiments, the time filter 108 is applied on a pixel-by-pixel basis. In some embodiments, the time filter 108 is an infinite impulse response (IIR) one-tap (i.e., first-order) filter. Equation (1) provides an example definition of a first-order high-pass filter. In equation (1), n ​​is each discrete-time step, α is the gain term, x is the pixel value output from the first step, and y is the output from the high-pass time filter.

number

[0048] The time filter 108 applied in the second stage 104 may be generalized to include another “tap,” that is, it may be a first-order filter or have a higher order. The generalized formula, equation (2), is given below.

number

[0049] In equation (2), n is each discrete-time step, α is the gain term, x is the pixel value output from step 1, y is the output at step 2, and βi is a constant selected for filter stability.

[0050] The time filter 108 is parameterizable and can be adjusted to suit applications where event-based detection is performed. Parameter α may be optimized to suit the image sensor used, or to suit visual conditions or a specific object of interest. For example, parameter α may be adjusted by adjusting the cutoff frequency based on the velocity of the edge configured to be captured by the filter. In addition or alternatively, the filter 108 may be more tuned to a particular frequency by increasing the order of the filter, i.e., increasing the steepness of the frequency cutoff. More generally, the adjustments that can be applied to the filtering in this stage 104 are primarily related to the dynamic performance of the system, i.e., how well the system performs in response to the dynamics of the object being captured by the system.

[0051] The use of memory containing multiple time steps from the time filter output helps time processing to account for recurring changes or patterns of change. This can be further enhanced by combining both the use of time steps and configurable parameters for feedback from those time steps. In this way, the filter can be configured to adapt to specific characteristics of the environment monitored by the event-based detection system. This can help suppress known or expected motion patterns and better identify information that is more likely to correspond to a true event, as opposed to background patterns or noise, such as the detection of a boat or vessel in a waterway under reduced influence from expected background or irrelevant changes, such as motion caused by waves.

[0052] The number of taps, i.e., the order of the filter, may be adjusted to control time accuracy. The number of taps, i.e., the order of the filter, is chosen to improve time accuracy (i.e., greater frequency selectivity) and assist in sub-pixel targeting. Higher-order IIRs can also help reduce noise. Increasing the number of "taps" in the filter in step 2 increases the order of the filter, which makes the filter more selective and therefore more sensitive to changes in the filter's frequency response bandwidth. However, increasing the number of taps in a digital filter requires more memory. Therefore, there is a trade-off between accuracy and resource usage.

[0053] Here, "subpixel" means that the image captured in the image frame can detect moving objects smaller than a pixel. Such objects appear smaller than a single pixel in the image because they are either small enough, far enough away, or both. Because the dynamic range of the image is extended by the nonlinearity in the first stage 102, during the second stage 104, it becomes easier to distinguish the intensity changes caused by such "subpixel" moving objects from noise. The "event" is detected at the pixel where the object is located.

[0054] When such "subpixel" objects transition across the pixel boundary between two pixels, the changes in the pixel intensity of those two pixels are inversely proportional. This is because the motion between pixels, and therefore the inverse relationship between the intensity changes in the two pixels, continues for a longer period and is captured over more frames. In other words, objects that are farther away and moving across the scene, occupying fewer pixels, move across the image more slowly than objects that occupy more pixels in the image. By increasing the order of the time filter (i.e., the number of taps), the frequency cutoff is reduced, allowing the algorithm to extract objects that move increasingly slowly. Alternatively, if only fast objects are considered, the frequency cutoff can be higher.

[0055] Theoretically, there is no limit to the number of taps in a selected filter, i.e., the filter order. However, in practice, the number of taps may be limited by practical constraints such as the associated increase in memory requirements. Furthermore, increasing the number of "tap" does not always lead to a significant improvement in filter performance. Therefore, the optimal number of taps may change in the future based on advances in computing memory and the associated decrease in costs, as well as the characteristics of the captured image.

[0056] The difference between the processing applied by the currently disclosed system and conventional methods is the use of time filtering at the pixel level to determine edge motion, i.e., edge movement. This filtering may be high-pass or band-pass. A common technique for approximating EBS is to use frame subtraction, for example, by equation (3), which does not utilize the output y from the previous sample point. This can be interpreted as a first-order, one-tap high-pass finite impulse response (FIR) filter and does not involve time processing utilizing the filter output, as shown below. Formula (3) y[n]=x[n]-x[n-1]

[0057] The processes employed in the systems described herein differ. The time filters used in these systems remove noise by correlating motion between pixels. For example, if pixel A changes, it is possible to determine whether the change was due to noise by examining whether what was in pixel A has moved to one of the adjacent pixels. If there is no change in the pixels in the vicinity of pixel A, the change in pixel A is likely to be noise and should be filtered out. In the disclosed embodiments, the time filter in stage 2 incorporates memory, enabling feedback and significantly reducing the impact of noise. Furthermore, the use of time filtering (as opposed to simply subtracting frames from different sample times) is advantageous in that the filter can be parameterized by frequency cutoff and number of taps, as previously stated. The output from the time filter 108 in stage 2 provides a frame-based pixel output that shows motion.

[0058] The frame-based pixel output from stage 2 can be considered a “pre-event” output, which is processed to generate event outputs. The event generation process is represented by block 114. Events are generated by comparing changes in pixel values ​​to the previous time step and thresholding the changes by a predetermined threshold to determine whether the output is “ON” or “OFF”. ON and OFF events can be color-coded in an event visualizer to show the movement of objects and their direction. ON and OFF events also help detection and tracking algorithms determine the direction and speed of motion. Optionally, each pixel may have its own programmable filter, rather than applying a global threshold as in the case of currently available custom EBS solutions. Furthermore, the speed of events can be programmed with a counter which may be specified for each pixel. This means that for slowly moving objects, the event may persist, or for fast-moving objects, the event may move on in the next clock cycle by resetting the counter to 0. This level of temporal programmability is not provided by existing EBS systems.

[0059] Referring again to Figure 4, the spatial high-pass filter 110 applied in the second stage 104 is intended to detect edges in the image data using spatial processing. By combining the temporal and spatial filters, the second stage 104 of system 100 detects motion only in the “edges” in the image.

[0060] In image processing, an edge is the boundary between a bright pixel and a dark pixel, or between pixel regions. In its simplest form, an edge detection algorithm compares all adjacent pixels in an image to determine whether a boundary exists between two pixels. More complex algorithms can use the derivative of pixel intensity to determine the intensity and direction of an edge. By detecting only edge motion, some of the artifacts that may appear on objects with uneven lighting, as in most natural scenes, can be reduced. Edge detection algorithms can be selected by those skilled in the art. Non-limiting examples include Sobel filtering and Canney edge detection.

[0061] In some embodiments, the spatial high-pass filter 110 is implemented using convolution operations. One example is the “sliding frame” approach, in which the input image is convolved with a kernel that “slides” across the input image. In the “sliding frame” approach, edges of a point of interest can be found by looking at the pixels around the point of interest and searching for the difference in data on both sides of the point of interest. For example, if the left side is dark and the right side is bright, or vice versa, a perpendicular edge exists at this point. This can be repeated for the entire image to create an image containing only edges. Mathematically, this is described by convolution, which provides a convenient way to generate these images. The convolution kernel used to “slide” across the image is configured to detect the required pixel intensity transitions so that edges can be identified. Figure 5 conceptually illustrates this approach. A 3x3 sliding kernel 202 is applied to a 6x6 input image 201, resulting in a 4x4 resulting image 203. The value of each pixel in the resulting image 203 is the result of the convolution between the kernel and a subset of pixels of the same size.

[0062] Some examples of kernels designed to detect edges are shown in Figure 6. In Figure 6-1, kernels 301 and 302 are configured to detect vertical and horizontal edges, respectively, and their results can be combined to provide a gradient image. In Figure 6-2, kernels 303 and 304 are also configured to detect vertical and horizontal edges, but have a weighting parameter "x" that can be adjusted to a value greater than 1 to emphasize the edges. The operators shown in Figure 6 are merely examples. It will be understood that other operators may be selected as appropriate by those skilled in the art depending on the specific application. For example, if the object of interest is known to have rotational symmetry, the operator or mask may be selected to emphasize the edges around the center.

[0063] The spatial filter 110 can be implemented as a software algorithm in some embodiments, but can also be implemented in hardware in other embodiments. In the hardware embodiment, the sliding frame approach may be simplified to a Fourier transform, which can greatly improve the hardware implementation. For example, this can utilize a hardware architecture for the Fast Fourier Transform. Given that the transition from dark to bright areas involves a sudden surge in intensity, which in the frequency domain is high-frequency information, the Fourier transform embodiment seeks high-frequency information in the image. Therefore, this embodiment attempts to remove low-frequency information via high-pass or band-pass filtering. The input signal is provided to the Fourier transform hardware pixel by pixel, which has the effect of applying a convolution (e.g., the kernel shown in Figure 6) to each pixel.

[0064] It should be noted that the size of the spatial filter is fully programmable. This is particularly advantageous for EBS originating from frames with a finite frame rate. Objects moving at very high speeds may traverse two or more pixels during frame transitions, and therefore, large spatial filters may be employed to track edges and ensure that pixel transitions correlate with changes in other pixels (whether adjacent or not).

[0065] The two-stage approach described above offers the advantage that event-based sensing systems are parameterizable. Conventional EBS have limitations in their dynamic response because they cannot change the physical properties of the diode (i.e., the image sensor). In terms of dynamic range, the only thing conventional EBS can do to add compressive nonlinearity is to use logarithmic operations on the magnitude of the diode current. Similarly, conventional EBS do not utilize memory, but rather use simple subtraction (e.g., Equation 3) to detect changes.

[0066] On the other hand, the event-based sensing system of this disclosure allows for flexibility in the selection of which type of compression nonlinearity and parameters to implement, and can be configured by adjusting the frequency response and the number of taps as described above. This is because the processing applied in the first stage 102 is configurable, rather than being limited to taking the logarithm of the photodiode current. The system further utilizes memory, which works to reduce the effects of noise as described above, and thus improves the ability to optimize the system for targets of various sizes and speeds.

[0067] Furthermore, unlike conventional EBS systems that generate images by combining events, this system enables near-simultaneous frame-based pixel output and separate event output. The displayed frame-based pixel output can be an edge image or its enhancement. Preferably, the frame-based pixel output and event output are synchronized to the frame rate clock cycle. Depending on the image sensor used, the user can view the complete image of what is being recorded with a latency of less than 1 microsecond (μs). The system described here can leverage pipeline-based processing to provide high-speed processing, as will be discussed later with reference to Figure 9.

[0068] It should be noted that the system described here is sensor technology independent. That is, it can be used with any sensor input, including electro-optical, near-infrared (NIR), short-wavelength infrared, and ultraviolet. This allows the system to provide the benefits offered by event-based processing using existing sensor technologies, without requiring specialized EBS (Event-Based Sensing) techniques. For example, event-based detection can be performed using an infrared sensor without developing an event-based infrared sensor.

[0069] Figures 7(a) and 7(b) schematically illustrate an embodiment of the processing performed in the first stage 102 of the event-based sensing system 100. In a multi-pipeline approach, this processing may be performed in one or more of the pipelines. This embodiment applies a different type of compressed nonlinearity to the input signal 106 than that shown in Figure 3, in that the nonlinear response changes depending on the lighting conditions.

[0070] First, the input signal 106 is processed by passing it through a low-pass filter 401 to remove low-frequency noise components. In this embodiment, the filtered signal is processed by a plurality of processing modules. The plurality of processing modules include a processing module 402 that prioritizes bright light or high-contrast signals and a processing module 403 that prioritizes low-light or low-contrast signals. Processing modules 402 and 403 each include a division-type low-pass filtering process to reduce the influence of quantum noise in the signal arising from stochastic noise in the image sensor. In the division-type low-pass filtering, the low-pass filter output is fed into a feedback loop and divided by the input signal to the module (provided by the output from 401), and the division result is fed as the input to the low-pass filter. In processing module 402, the filter output (from filter 405) is further passed through a higher-frequency-prioritizing nonlinear gain 404 before being divided by the processing module input (provided by filter 401). The bandpass behavior of this processing module is brought about by the nonlinear gain 404 following the low-pass filter 405. In the example discussed below (e.g., Equation 5), the nonlinearity is accompanied by an exponential gain.

[0071] The filter bandwidth is selected to match the frequency range preferred by each module 402, 403. The outputs from processing modules 402, 403 are fed to a nonlinear compression stage 409, and as a result, the overall output signal of the first stage 102 is compressed to a specific range. That is, the nonlinear compression stage 409 operates on the data that has passed through processing modules 402, 403. The specific processing architecture of the data processed by modules 402, 403 before being processed by the nonlinear compression stage 409 can be selected by those skilled in the art depending on the application.

[0072] For example, as a person skilled in the art will understand, one possibility is parallel processing, in which modules 402, 403 process the output from the low-pass filter 401 in parallel, and then enter an arithmetic unit 411 such as an adder or averager, which can then be fed from that arithmetic unit to the nonlinear compression stage 409. This is shown in Figure 7(a). Naturally, a person skilled in the art will also understand that another way of combining the processing is to process the data in series through these modules 402, 403, and then further in series to the nonlinear compression stage 409, as shown in Figure 7(b).

[0073] In this embodiment, the nonlinear compression step 409 is implemented using Naka-Rushton / Gamma compression or an approximation thereof. In other embodiments, other functions that asymptotically approach the limit value may be used as alternatives and can be selected by those skilled in the art.

[0074] Examples of applicable filters are provided below. However, it will be understood that these may be appropriately modified by those skilled in the art to suit the scenario and application.

[0075] An example of an initial low-pass filtering path in filter 401 (see Figure 7) is given by equation (4) below, which defines the transfer function H1 in different complex frequency domains (the z domain refers to the complex frequency domain). In equation (4), α sets the time constant. When α=0, the output does not change, i.e., the output is independent of the input; when α=1, the output is equal to the input, and the filter becomes a finite impulse response (FIR) filter. When α=-1, the filter is unstable. In general, α is selected in the range 0 < α < 1.

number

[0076] While low-pass filtering here is expressed as a first-order low-pass filter, it should be noted that higher-order filters may also be used. Using higher-order filters is expected to improve image contrast in low-light conditions, but it requires more memory. With currently available computing systems, memory requirements tend to decrease the efficiency of third-order and higher-order algorithms.

[0077] An example of the processing performed by the high-intensity light processing path 402 is represented by the following equation (5), which defines the division-type low-pass transfer function H2. Here, β has the same function as α described above and sets the time constant of the filter. The variable a determines the rate of change of the exponential function. T scales the rate of change in the discrete-time domain (z domain).

number

[0078] The low-light processing pass applies a low-pass filter 406. An example of the processing in pass 403 is expressed as a transfer function defined by equation (6) below, where γ has the same function as α and β described above.

number

[0079] The outputs from processing paths 402 and 403 are taken out at the outputs of dividers 407 and 408 and then fed to the nonlinear compression stage 409. An example of nonlinear compression 409 is implemented using the Naka-Rushton function, which can be expressed by the following equation (7): where R is the response to contrast C (which is the signal output of H2 and H3, where C=H2 in the high-contrast path and C=H3 in the low-contrast path), K is the asymptotically maximum response amplitude, and n is proportional to the slope of the curve at the point where the contrast is K. b is the bias point around which compression takes place. b is set by the DC operating bias of the sensor.

number

[0080] As described above, nonlinear compression may be performed by other functions in different embodiments. For example, the Naka-Rushton function may be replaced by the inverse tan function. See equation (8) below. Equation (8) is not a direct transformation of equation (7) in that the parameters K, n, and b in equation (8) are not the same as those in equation (7). However, the similar symbols in both equations are corresponding constants that have a similar effect on the response R. That is, K sets the maximum contrast, the slope is governed by n, and b is the offset. Formula (8) R(C) = R max tan -1 (nC-K)+b

[0081] Figure 8 shows an example of the “pre-event” output of the currently described system embodiment, i.e., the output of stage 2 before the event is generated. Figure 8-1 is a pre-event output image showing a moving hand. Figure 8-2 is a pre-event output image showing a moving camera. Figure 8-3 is a pre-event output image where no motion was captured in the original input image.

[0082] The foregoing parts may be modified and altered without departing from the intent or scope of this disclosure.

[0083] For example, in the embodiment of the system currently being described, the time high-pass filter included in the second stage 104 may instead be a time IIR band-pass filter. This provides or improves upon more targeted noise reduction before event generation. For example, a band-pass filter can be tuned to specifically filter out moving objects at known frequencies or known frequency ranges. For example, the system currently being described, as well as a custom EBS system, picks up the flickering of fluorescent lights. In the system currently being described, a band-pass filter can be used to remove all motion at known frequencies, thus reducing the impact on the image of the fluorescent lights without requiring a special algorithm to remove this noise after the event. Another example is removing pixel changes in the frame caused by ocean waves when searching for objects at sea. Wave motion is inherently noise, and therefore a time band-pass filter can remove this noise before event generation. This selective removal of "noise" from known objects at specific frequencies reduces the amount of data generated through event generation and thus improves the efficiency of the system. This also means that the special algorithms commonly used to remove this noise after event generation become unnecessary, leading to a more efficient system. An example of a time bandpass filter is given by the following equation:

number

[0084] In equation (9), K is the scaling factor, N is the number of zeros at 0 and infinity, and p x represents the pole position of the filter.

[0085] As previously mentioned, in addition to implementing spatial filters using FFT, spatial filters can also be implemented using a pipeline structure. Here, memory is maintained for previous pixels processed by the pipeline in the previous clock, and for pixels from the current frame that have not yet been processed. Figure 9 shows an example. In this example, there are three parallel pipelines 911, 912, and 913 for processing a frame-based image output 901, which is a series of frames. In this example, the filtered value of pixel X is determined using 1) the pixel values ​​of pixels W-1, X-1, and Y-1 processed by pipelines 911, 912, and 913 in the previous clock, 2) the pixel values ​​of pixels W, X, and Y processed by pipelines 911, 912, and 913 in the current clock, and 3) the pixel values ​​of pixels W+1, X+1, and Y+1 from the current frame that have not yet been processed at the current clock cycle. Therefore, the pixel values ​​of pixels W, X, Y, W-1, X-1, and Y-1 are values ​​from the current frame, but processed in the current and previous clock cycles by the pipeline structure, while the pixel values ​​of pixels W+1, X+1, and Y+1 are values ​​from the current frame. For example,

number

[0086] In more generalized examples, the processing in the second step 104 in some embodiments may omit all high-pass spatial filters. Thus, the result from step 2 is a motion-indicating frame-based pixel output. This still has the advantage of allowing configurable compression nonlinearity to be incorporated into the system and the advantage of allowing event-based detection using any commercially available sensor.

[0087] Furthermore, embodiments may be implemented in hardware such as analog circuits, which may take the form of application-specific integrated circuits or field-programmable gate arrays. Embodiments may utilize a digital signal processor (DSP), in which the analog input signal is converted to a digital signal and then processed by the DSP. In other embodiments, a mixture of analog and digital processing may be used, provided that the analog-to-digital conversion is performed at an appropriate stage in the overall processing.

[0088] The event-based detection system disclosed herein is a novel embodiment of EBS. The method of manipulating images is parameterizable, which also allows for flexibility in the application of the event-based detection system, the lighting conditions under which the system operates, and the implementation. For example, a software implementation may be used in part or all, depending on the feasibility of incorporating hardware.

[0089] The event-based detection systems described herein have practical applications within a range of scenarios. For example, in surveillance applications, event-based detection may be used to locate moving objects in complex scenes. In defense systems, event-based detection may be used to track moving targets or objects under surveillance, such as unmanned aerial vehicles or missiles. In particular, embodiments of event-based detection systems tuned to provide "sub-pixel" level event detection have the ability to detect objects when they are far from the sensor, such as when the object's image (or data) does not yet occupy an entire pixel. This differs from detectable radar systems, where aircraft, etc., are less likely to be controlled to take evasive action to avoid detection.

[0090] Figure 10 shows a rendered image of event output processed from an image input showing a flying drone. The processing algorithm used in this example included a first-order time filter but did not include spatial filtering in the second processing stage. Crosses represent "OFF" events (locations where the drone was present but is no longer detected), and circles represent "ON" events (locations where the drone is currently detected). Here, the clusters of circles for "ON" events indicate the leading edge moving towards the upper right of the image. The clusters of crosses for "OFF" events indicate locations where the drone was present.

[0091] Another area where event-based detection is potentially useful is space applications. For example, it can be used to detect and track space debris or satellites. Embodiments in particular, where the input is provided by an infrared imaging sensor, could be useful for detecting “events” to provide space situational awareness, for example, by tracking the thermal signature of orbiting objects such as satellites.

[0092] Another area where event-based detection can be useful is in autonomous vehicle applications. For example, event-based detection may be used to detect moving objects in the path of an autonomous vehicle to provide situational awareness. Another example is lane tracking to track whether there has been relative movement between the position of the lane and the position of the vehicle. This can help detect when the vehicle has deviated from its lane or has begun to deviate. A further example is stabilization control. Event outputs can be used to algorithmically determine the movement of a camera and use this to remove the camera's own movement from the output. This eliminates the need for a gimbal.

[0093] The matters described above and in the accompanying drawings are provided for illustrative purposes only and are not limiting. While specific embodiments have been shown and described, it will be apparent to those skilled in the art that changes and modifications may be made without departing from a broader aspect of the inventors' contribution. The scope of protection actually sought is intended to be defined in the following claims from an appropriate viewpoint based on the prior art.

[0094] In the subsequent claims and the foregoing description of the invention, unless the context requires it to be interpreted in another sense for the sake of explicit wording or necessary implied meaning, the word “comprise” or variations such as “comprises” or “comprising” are used in a comprehensive sense, i.e., to identify the presence of the described features, but not to preclude the presence or addition of further features in the various embodiments of the invention.

Claims

1. This involves acquiring time-series input signal data that characterizes the information from the input source, and Applying a first processing step, wherein the first processing step comprises applying compression nonlinearity to the input signal data, Applying a second processing stage to the output from the first processing stage, wherein the second processing stage includes spatial and / or temporal processing within a feedback loop, and applies a high-pass or band-pass spatial and / or temporal filter to the output from the first processing stage, thereby suppressing recurring changes and improving the ratio of events to recurring changes in the output from the first processing stage. An image processing method comprising:

2. The method according to claim 1, further comprising detecting an event in the output of the second processing stage.

3. The method according to claim 2, wherein detecting the event comprises applying at least one threshold value to the pixel value of the output in the second processing step.

4. The method according to claim 3, wherein the at least one threshold applied includes a plurality of thresholds, and each threshold is applied to each pixel or each set of pixels.

5. The method according to any one of claims 2 to 4, comprising setting an event rate for a detected event.

6. The method according to claim 5, comprising setting two different event rates for each of two different detected events.

7. The method according to claim 5 or 6, wherein the event rate is set by setting a counter which is the number of clock cycles during which the detected event is expected to persist.

8. The method according to any one of claims 1 to 7, wherein the time high-pass or band-pass filter is an n-th order filter, where n corresponds to the number of time steps incorporated into the time processing.

9. The method according to claim 8, wherein n is equal to 1.

10. The method according to claim 8, wherein n is greater than 1.

11. The method according to any one of claims 1 to 10, wherein the time high-pass or band-pass filter is an infinite impulse response filter.

12. The method according to any one of claims 1 to 11, wherein the frequency cutoff of the time high-pass or band-pass filter is variable based on at least one or more of the characteristics of the input source and the characteristics of the environment in which the input source acquires the input data.

13. The method according to any one of claims 1 to 12, wherein the second processing step further comprises spatial processing configured to detect edges.

14. The method according to claim 13, wherein the spatial processing is configured to be performed before or after the temporal processing.

15. The method according to claim 13 or 14, wherein the spatial processing applies a high-pass filter implemented using convolution.

16. The method according to any one of claims 1 to 15, wherein the first processing step comprises a time feedback loop, in which current and previous samples from the input signal data and at least one previous sample of the output from the first processing step are used to obtain the current sample of the output from the first processing step.

17. The method according to any one of claims 1 to 16, wherein applying the compression nonlinearity to the input data comprises applying a gain to the input signal data, and the gain is variable based on the magnitude of the input signal data.

18. The method according to any one of claims 1 to 17, wherein the first processing step comprises a plurality of processing modules, each configured to favorably process input data with different characteristics.

19. The method according to claim 18, wherein the processing path comprises a first processing module for favorably processing lower size or lower contrast input data, and a second processing module for favorably processing higher size or higher contrast input data.

20. The method according to claim 19, wherein each processing module in the first processing step comprises a division-type low-pass filter.

21. The method according to any one of claims 1 to 20, wherein the input source comprises an image sensor, the time-series input data is a series of frames, each frame comprises a plurality of pixels, and the first and second processing steps are applied to each pixel.

22. The method according to claim 21, wherein the image sensor is an electro-optical sensor, an infrared sensor, or another sensor having one or more pixels.

23. A signal processing device comprising multiple processing modules, wherein the modules are A first processing module configured to receive and process a time-series input signal from an input source, and configured to apply compression nonlinearity to the received input signal, A second processing module configured to receive the output from the first processing module, and configured to perform spatial and / or temporal processing in a feedback loop, and to apply a high-pass or band-pass spatial and / or temporal filter to the output from the first processing stage, thereby suppressing recurring changes and improving the ratio of events to recurring changes in the output from the first processing stage, A signal processing device equipped with the following features.

24. The device according to claim 23, further comprising an event detection module configured to perform event detection on the output from the second processing module.

25. The device according to claim 23 or 24, wherein the second processing module is further configured to apply a spatial filtering process to detect edges in the output of the first processing module.

26. The device according to claim 24 or 25, wherein the first processing module or the second processing module, or both, are implemented in hardware.

27. The device according to any one of claims 24 to 26, wherein the input signal comprises signal data for a set of pixels, and the processing module processes the input signal pixel by pixel.