System and method for collision prediction in a dynamic scene

The implementation of an event-based vision sensor for tracking objects and shadows in dynamic scenes addresses inefficiencies in collision detection, enabling precise collision prediction and improved interaction accuracy in augmented reality and robotics.

WO2026099099A1PCT designated stage Publication Date: 2026-05-15SONY SEMICON SOLUTIONS CORP +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2025-11-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing systems struggle to accurately detect and predict collisions between objects and surfaces in dynamic scenes, particularly in augmented reality and robotics applications, due to inefficiencies in collision detection and prediction.

Method used

The use of an event-based vision sensor (EVS) to track objects and their shadows with high temporal resolution, processing event-based data to detect and predict collisions by analyzing motion vectors and shadow changes.

Benefits of technology

Enables precise detection and prediction of collisions, enhancing interaction accuracy in dynamic scenes and improving safety and precision in augmented reality and robotics systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025081618_15052026_PF_FP_ABST
    Figure EP2025081618_15052026_PF_FP_ABST
Patent Text Reader

Abstract

An imaging system comprising an event-based vision sensor, EVS, configured to obtain event-based data of a dynamic scene, wherein the EVS detects as an event the occurrence of a change in intensities of light; and a processing unit configured to process the event-based data, wherein, depending on the processed event-based data, the imaging system is configured to evaluate how a shadow changes in the dynamic scene to determine whether an object casting the shadow collides with a surface and / or another object in the dynamic scene.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 73785

[0002] 1

[0003] SYSTEM AND METHOD FOR COLLISION PREDICTION IN A DYNAMIC SCENE

[0004] FIELD OF INVENTION

[0005] The present technology relates to a system and a method for performing collision prediction in a dynamic scene, in particular, to a system and a method providing an improved estimation of a possible collision or contact between an object and a surface and / or another object in the dynamic scene.

[0006] BACKGROUND

[0007] Understanding dynamic scenes by unraveling the spatio-temporal visual patterns of a specific scenario under consideration requires a realistic interpretation of the processed scenes. Interactive systems for displaying and interacting with such a dynamic scene may use a combination of real-world data and computer-generated virtual representations of three-dimensional content. For example, Augmented Reality (AR) systems may incorporate three basic features: a combination of real and virtual worlds, real-time interaction, and accurate 3D registration of virtual and real objects. As these interactive systems advance, requirements regarding display, interaction and tracking technology also increase.

[0008] In this context, the accurate detection of a collision or contact between objects and / or surfaces as well as monitoring a proximity between objects and surfaces that may exist in both, the real- and the virtual worlds, is still an object for improvement. For example, detecting hand-object and hand-surface contact in AR systems is important for correctly interpreting human interactions. Use cases include, for example, the provision of user interfaces, where an interaction of a hand of a user with a virtual controller serves as input to an interactive system.

[0009] In addition, predicting future positions of objects and possible collisions between objects and / or surfaces may help to improve compensating latency between real-world motion and the rendering in an AR display and / or providing any general feedback like haptic or audio feedback. In greater detail, predicting collisions between objects may enhance the accuracy of future position prediction.

[0010] Moreover, predicting collisions in a dynamic scene by processing image data may also play a crucial role in other fields of technology such as robotics. For example, correctly detecting, interpreting, and forecasting a robotic arm movement and its interaction with its surroundings is necessary to guarantee a safe and accurate operation of robots, e.g. in industrial or service applications.

[0011] It is therefore desirable to improve a system and a method for performing collision detection in a dynamic scene providing an optimized estimation of a possible collision or contact between an object and a surface and / or another object in the dynamic scene.

[0012] SUMMARY OF INVENTION 73785

[0013] 2

[0014] To this end, an imaging system is provided that comprises an event-based vision sensor (EVS) configured to obtain event-based data of a dynamic scene. The EVS detects as an event the occurrence of a change in intensities of light. Furthermore, the imaging system comprises a processing unit configured to process the event-based data. More specifically, depending on the processed event-based data, the imaging system is configured to evaluate how a shadow changes in the dynamic scene to determine whether an object casting the shadow collides with a surface and / or another object in the dynamic scene.

[0015] Furthermore, a method for the imaging system is provided, which comprises the following operations: obtaining, by the EVS, event-based data of the dynamic scene, wherein the EVS detects as an event the occurrence of a change in intensities of light. The method further includes processing, by the processing unit, the event-based data, and evaluating, by the imaging system, how a shadow changes in the dynamic scene to determine whether the object casting the shadow collides with a surface and / or another object in the dynamic scene depending on the processed event-based data.

[0016] The use of a EVS sensor allows to efficiently track both the object and its shadow with high temporal resolution. By processing the generated event-based data using both object and shadow edges, it is possible to detect if the object is in collision with a surface and / or another object. Moreover, the separation between the object and its shadow can be inferred as a function of time, and by extrapolating this inference into the future, it is possible to predict the moment of collision with the surface and / or the other object. This enables detecting and determining movements of an object such as a hand or a ball in a dynamic scene more precisely. Furthermore, this allows interacting with virtual objects more accurately.

[0017] BRIEF DESCRIPTION OF DRAWINGS

[0018] Fig. 1 is a schematic diagram of a sensor device.

[0019] Fig. 2 is a schematic block diagram of a sensor section.

[0020] Fig. 3 is a schematic block diagram of a pixel array section.

[0021] Fig. 4 is a schematic circuit diagram of a pixel block.

[0022] Fig. 5 is a schematic block diagram illustrating of an event detecting section.

[0023] Fig. 6 is a schematic circuit diagram of a current-voltage converting section.

[0024] Fig. 7 is a schematic circuit diagram of a subtraction section and a quantization section.

[0025] Fig. 8 is a schematic timing chart of an example operation of the sensor section.

[0026] Fig. 9 is a schematic diagram of a frame data generation method based on event data. 3

[0027] Fig. 10 is a schematic block diagram of another quantization section.

[0028] Fig. 11 is a schematic diagram of another event detecting section.

[0029] Fig. 12 is a schematic block diagram of another pixel array section.

[0030] Fig. 13 is a schematic circuit diagram of another pixel block.

[0031] Fig. 14 is a schematic block diagram of a scan-type imaging device.

[0032] Fig. 15A is a schematic diagram showing an exemplary imaging system.

[0033] Fig. 15B is a schematic diagram showing an exemplary dynamic scene.

[0034] Fig. 16A is a schematic diagram showing the exemplary dynamic scene in more detail.

[0035] Fig. 16B is a schematic diagram showing event-based data produced by the imaging system from the dynamic scene at different points in time.

[0036] Fig. 17 is a schematic diagram showing an exemplary dynamic scene.

[0037] Fig. 18 illustrates schematically a process flow of a method for the imaging system.

[0038] Figs. 19A and 19B are schematic illustrations of a mobile device and a head mounted display comprising a sensor device.

[0039] DETAILED DESCRIPTION

[0040] The present disclosure is directed to mitigating problems regarding efficiency and accuracy of collision detection between an object and a surface and / or other objects in a dynamic scene. In particular, the problem is addressed how to improve predictions of possible collisions between the object and the surface and / or another object in the dynamic scene. The solution to this problem comprises the appliance of an eventbased vision sensor (EVS) that provides a high temporal resolution of the object in the dynamic scene and a corresponding shadow that this object casts. According to a first aspect, the EVS monitors the object and its shadow. More specifically, it detects and tracks edges of the object and its corresponding shadow. In addition, according to a second aspect, the event-based data is processed by analyzing features in the eventbased data to predict motion vectors of the object and its shadow. In greater detail, a distance between the object and its shadow is calculated as a function of time to predict a possible collision. The present disclosure is thus based on the operation of a conventional event based / dynamic vision sensor (EVS / DVS).

[0041] Thus, at first a possible implementation of an EVS / DVS will be described. This is of course purely exemplary. It is to be understood that EVSs / DVSs could also be implemented differently. 4

[0042] Fig. 1 is a diagram illustrating a configuration example of a sensor device 10, which is in the example of Fig. 1 constituted by a sensor chip.

[0043] The sensor device 10 is a single-chip semiconductor chip and includes a sensor die (substrate) 11, which serves as a plurality of dies (substrates), and a logic die 12 that are stacked. Note that, the sensor device 10 can also include only a single die or three or more stacked dies.

[0044] In the sensor device 10 of Fig. 1, the sensor die 11 includes (a circuit serving as) a sensor section 21, and the logic die 12 includes a logic section 22. Note that, the sensor section 21 can be partly formed on the logic die 12. Further, the logic section 22 can be partly formed on the sensor die 11.

[0045] The sensor section 21 includes pixels configured to perform photoelectric conversion on incident light to generate electrical signals, and generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels. The sensor section 21 supplies the event data to the logic section 22. That is, the sensor section 21 performs imaging of performing, in the pixels, photoelectric conversion on incident light to generate electrical signals, similarly to a synchronous image sensor, for example. The sensor section 21, however, generates event data indicating the occurrence of events that are changes in the electrical signal of the pixels instead of generating image data in a frame format (frame data). The sensor section 21 outputs, to the logic section 22, the event data obtained by the imaging.

[0046] Here, the synchronous image sensor is an image sensor configured to perform imaging in synchronization with a vertical synchronization signal and output frame data that is image data in a frame format. The sensor section 21 can be regarded as asynchronous (an asynchronous image sensor) in contrast to the synchronous image sensor since the sensor section 21 does not operate in synchronization with a vertical synchronization signal when outputting event data. Event detection may also be performed synchronously using a row scanner that determines which pixels generated an event during the past “frame” and assigns these events the same timestamp. Scan-type event detection is useful for moderate to high activity scenes since it incurs less readout overhead than arbiter-type, i.e. asynchronous, event detection.

[0047] Note that, the sensor section 21 can generate and output, other than event data, frame data, similarly to the synchronous image sensor. In addition, the sensor section 21 can output, together with event data, electrical signals of pixels in which events have occurred, as pixel signals that are pixel values of the pixels in frame data.

[0048] The logic section 22 controls the sensor section 21 as needed. Further, the logic section 22 performs various types of data processing, such as data processing of generating frame data on the basis of event data from the sensor section 21 and image processing on frame data from the sensor section 21 or frame data generated on the basis of the event data from the sensor section 21, and outputs data processing results obtained by performing the various types of data processing on the event data and the frame data.

[0049] Fig. 2 is a block diagram illustrating a configuration example of the sensor section 21 of Fig. 1. 5

[0050] The sensor section 21 includes a pixel array section 31, a driving section 32, an arbiter 33, an AD (Analog to Digital) conversion section 34, and an output section 35.

[0051] The pixel array section 31 includes a plurality of pixels 51 (Fig. 3) arrayed in a two-dimensional lattice pattern. The pixel array section 31 detects, in a case where a change larger than a predetermined threshold (including a change equal to or larger than the threshold as needed) has occurred in (a voltage corresponding to) a photocurrent that is an electrical signal generated by photoelectric conversion in the pixel 51, the change in the photocurrent as an event. In a case of detecting an event, the pixel array section 31 outputs, to the arbiter 33, a request for requesting the output of event data indicating the occurrence of the event. Then, in a case of receiving a response indicating event data output permission from the arbiter 33, the pixel array section 31 outputs the event data to the driving section 32 and the output section 35. In addition, the pixel array section 31 outputs an electrical signal of the pixel 51 in which the event has been detected to the AD conversion section 34, as a pixel signal.

[0052] The driving section 32 supplies control signals to the pixel array section 31 to drive the pixel array section 31. For example, the driving section 32 drives the pixel 51 regarding which the pixel array section 31 has output event data, so that the pixel 51 in question supplies (outputs) a pixel signal to the AD conversion section 34.

[0053] The arbiter 33 arbitrates the requests for requesting the output of event data from the pixel array section 31, and returns responses indicating event data output permission or prohibition to the pixel array section 31.

[0054] The AD conversion section 34 includes, for example, a single-slope ADC (AD converter) (not illustrated) in each column of pixel blocks 41 (Fig. 3) described later, for example. The AD conversion section 34 performs, with the ADC in each column, AD conversion on pixel signals of the pixels 51 of the pixel blocks 41 in the column, and supplies the resultant to the output section 35. Note that, the AD conversion section 34 can perform CDS (Correlated Double Sampling) together with pixel signal AD conversion.

[0055] The output section 35 performs necessary processing on the pixel signals from the AD conversion section 34 and the event data from the pixel array section 31 and supplies the resultant to the logic section 22 (Fig. 1).

[0056] Here, a change in the photocurrent generated in the pixel 51 can be recognized as a change in the amount of light entering the pixel 51, so that it can also be said that an event is a change in light amount (a change in light amount larger than the threshold) in the pixel 51.

[0057] Event data indicating the occurrence of an event at least includes location information (coordinates or the like) indicating the location of a pixel block in which a change in light amount, which is the event, has occurred. Besides, the event data can also include the polarity (positive or negative) of the change in light amount. 6

[0058] With regard to the series of event data that is output from the pixel array section 31 at timings at which events have occurred, it can be said that, as long as the event data interval is the same as the event occurrence interval, the event data implicitly includes time point information indicating (relative) time points at which the events have occurred. However, for example, when the event data is stored in a memory and the event data interval is no longer the same as the event occurrence interval, the time point information implicitly included in the event data is lost. Thus, the output section 35 includes, in event data, time point information indicating (relative) time points at which events have occurred, such as timestamps, before the event data interval is changed from the event occurrence interval. The processing of including time point information in event data can be performed in any block other than the output section 35 as long as the processing is performed before time point information implicitly included in event data is lost.

[0059] Fig. 3 is a block diagram illustrating a configuration example of the pixel array section 31 of Fig. 2.

[0060] The pixel array section 31 includes the plurality of pixel blocks 41. The pixel block 41 includes the IxJ pixels 51 that are one or more pixels arrayed in I rows and J columns (I and J are integers), an event detecting section 52, and a pixel signal generating section 53. The one or more pixels 51 in the pixel block 41 share the event detecting section 52 and the pixel signal generating section 53. Further, in each column of the pixel blocks 41, a VSL (Vertical Signal Line) for connecting the pixel blocks 41 to the ADC of the AD conversion section 34 is wired.

[0061] The pixel 51 receives light incident from an object and performs photoelectric conversion to generate a photocurrent serving as an electrical signal. The pixel 51 supplies the photocurrent to the event detecting section 52 under the control of the driving section 32.

[0062] The event detecting section 52 detects, as an event, a change larger than the predetermined threshold in photocurrent from each of the pixels 51, under the control of the driving section 32. In a case of detecting an event, the event detecting section 52 supplies, to the arbiter 33 (Fig. 2), a request for requesting the output of event data indicating the occurrence of the event. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the event detecting section 52 outputs the event data to the driving section 32 and the output section 35.

[0063] The pixel signal generating section 53 generates, in the case where the event detecting section 52 has detected an event, a voltage corresponding to a photocurrent from the pixel 51 as a pixel signal, and supplies the voltage to the AD conversion section 34 through the VSL, under the control of the driving section 32.

[0064] Here, detecting a change larger than the predetermined threshold in photocurrent as an event can also be recognized as detecting, as an event, absence of change larger than the predetermined threshold in photocurrent. The pixel signal generating section 53 can generate a pixel signal in the case where absence of change larger than the predetermined threshold in photocurrent has been detected as an event as well as in the case where a change larger than the predetermined threshold in photocurrent has been detected as an event. 7

[0065] Fig. 4 is a circuit diagram illustrating a configuration example of the pixel block 41.

[0066] The pixel block 41 includes, as described with reference to Fig. 3, the pixels 51, the event detecting section 52, and the pixel signal generating section 53.

[0067] The pixel 51 includes a photoelectric conversion element 61 and transfer transistors 62 and 63.

[0068] The photoelectric conversion element 61 includes, for example, a PD (Photodiode). The photoelectric conversion element 61 receives incident light and performs photoelectric conversion to generate charges.

[0069] The transfer transistor 62 includes, for example, an N (Negative)-type MOS (Metal-Oxide-Semiconductor) FET (Field Effect Transistor). The transfer transistor 62 of the n-th pixel 51 of the lx J pixels 51 in the pixel block 41 is turned on or off in response to a control signal OFGn supplied from the driving section 32 (Fig. 2). When the transfer transistor 62 is turned on, charges generated in the photoelectric conversion element 61 are transferred (supplied) to the event detecting section 52, as a photocurrent.

[0070] The transfer transistor 63 includes, for example, an N-type MOSFET. The transfer transistor 63 of the n-th pixel 51 of the IxJ pixels 51 in the pixel block 41 is turned on or off in response to a control signal TRGn supplied from the driving section 32. When the transfer transistor 63 is turned on, charges generated in the photoelectric conversion element 61 are transferred to an FD 74 of the pixel signal generating section 53.

[0071] The IxJ pixels 51 in the pixel block 41 are connected to the event detecting section 52 of the pixel block 41 through nodes 60. Thus, photocurrents generated in (the photoelectric conversion elements 61 of) the pixels 51 are supplied to the event detecting section 52 through the nodes 60. As a result, the event detecting section 52 receives the sum of photocurrents from all the pixels 51 in the pixel block 41. Thus, the event detecting section 52 detects, as an event, a change in sum of photocurrents supplied from the IxJ pixels 51 in the pixel block 41.

[0072] The pixel signal generating section 53 includes a reset transistor 71, an amplification transistor 72, a selection transistor 73, and the FD (Floating Diffusion) 74.

[0073] The reset transistor 71, the amplification transistor 72, and the selection transistor 73 include, for example, N-type MOSFETs.

[0074] The reset transistor 71 is turned on or off in response to a control signal RST supplied from the driving section 32 (Fig. 2). When the reset transistor 71 is turned on, the FD 74 is connected to a power supply VDD, and charges accumulated in the FD 74 are thus discharged to the power supply VDD. With this, the FD 74 is reset.

[0075] The amplification transistor 72 has a gate connected to the FD 74, a drain connected to the power supply VDD, and a source connected to the VSL through the selection transistor 73. The amplification transistor 8

[0076] 72 is a source follower and outputs a voltage (electrical signal) corresponding to the voltage of the FD 74 supplied to the gate to the VSL through the selection transistor 73.

[0077] The selection transistor 73 is turned on or off in response to a control signal SEL supplied from the driving section 32. When the selection transistor 73 is turned on, a voltage corresponding to the voltage of the FD 74 from the amplification transistor 72 is output to the VSL.

[0078] The FD 74 accumulates charges transferred from the photoelectric conversion elements 61 of the pixels 51 through the transfer transistors 63, and converts the charges to voltages.

[0079] With regard to the pixels 51 and the pixel signal generating section 53, which are configured as described above, the driving section 32 turns on the transfer transistors 62 with control signals OFGn, so that the transfer transistors 62 supply, to the event detecting section 52, photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51. With this, the event detecting section 52 receives a current that is the sum of the photocurrents from all the pixels 51 in the pixel block 41, which might also be only a single pixel.

[0080] When the event detecting section 52 detects, as an event, a change in photocurrent (sum of photocurrents) in the pixel block 41, the driving section 32 turns off the transfer transistors 62 of all the pixels 51 in the pixel block 41, to thereby stop the supply of the photocurrents to the event detecting section 52. Then, the driving section 32 sequentially turns on, with the control signals TRGn, the transfer transistors 63 of the pixels 51 in the pixel block 41 in which the event has been detected, so that the transfer transistors 63 transfers charges generated in the photoelectric conversion elements 61 to the FD 74. The FD 74 accumulates the charges transferred from (the photoelectric conversion elements 61 of) the pixels 51. Voltages corresponding to the charges accumulated in the FD 74 are output to the VSL, as pixel signals of the pixels 51, through the amplification transistor 72 and the selection transistor 73.

[0081] As described above, in the sensor section 21 (Fig. 2), only pixel signals of the pixels 51 in the pixel block 41 in which an event has been detected are sequentially output to the VSL. The pixel signals output to the VSL are supplied to the AD conversion section 34 to be subjected to AD conversion.

[0082] Here, in the pixels 51 in the pixel block 41, the transfer transistors 63 can be turned on not sequentially but simultaneously. In this case, the sum of pixel signals of all the pixels 51 in the pixel block 41 can be output.

[0083] In the pixel array section 31 of Fig. 3, the pixel block 41 includes one or more pixels 51, and the one or more pixels 51 share the event detecting section 52 and the pixel signal generating section 53. Thus, in the case where the pixel block 41 includes a plurality of pixels 51, the numbers of the event detecting sections 52 and the pixel signal generating sections 53 can be reduced as compared to a case where the event detecting section 52 and the pixel signal generating section 53 are provided for each of the pixels 51, with the result that the scale of the pixel array section 31 can be reduced. 9

[0084] Note that, in the case where the pixel block 41 includes a plurality of pixels 51, the event detecting section 52 can be provided for each of the pixels 51. In the case where the plurality of pixels 51 in the pixel block 41 share the event detecting section 52, events are detected in units of the pixel blocks 41. In the case where the event detecting section 52 is provided for each of the pixels 51, however, events can be detected in units of the pixels 51.

[0085] Yet, even in the case where the plurality of pixels 51 in the pixel block 41 share the single event detecting section 52, events can be detected in units of the pixels 51 when the transfer transistors 62 of the plurality of pixels 51 are temporarily turned on in a time-division manner.

[0086] Further, in a case where there is no need to output pixel signals, the pixel block 41 can be formed without the pixel signal generating section 53. In the case where the pixel block 41 is formed without the pixel signal generating section 53, the sensor section 21 can be formed without the AD conversion section 34 and the transfer transistors 63. In this case, the scale of the sensor section 21 can be reduced. The sensor will then output the address of the pixel (block) in which the event occurred, if necessary with a time stamp.

[0087] Fig. 5 is a block diagram illustrating a configuration example of the event detecting section 52 of Fig. 3.

[0088] The event detecting section 52 includes a current-voltage converting section 81, a buffer 82, a subtraction section 83, a quantization section 84, and a transfer section 85.

[0089] The current-voltage converting section 81 converts (a sum of) photocurrents from the pixels 51 to voltages corresponding to the logarithms of the photocurrents (hereinafter also referred to as a "photovoltage") and supplies the voltages to the buffer 82.

[0090] The buffer 82 buffers photovoltages from the current-voltage converting section 81 and supplies the resultant to the subtraction section 83.

[0091] The subtraction section 83 calculates, at a timing instructed by a row driving signal that is a control signal from the driving section 32, a difference between the current photovoltage and a photovoltage at a timing slightly shifted from the current time, and supplies a difference signal corresponding to the difference to the quantization section 84.

[0092] The quantization section 84 quantizes difference signals from the subtraction section 83 to digital signals and supplies the quantized values of the difference signals to the transfer section 85 as event data.

[0093] The transfer section 85 transfers (outputs), on the basis of event data from the quantization section 84, the event data to the output section 35. That is, the transfer section 85 supplies a request for requesting the output of the event data to the arbiter 33. Then, when receiving a response indicating event data output permission to the request from the arbiter 33, the transfer section 85 outputs the event data to the output section 35. 73785

[0094] 10

[0095] Fig. 6 is a circuit diagram illustrating a configuration example of the current-voltage converting section 81 of Fig. 5.

[0096] The current-voltage converting section 81 includes transistors 91 to 93. As the transistors 91 and 93, for example, N-type MOSFETs can be employed. As the transistor 92, for example, a P-type MOSFET can be employed.

[0097] The transistor 91 has a source connected to the gate of the transistor 93, and a photocurrent is supplied from the pixel 51 to the connecting point between the source of the transistor 91 and the gate of the transistor 93. The transistor 91 has a drain connected to the power supply VDD and a gate connected to the drain of the transistor 93.

[0098] The transistor 92 has a source connected to the power supply VDD and a drain connected to the connecting point between the gate of the transistor 91 and the drain of the transistor 93. A predetermined bias voltage Vbias is applied to the gate of the transistor 92. With the bias voltage Vbias, the transistor 92 is turned on or off, and the operation of the current-voltage converting section 81 is turned on or off depending on whether the transistor 92 is turned on or off.

[0099] The source of the transistor 93 is grounded.

[0100] In the current-voltage converting section 81, the transistor 91 has the drain connected on the power supply VDD side and is thus a source follower. The source of the transistor 91, which is the source follower, is connected to the pixels 51 (Fig. 4), so that photocurrents based on charges generated in the photoelectric conversion elements 61 of the pixels 51 flow through the transistor 91 (from the drain to the source). The transistor 91 operates in a subthreshold region, and at the gate of the transistor 91, photovoltages corresponding to the logarithms of the photocurrents flowing through the transistor 91 are generated. As described above, in the current-voltage converting section 81, the transistor 91 converts photocurrents from the pixels 51 to photovoltages corresponding to the logarithms of the photocurrents.

[0101] In the current-voltage converting section 81 , the transistor 91 has the gate connected to the connecting point between the drain of the transistor 92 and the drain of the transistor 93, and the photovoltages are output from the connecting point in question.

[0102] Fig. 7 is a circuit diagram illustrating configuration examples of the subtraction section 83 and the quantization section 84 of Fig. 5.

[0103] The subtraction section 83 includes a capacitor 101, an operational amplifier 102, a capacitor 103, and a switch 104. The quantization section 84 includes a comparator 111.

[0104] The capacitor 101 has one end connected to the output terminal of the buffer 82 (Fig. 5) and the other end connected to the input terminal (inverting input terminal) of the operational amplifier 102. Thus, photovoltages are input to the input terminal of the operational amplifier 102 through the capacitor 101. 73785

[0105] 11

[0106] The operational amplifier 102 has an output terminal connected to the non-inverting input terminal (+) of the comparator 111.

[0107] The capacitor 103 has one end connected to the input terminal of the operational amplifier 102 and the other end connected to the output terminal of the operational amplifier 102.

[0108] The switch 104 is connected to the capacitor 103 to switch the connections between the ends of the capacitor 103. The switch 104 is turned on or off in response to a row driving signal that is a control signal from the driving section 32, to thereby switch the connections between the ends of the capacitor 103.

[0109] A photovoltage on the buffer 82 (Fig. 5) side of the capacitor 101 when the switch 104 is on is denoted by Vinit, and the capacitance (electrostatic capacitance) of the capacitor 101 is denoted by Cl. The input terminal of the operational amplifier 102 serves as a virtual ground terminal, and a charge Qinit that is accumulated in the capacitor 101 in the case where the switch 104 is on is expressed by Expression (1).

[0110] Qinit = Cl x Vinit (1)

[0111] Further, in the case where the switch 104 is on, the connection between the ends of the capacitor 103 is cut (short-circuited), so that no charge is accumulated in the capacitor 103.

[0112] When a photovoltage on the buffer 82 (Fig. 5) side of the capacitor 101 in the case where the switch 104 has thereafter been turned off is denoted by Vafter, a charge Qafter that is accumulated in the capacitor 101 in the case where the switch 104 is off is expressed by Expression (2).

[0113] Qafter = C 1 x Vafter (2)

[0114] When the capacitance of the capacitor 103 is denoted by C2 and the output voltage of the operational amplifier 102 is denoted by Vout, a charge Q2 that is accumulated in the capacitor 103 is expressed by Expression (3).

[0115] Q2 = -C2 x Vout (3)

[0116] Since the total amount of charges in the capacitors 101 and 103 does not change before and after the switch 104 is turned off, Expression (4) is established.

[0117] Qinit = Qafter + Q2 (4)

[0118] When Expression (1) to Expression (3) are substituted for Expression (4), Expression (5) is obtained.

[0119] Vout = -(C1 / C2) x (Vafter - Vinit) (5) 73785

[0120] 12

[0121] With Expression (5), the subtraction section 83 subtracts the photovoltage Vinit from the photovoltage Vafter, that is, calculates the difference signal (Vout) corresponding to a difference Vafter - Vinit between the photovoltages Vafter and Vinit. With Expression (5), the subtraction gain of the subtraction section 83 is C1 / C2. Since the maximum gain is normally desired, Cl is preferably set to a large value and C2 is preferably set to a small value. Meanwhile, when C2 is too small, kTC noise increases, resulting in a risk of deteriorated noise characteristics. Thus, the capacitance C2 can only be reduced in a range that achieves acceptable noise. Further, since the pixel blocks 41 each have installed therein the event detecting section 52 including the subtraction section 83, the capacitances Cl and C2 have space constraints. In consideration of these matters, the values of the capacitances Cl and C2 are determined.

[0122] The comparator 111 compares a difference signal from the subtraction section 83 with a predetermined threshold (voltage) Vth (>0) applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 111 outputs the quantized value obtained by the quantization to the transfer section 85 as event data.

[0123] For example, in a case where a difference signal is larger than the threshold Vth, the comparator 111 outputs an H (High) level indicating 1, as event data indicating the occurrence of an event. In a case where a difference signal is not larger than the threshold Vth, the comparator 111 outputs an L (Low) level indicating 0, as event data indicating that no event has occurred.

[0124] The transfer section 85 supplies a request to the arbiter 33 in a case where it is confirmed on the basis of event data from the quantization section 84 that a change in light amount that is an event has occurred, that is, in the case where the difference signal (Vout) is larger than the threshold Vth. When receiving a response indicating event data output permission, the transfer section 85 outputs the event data indicating the occurrence of the event (for example, H level) to the output section 35.

[0125] The output section 35 includes, in event data from the transfer section 85, location / address information regarding (the pixel block 41 including) the pixel 51 in which an event indicated by the event data has occurred and time point information indicating a time point at which the event has occurred, and further, as needed, the polarity of a change in light amount that is the event, i.e. whether the intensity did increase or decrease. The output section 35 outputs the event data.

[0126] As the data format of event data including location information regarding the pixel 51 in which an event has occurred, time point information indicating a time point at which the event has occurred, and the polarity of a change in light amount that is the event, for example, the data format called "AER (Address Event Representation)" can be employed.

[0127] Note that, a gain A of the entire event detecting section 52 is expressed by the following expression where the gain of the current-voltage converting section 81 is denoted by CGlog and the gain of the buffer 82 is 1.

[0128] A = CGlogCl / C2 (Eiphoto n) (6) 73785

[0129] 13

[0130] Here, iphoto n denotes a photocurrent of the n-th pixel 51 of the I* J pixels 51 in the pixel block 41. In Expression (6), E denotes the summation of n that takes integers ranging from 1 to IxJ.

[0131] Note that, the pixel 51 can receive any light as incident light with an optical fdter through which predetermined light passes, such as a color filter. For example, in a case where the pixel 51 receives visible light as incident light, event data indicates the occurrence of changes in pixel value in images including visible objects. Further, for example, in a case where the pixel 51 receives, as incident light, infrared light, millimeter waves, or the like for ranging, event data indicates the occurrence of changes in distances to objects. In addition, for example, in a case where the pixel 51 receives infrared light for temperature measurement, as incident light, event data indicates the occurrence of changes in temperature of objects. In the present embodiment, the pixel 51 is assumed to receive visible light as incident light.

[0132] Fig. 8 is a timing chart illustrating an example of the operation of the sensor section 21 of Fig. 2.

[0133] At Timing TO, the driving section 32 changes all the control signals OFGn from the L level to the H level, thereby turning on the transfer transistors 62 of all the pixels 51 in the pixel block 41 . With this, the sum of photocurrents from all the pixels 51 in the pixel block 41 is supplied to the event detecting section 52. Here, the control signals TRGn are all at the L level, and hence the transfer transistors 63 of all the pixels 51 are off.

[0134] For example, at Timing Tl, when detecting an event, the event detecting section 52 outputs event data at the H level in response to the detection of the event.

[0135] At Timing T2, the driving section 32 sets all the control signals OFGn to the L level on the basis of the event data at the H level, to stop the supply of the photocurrents from the pixels 51 to the event detecting section 52. Further, the driving section 32 sets the control signal SEL to the H level, and sets the control signal RST to the H level over a certain period of time, to control the FD 74 to discharge the charges to the power supply VDD, thereby resetting the FD 74. The pixel signal generating section 53 outputs, as a reset level, a pixel signal corresponding to the voltage of the FD 74 when the FD 74 has been reset, and the AD conversion section 34 performs AD conversion on the reset level.

[0136] At Timing T3 after the reset level AD conversion, the driving section 32 sets a control signal TRG1 to the H level over a certain period to control the first pixel 51 in the pixel block 41 in which the event has been detected to transfer, to the FD 74, charges generated by photoelectric conversion in (the photoelectric conversion element 61 of) the first pixel 51. The pixel signal generating section 53 outputs, as a signal level, a pixel signal corresponding to the voltage of the FD 74 to which the charges have been transferred from the pixel 51, and the AD conversion section 34 performs AD conversion on the signal level.

[0137] The AD conversion section 34 outputs, to the output section 35, a difference between the signal level and the reset level obtained after the AD conversion, as a pixel signal serving as a pixel value of the image (frame data). 73785

[0138] 14

[0139] Here, the processing of obtaining a difference between a signal level and a reset level as a pixel signal serving as a pixel value of an image is called "CDS." CDS can be performed after the AD conversion of a signal level and a reset level, or can be simultaneously performed with the AD conversion of a signal level and a reset level in a case where the AD conversion section 34 performs single-slope AD conversion. In the latter case, AD conversion is performed on the signal level by using the AD conversion result of the reset level as an initial value.

[0140] At Timing T4 after the AD conversion of the pixel signal of the first pixel 51 in the pixel block 41, the driving section 32 sets a control signal TRG2 to the H level over a certain period of time to control the second pixel 51 in the pixel block 41 in which the event has been detected to output a pixel signal.

[0141] In the sensor section 21, similar processing is executed thereafter, so that pixel signals of the pixels 51 in the pixel block 41 in which the event has been detected are sequentially output.

[0142] When the pixel signals of all the pixels 51 in the pixel block 41 are output, the driving section 32 sets all the control signals OFGn to the H level to turn on the transfer transistors 62 of all the pixels 51 in the pixel block 41.

[0143] Fig. 9 is a diagram illustrating an example of a frame data generation method based on event data.

[0144] The logic section 22 sets a frame interval and a frame width on the basis of an externally input command, for example. Here, the frame interval represents the interval of frames of frame data that is generated on the basis of event data. The frame width represents the time width of event data that is used for generating frame data on a single frame. A frame interval and a frame width that are set by the logic section 22 are also referred to as a "set frame interval" and a "set frame width," respectively.

[0145] The logic section 22 generates, on the basis of the set frame interval, the set frame width, and event data from the sensor section 21, frame data that is image data in a frame format, to thereby convert the event data to the frame data.

[0146] That is, the logic section 22 generates, in each set frame interval, frame data on the basis of event data in the set frame width from the beginning of the set frame interval.

[0147] Here, it is assumed that event data includes time point information ti indicating a time point at which an event has occurred (hereinafter also referred to as an "event time point") and coordinates (x, y) serving as location information regarding (the pixel block 41 including) the pixel 51 in which the event has occurred (hereinafter also referred to as an "event location").

[0148] In Fig. 9, in a three-dimensional space (time and space) with the x axis, the y axis, and the time axis t, points representing event data are plotted on the basis of the event time point t and the event location (coordinates) (x, y) included in the event data. 73785

[0149] 15

[0150] That is, when a location (x, y, t) on the three-dimensional space indicated by the event time point t and the event location (x, y) included in event data is regarded as the space-time location of an event, in Fig. 9, the points representing the event data are plotted on the space-time locations (x, y, t) of the events.

[0151] The logic section 22 starts to generate frame data on the basis of event data by using, as a generation start time point at which frame data generation starts, a predetermined time point, for example, a time point at which frame data generation is externally instructed or a time point at which the sensor device 10 is powered on.

[0152] Here, cuboids each having the set frame width in the direction of the time axis t in the set frame intervals, which appear from the generation start time point, are referred to as a "frame volume." The size of the frame volume in the x-axis direction or the y-axis direction is equal to the number of the pixel blocks 41 or the pixels 51 in the x-axis direction or the y-axis direction, for example.

[0153] The logic section 22 generates, in each set frame interval, frame data on a single frame on the basis of event data in the frame volume having the set frame width from the beginning of the set frame interval.

[0154] Frame data can be generated by, for example, setting white to a pixel (pixel value) in a frame at the event location (x, y) included in event data and setting a predetermined color such as gray to pixels at other locations in the frame.

[0155] Besides, in a case where event data includes the polarity of a change in light amount that is an event, frame data can be generated in consideration of the polarity included in the event data. For example, white can be set to pixels in the case a positive polarity, while black can be set to pixels in the case of a negative polarity.

[0156] In addition, in the case where pixel signals of the pixels 51 are also output when event data is output as described with reference to Fig. 3 and Fig. 4, frame data can be generated on the basis of the event data by using the pixel signals of the pixels 51. That is, frame data can be generated by setting, in a frame, a pixel at the event location (x, y) (in a block corresponding to the pixel block 41) included in event data to a pixel signal of the pixel 51 at the location (x, y) and setting a predetermined color such as gray to pixels at other locations.

[0157] Note that, in the frame volume, there are a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) in some cases. In this case, for example, event data at the latest or oldest event time point t can be prioritized. Further, in the case where event data includes polarities, the polarities of a plurality of pieces of event data that are different in the event time point t but the same in the event location (x, y) can be added together, and a pixel value based on the added value obtained by the addition can be set to a pixel at the event location (x, y). 73785

[0158] 16

[0159] Here, in a case where the frame width and the frame interval are the same, the frame volumes are adjacent to each other without any gap. Further, in a case where the frame interval is larger than the frame width, the frame volumes are arranged with gaps. In a case where the frame width is larger than the frame interval, the frame volumes are arranged to be partly overlapped with each other.

[0160] Fig. 10 is a block diagram illustrating another configuration example of the quantization section 84 of Fig. 5.

[0161] Note that, in Fig. 10, parts corresponding to those in the case of Fig. 7 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.

[0162] In Fig. 10, the quantization section 84 includes comparators 111 and 112 and an output section 113.

[0163] Thus, the quantization section 84 of Fig. 10 is similar to the case of Fig. 7 in including the comparator 111. However, the quantization section 84 of Fig. 10 is different from the case of Fig. 7 in newly including the comparator 112 and the output section 113.

[0164] The event detecting section 52 (Fig. 5) including the quantization section 84 of Fig. 10 detects, in addition to events, the polarities of changes in light amount that are events.

[0165] In the quantization section 84 of Fig. 10, the comparator 111 outputs, in the case where a difference signal is larger than the threshold Vth, the H level indicating 1, as event data indicating the occurrence of an event having the positive polarity. The comparator 111 outputs, in the case where a difference signal is not larger than the threshold Vth, the L level indicating 0, as event data indicating that no event having the positive polarity has occurred.

[0166] Further, in the quantization section 84 of Fig. 10, a threshold Vth' (<Vth) is supplied to the non-inverting input terminal (+) of the comparator 112, and difference signals are supplied to the inverting input terminal (-) of the comparator 112 from the subtraction section 83. Here, for the sake of simple description, it is assumed that the threshold Vth' is equal to -Vth, for example, which needs however not to be the case.

[0167] The comparator 112 compares a difference signal from the subtraction section 83 with the threshold Vth' applied to the inverting input terminal (-), thereby quantizing the difference signal. The comparator 112 outputs, as event data, the quantized value obtained by the quantization.

[0168] For example, in a case where a difference signal is smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is larger than the threshold Vth), the comparator 112 outputs the H level indicating 1, as event data indicating the occurrence of an event having the negative polarity. Further, in a case where a difference signal is not smaller than the threshold Vth' (the absolute value of the difference signal having a negative value is not larger than the threshold Vth), the comparator 112 outputs the L level indicating 0, as event data indicating that no event having the negative polarity has occurred. 73785

[0169] 17

[0170] The output section 113 outputs, on the basis of event data output from the comparators 111 and 112, event data indicating the occurrence of an event having the positive polarity, event data indicating the occurrence of an event having the negative polarity, or event data indicating that no event has occurred to the transfer section 85.

[0171] For example, the output section 113 outputs, in a case where event data from the comparator 111 is the H level indicating 1, +V volts indicating +1, as event data indicating the occurrence of an event having the positive polarity, to the transfer section 85. Further, the output section 113 outputs, in a case where event data from the comparator 112 is the H level indicating 1, -V volts indicating -1, as event data indicating the occurrence of an event having the negative polarity, to the transfer section 85. In addition, the output section 113 outputs, in a case where each event data from the comparators 111 and 112 is the L level indicating 0, 0 volts (GND level) indicating 0, as event data indicating that no event has occurred, to the transfer section 85.

[0172] The transfer section 85 supplies a request to the arbiter 33 in the case where it is confirmed on the basis of event data from the output section 113 of the quantization section 84 that a change in light amount that is an event having the positive polarity or the negative polarity has occurred. After receiving a response indicating event data output permission, the transfer section 85 outputs event data indicating the occurrence of the event having the positive polarity or the negative polarity (+V volts indicating 1 or -V volts indicating -1) to the output section 35.

[0173] Preferably, the quantization section 84 has a configuration as illustrated in Fig. 10.

[0174] Fig. 11 is a diagram illustrating another configuration example of the event detecting section 52.

[0175] In Fig. 11, the event detecting section 52 includes a subtractor 430, a quantizer 440, a memory 451, and a controller 452. The subtractor 430 and the quantizer 440 correspond to the subtraction section 83 and the quantization section 84, respectively.

[0176] Note that, in Fig. 11, the event detecting section 52 further includes blocks corresponding to the currentvoltage converting section 81 and the buffer 82, but the illustrations of the blocks are omitted in Fig. 11.

[0177] The subtractor 430 includes a capacitor 431, an operational amplifier 432, a capacitor 433, and a switch 434. The capacitor 431, the operational amplifier 432, the capacitor 433, and the switch 434 correspond to the capacitor 101, the operational amplifier 102, the capacitor 103, and the switch 104, respectively.

[0178] The quantizer 440 includes a comparator 441. The comparator 441 corresponds to the comparator 111.

[0179] The comparator 441 compares a voltage signal (difference signal) from the subtractor 430 with the predetermined threshold voltage Vth applied to the inverting input terminal (-). The comparator 441 outputs a signal indicating the comparison result, as a detection signal (quantized value). The voltage signal from the subtractor 430 may be input to the input terminal (-) of the comparator 441, and the predetermined threshold voltage Vth may be input to the input terminal (+) of the comparator 441.

[0180] The controller 452 supplies the predetermined threshold voltage Vth applied to the inverting input terminal (-) of the comparator 441. The threshold voltage Vth which is supplied may be changed in a time-division manner. For example, the controller 452 supplies a threshold voltage Vthl corresponding to ON events (for example, positive changes in photocurrent) and a threshold voltage Vth2 corresponding to OFF events (for example, negative changes in photocurrent) at different timings to allow the single comparator to detect a plurality of types of address events (events).

[0181] The memory 451 accumulates output from the comparator 441 on the basis of Sample signals supplied from the controller 452. The memory 451 may be a sampling circuit, such as a switch, plastic, or capacitor, or a digital memory circuit, such as a latch or flip-flop. For example, the memory 451 may hold, in a period in which the threshold voltage Vth2 corresponding to OFF events is supplied to the inverting input terminal (-) of the comparator 441, the result of comparison by the comparator 441 using the threshold voltage Vthl corresponding to ON events. Note that, the memory 451 may be omitted, may be provided inside the pixel (pixel block 41), or may be provided outside the pixel.

[0182] Fig. 12 is a block diagram illustrating another configuration example of the pixel array section 31 of Fig. 2.

[0183] Note that, in Fig. 12, parts corresponding to those in the case of Fig. 3 are denoted by the same reference signs, and the description thereof is omitted as appropriate below.

[0184] In Fig. 12, the pixel array section 31 includes the plurality of pixel blocks 41. The pixel block 41 includes the lx J pixels 51 that are one or more pixels and the event detecting section 52.

[0185] Thus, the pixel array section 31 of Fig. 12 is similar to the case of Fig. 3 in that the pixel array section 31 includes the plurality of pixel blocks 41 and that the pixel block 41 includes one or more pixels 51 and the event detecting section 52. However, the pixel array section 31 of Fig. 12 is different from the case of Fig. 3 in that the pixel block 41 does not include the pixel signal generating section 53.

[0186] As described above, in the pixel array section 31 of Fig. 12, the pixel block 41 does not include the pixel signal generating section 53, so that the sensor section 21 (Fig. 2) can be formed without the AD conversion section 34.

[0187] Fig. 13 is a circuit diagram illustrating a configuration example of the pixel block 41 of Fig. 12.

[0188] As described with reference to Fig. 12, the pixel block 41 includes the pixels 51 and the event detecting section 52, but does not include the pixel signal generating section 53. 73785

[0189] 19

[0190] In this case, the pixel 51 can only include the photoelectric conversion element 61 without the transfer transistors 62 and 63.

[0191] Note that, in the case where the pixel 51 has the configuration illustrated in Fig. 13, the event detecting section 52 can output a voltage corresponding to a photocurrent from the pixel 51, as a pixel signal.

[0192] Above, the sensor device 10 was described to be an asynchronous imaging device configured to read out events by the asynchronous readout system. However, the event readout system is not limited to the asynchronous readout system and may be the synchronous readout system. An imaging device to which the synchronous readout system is applied is a scan type imaging device that is the same as a general imaging device configured to perform imaging at a predetermined frame rate.

[0193] Fig. 14 is a block diagram illustrating a configuration example of a scan type imaging device.

[0194] As illustrated in Fig. 14, an imaging device 510 includes a pixel array section 521, a driving section 522, a signal processing section 525, a read-out region selecting section 527, and a signal generating section 528.

[0195] The pixel array section 521 includes a plurality of pixels 530. The plurality of pixels 530 each output an output signal in response to a selection signal from the read-out region selecting section 527. The plurality of pixels 530 can each include an in-pixel quantizer as illustrated in Fig. 11, for example. The plurality of pixels 530 output output signals corresponding to the amounts of change in light intensity. The plurality of pixels 530 may be two-dimensionally disposed in a matrix as illustrated in Fig. 14.

[0196] The driving section 522 drives the plurality of pixels 530, so that the pixels 530 output pixel signals generated in the pixels 530 to the signal processing section 525 through an output line 514. Note that, the driving section 522 and the signal processing section 525 are circuit sections for acquiring grayscale information. Thus, in a case where only event information (event data) is acquired, the driving section 522 and the signal processing section 525 may be omitted.

[0197] The read-out region selecting section 527 selects some of the plurality of pixels 530 included in the pixel array section 521. For example, the read-out region selecting section 527 selects one or a plurality of rows included in the two-dimensional matrix structure corresponding to the pixel array section 521. The readout region selecting section 527 sequentially selects one or a plurality of rows on the basis of a cycle set in advance. Further, the read-out region selecting section 527 may determine a selection region on the basis of requests from the pixels 530 in the pixel array section 521.

[0198] The signal generating section 528 generates, on the basis of output signals of the pixels 530 selected by the read-out region selecting section 527, event signals corresponding to active pixels in which events have been detected of the selected pixels 530. The events mean an event that the intensity of light changes. The active pixels mean the pixel 530 in which the amount of change in light intensity corresponding to an output signal exceeds or falls below a threshold set in advance. For example, the signal generating section 528 compares output signals from the pixels 530 with a reference signal, and detects, as an active pixel, a pixel 73785

[0199] 20 that outputs an output signal larger or smaller than the reference signal. The signal generating section 528 generates an event signal (event data) corresponding to the active pixel.

[0200] The signal generating section 528 can include, for example, a column selecting circuit configured to arbitrate signals input to the signal generating section 528. Further, the signal generating section 528 can output not only information regarding active pixels in which events have been detected, but also information regarding non-active pixels in which no event has been detected.

[0201] The signal generating section 528 outputs, through an output line 515, address information and timestamp information (for example, (X, Y, T)) regarding the active pixels in which the events have been detected. However, the data that is output from the signal generating section 528 may not only be the address information and the timestamp information, but also information in a frame format (for example, (0, 0, 1, o, -)).

[0202] In the following description reference will mainly be made to sensor devices of the EVS type as described above in order to ease the description and to cover an important application example. However, the principles explained below apply just as well to any event-based vision sensor that is capable to generate events based on the occurrence of temporal intensity changes. In the following it will be described how this basic concept of event-based vision sensors can be extended to color variations.

[0203] Fig. 15A shows an exemplary imaging system 300 comprising an event-based vision sensor (EVS) 310 and a processing unit 320.

[0204] The EVS 310 obtains event-based data EVD of a dynamic scene DYS. It detects as an event the occurrence of a change in intensities of light. As described above, the EVS 310 outputs pixel-level brightness changes instead of standard intensity frames and may have a configuration as explained in the previous paragraphs.

[0205] The EVS 310 may be implemented, for example, in an event camera that may capture the dynamic scene DYS. The dynamic scene DYS may capture a time sequence of information It_2, It-i, It, which will be explained further in combination with Figs. 16A and 16B. In Fig. 15A, a simplified example of a dynamic scene DYS is illustrated, where an object 340 such as a ball performs a certain movement.

[0206] This movement is further indicated in Fig. 15B, which shows the ball at different times t-1 and t. The ball describes a specific trajectory and casts a shadow 360 when it is close enough to a surface. The ball may come into contact with the surface at a time indicated as t-coll. The details for predicting such a collision will be explained in detail later on.

[0207] The processing unit 320 processes the event-based data EVD. It may comprise a central processing unit (CPU) that allows analyzing the event-based data EVD received from the EVS 310 as indicated by an arrow in Fig. 15A. 73785

[0208] 21

[0209] In general, the prediction of trajectories of moving objects from image data involves solving of motion equations considering environmental influences such as interactions with other objects and / or surfaces. In this regard, the dynamics of interactions can be hard to predict. In addition, in the context of augmented reality (AR), the movement of objects such as a hand of a user requires high accuracy to correctly interpret gestures and to predict interactions with other objects (virtual or real). For example, for a user that intends to interact with a virtual controller, a corresponding contact of a finger on the virtual controller needs to be correctly predicted to ensure an optimal user experience. Furthermore, an accurate prediction of an object’s trajectory and corresponding collisions / contacts plays also a crucial role in the field of robotics, where detecting / predicting robot-environment contact needs to be performed with high precision.

[0210] The imaging system 300 of the present application addresses these problems by using the event-based data EVD obtained by the EVS 310.

[0211] More specifically, depending on the processed event-based data EVD, the imaging system 300 evaluates how the shadow 360 changes in the dynamic scene DYS to determine whether the object 340 casting the shadow 360 collides with the surface and / or another object in the dynamic scene DYS. Considering eventbased data EVD of both the object 340 and its shadow 360 allows determining if a collision with the shadow or the surface on which the shadow is cast takes place, and if this the case, when and where that collision might happen. The use of the EVS 310 enables analyzing the spatio-temporal movement of the object 340 and its shadow 360 with high temporal precision.

[0212] It is noted that in the following examples described regarding Figs. 15A to 17, only the collision between an object 340 and a surface is considered for simplicity reasons. However, the described mechanisms may also be applied to colliding objects 340.

[0213] As can be seen in Fig. 15B, the object 340 (e.g., the ball as described above) and its shadow 360 may converge in space-time. The movement of the object 340 and its shadow 360 give rise to a distinct signal in the EVS 310 due to the apparent intensity changes caused by the movement. More precisely, the corresponding changes may be reflected as events and obtained (as the event-based data EVD) by the EVS 310. Processing this event-based data 310 provides the processing unit 310 with information that allows it to predict the moment of contact between the object 340 and the surface. Thus, by monitoring the object 340 and the shadow 360 cast by the object 340 in the dynamic scene DYS via the EVS 310, it may be possible to predict a collision with high accuracy.

[0214] In greater detail, the EVS 310 may detect and track an edge of the object 340 and an edge of the shadow 360 cast by the object 340 in the dynamic scene DYS with high precision.

[0215] In this context, Figs. 16A and 16B show the dynamic scene DYS from the perspective of the EVS 310 implemented, e.g., in the event camera as described above. Specifically, Fig. 16A shows information It_2, It- i, E captured by the EVS 310 at different points in time t-2, t-1, t. The event-based data EVD associated with each one of this information It_2, It-i , E may be used to predict the possibility of a collision and a possible time t-coll and position of the collision of the object 340 with the surface. 73785

[0216] 22

[0217] This is indicated in Fig. 16B, where corresponding edges of the object 340 and the shadow 360 (e.g., the event-based data EVD) are illustrated for information It.i and It. The processing unit 320 may process the event-based data EVD pertaining to each specific information It_i, Itby analyzing features such as the edges of the object 340 and its shadow 360 in the event-based data EVD. This is indicated in a simplified manner in Fig. 16B. Depending on this processed event-based data, the processing unit 320 may predict a motion vector of the object 340 and a motion vector of the shadow 360 cast by the object 340.

[0218] More specifically, processing the event-based data EVD by using both the object and the shadow edges may be performed by a first algorithm of the processing unit 320, which may detect if the object 340 is in collision with the surface.

[0219] As shown in Fig 16B for example, the event-based data EVD due to the object 340 and its shadow 360 follows a specific three-dimensional pattern, which typically resembles a pair of tubular structures that converge in space-time. To detect whether the object 340 is in collision with the surface, different approaches may be used as the first algorithm. For example, one example of such an algorithm may compute the density of points of the converging tubular structures, as a function of time, and detect when the density increases significantly, e.g., above a predefined threshold. Such an event would trigger the onset of a collision.

[0220] Alternatively, another algorithm may use a parametric-based approach, by fitting a spatio-temporal curve to the 3D point clouds of the two structures (updating continuously this fit with new data-points) and apply a threshold condition between the distance of the spatiotemporal curves.

[0221] Furthermore, as another alternative, it is possible to use a third type of algorithm based on a machine learning approach. A machine learning model may be pre-trained on a dataset consisting of event-based data EVD corresponding to collision scenarios as the ones described in the present application, but with a diversity of objects, illumination conditions and velocities, and where the target label would be a flag signaling whether there is a collision or not.

[0222] In this regard, the processing unit 320 may calculate, depending on the processed event-based data, a distance dist between the object 340 and the shadow 360 cast by the object 340 as a function of time to predict a position and the time of collision t-coll with the surface and / or the other object (as mentioned above the other object is not shown in Figs. 15Ato 17).

[0223] For example, the processing unit 320 may infer the separation / distance dist between the object 340 and the surface as a function of time. By extrapolating this inference into the future, e.g., by applying a second algorithm, the processing unit 320 may predict the moment of collision, i.e., the position and the time of collision t-coll with the surface. This predicted collision is shown in Fig. 16B by E-coii indicated through dashed lines. The dashed dotted lines show how the distance dist between the edge of the object 340 and the edge of the shadow 360 decreases until they may converge in space-time. 73785

[0224] 23

[0225] A parametric-algorithm, similar to the second alternative described above, may be applied as the second algorithm before the collision happens, for example, by fitting the point-clouds to spatio-temporal curves. In this case, good extrapolation properties may be achieved, e.g., by applying orthogonal polynomials. By then extrapolating these curves in time, solving for the time at which a certain distance condition between the curves is satisfied, e.g., the point in time where the curves distances are lower than a predefined- threshold, may provide an estimate of the future collision time t-coll of the object 340 with the surface.

[0226] Alternatively, it is possible to use a machine learning approach, where the point clouds are input into a machine learning model that was previously trained with a multitude of data generated with a variety of objects, illumination conditions and velocities, to predict the time of collision t-coll between the two objects 340.

[0227] The imaging system 300 may further comprise a light source 380 indicated in Figs. 15A and 16A, and 17. The light source 380 may provide artificial light. According to one example, passive, natural ambient light may be supplied by the sun as illustrated in Figs. 15A and 16A. For example, when objects 340 are in proximity, they will cast shadows 360 on each other by obstructing the ambient light, which may be used for determining and predicting a possible collision between these objects 340.

[0228] In addition, or alternatively, active, e.g., fixed LED / laser light may be provided for better control of the shadow’s 360 geometry. Therefore, the light source 380 may comprise, e.g., a light emitting diode (LED).

[0229] According to examples, the light source 380 may be installed at the top, while the event camera comprising the EVS 310 may be installed at a side to provide an optimal illuminated overview of the dynamic scene DYS. Different configurations are possible, depending on the specific requirements.

[0230] The processing unit 320 may further apply a software framework for developing and providing machine learning tools to process the event-based data EVD. For example, the above described first / second algorithms may be developed based on machine learning architectures to learn from data and experience in an in principle known manner. This may provide a better detection and interpretation of object movements such as a hand motion or a robotic arm movement. Furthermore, applying machine learning tools may improve processing event-based data EVD in complex situations involving, e.g., a willful movement of an object 340. For example, it may allow to add a reaction time to a collision prediction.

[0231] Fig. 17 illustrates a further exemplary dynamic scene DYS, where an object 340 represented by a finger of a user performs a willful movement and contacts a surface. The left part shows the dynamic scene DYS at time t-1. The middle part at time t. The right part shows the time of collision t-coll, i.e., when the finger contacts the surface and the shadow 360 of the finger merges with the finger. Processing the resulting eventbased data EVD including object- and shadow movement of this dynamic scene DSY allows an accurate prediction of the possible collision time t-coll and position.

[0232] Here, for willful movements of a human the processing unit 320 may also take into account typical reaction times of humans for the collision prediction. In particular, the processing unit 320 may carry out the 73785

[0233] 24 collision prediction for a human hand (or any other body part) by assuming that it is not possible for a human to decelerate or accelerate a hand (or body part) by more than a specific rate. This allows to exclude overly abrupt movements from consideration.

[0234] Fig. 18 shows a process flow of a method for the imaging system 300 described above.

[0235] At S201, the event-based data EVD of the dynamic scene DYS is obtained by the EVS 310. As described above, the EVS 310 detects as an event the occurrence of a change of a difference in intensities of light.

[0236] At S202, the event-based data EVD is processed by the processing unit 320.

[0237] At S203, the imaging system 300 evaluates, depending on the processed event-based data EVD, how the shadow 360 changes in the dynamic scene DYS to determine whether the object 340 casting the shadow 360 collides with the surface and / or another object in the dynamic scene SYZ.

[0238] Moreover, at S201 detecting and tracking, by the EVS 310, the edge of the object 340 and the edge of the shadow 360 cast by the object 340 in the dynamic scene DYS may be carried out.

[0239] Furthermore, at S202 processing, by the processing unit 320, the event-based data EVD by analyzing features in the event-based data EVD may be carried out. The features may comprise an edge of the object 340 and an edge of the shadow 360 cast by the object 340.

[0240] In addition, at S202 deriving, by the processing unit 320, a motion vector of the object 340 and a motion vector of the shadow 360 cast by the object 340 depending on the processed event-based data may be carried out.

[0241] At S204 calculating, by the processing unit 320, the distance dist between the object 340 and the shadow 360 cast by the object 340 as a function of time to predict the position and time of collision t-coll with the surface and / or the other object depending on the processed event-based data may be carried out.

[0242] Further possible implementations of the sensor device 100 are mobile devices 3000 such as cell phones, tablets, smart watches and the like as shown in Fig. 19A or head-mounted displays 4000 as shown in Fig. 19B. Further, the sensor device 10 is useable in augmented and / or virtual reality applications / cameras or in surveillance systems like 360° cameras.

[0243] Note that, the embodiments of the present technology are not limited to the above-mentioned embodiment, and various modifications can be made without departing from the gist of the present technology.

[0244] Further, the effects described herein are only exemplary and not limited, and other effects may be provided.

[0245] Note that, the present technology can also take the following configurations. 73785

[0246] 25

[0247] [1] An imaging system (300) comprising: an event-based vision sensor, EVS (310), configured to obtain event-based data (EVD) of a dynamic scene (DYS), wherein the EVS (310) detects as an event the occurrence of a change in intensities of light; and a processing unit (320) configured to process the event-based data (EVD), wherein, depending on the processed event-based data, the imaging system (300) is configured to evaluate how a shadow (360) changes in the dynamic scene (DYS) to determine whether an object (340) casting the shadow (360) collides with a surface and / or another object in the dynamic scene (DYS).

[0248] [2] The imaging system (300) according to [1], wherein the EVS (310) is further configured to monitor the object (340) and the shadow (360) cast by the object (340) in the dynamic scene (DYS).

[0249] [3] The imaging system (300) according to [1] or [2], wherein the EVS (310) is further configured to detect and track an edge of the object (340) and an edge of the shadow (360) cast by the object (340) in the dynamic scene (DYS).

[0250] [4] The imaging system (300) according to any one of [1] to [3], wherein the processing unit (320) is further configured to: process the event-based data (EVD) by analyzing features in the event-based data (EVD), the features comprising an edge of the object (340) and an edge of the shadow (360) cast by the object (340), and predict, depending on the processed event-based data, a motion vector of the object (340) and a motion vector of the shadow (360) cast by the object (340).

[0251] [5] The imaging system (300) according to any one of [1] to [4], wherein the processing unit (320) is further configured to calculate, depending on the processed event-based data, a distance (dist) between the object (340) and the shadow (360) cast by the object (340) as a function of time to predict a position and time of collision (t-coll) with the surface and / or the other object.

[0252] [6] The imaging system (300) according to any one of [1] to [5], further comprising a light source (380) which provides artificial light.

[0253] [7] The imaging system (300) according to [6], wherein the light source (380) comprises a light emitting diode, LED.

[0254] [8] The imaging system (300) according to any one of [1] to [7], wherein the processing unit (320) is further configured to apply a software framework for developing and providing machine learning tools to process the event-based data (EVD).

[0255] [9] A head mounted display comprising the imaging system according to any one of [1] to [8], 73785

[0256] 26

[0257]

[0010] A method for an imaging system (300) according to any one of [1] to [8] or for a head mounted display according to [9], the method comprising: obtaining (S201), by an event-based vision sensor, EVS (310), event-based data (EVD) of a dynamic scene (DYS), wherein the EVS (310) detects as an event the occurrence of a change of a difference in intensities of light; processing (S202), by a processing unit (320), the event-based data (EVD); and evaluating (S203), by the imaging system (300), how a shadow (360) changes in the dynamic scene (DYS) to determine whether an object (340) casting the shadow (360) collides with a surface and / or another object in the dynamic scene (DYS) depending on the processed event-based data.

[0258]

[0011] The method according to

[0010] , wherein obtaining the event-based data (EVD) of the dynamic scene (DYS) further comprises detecting and tracking (S201), by the EVS (310), an edge of the object (340) and an edge of the shadow (360) cast by the object (340) in the dynamic scene (DYS).

[0259]

[0012] The method according to

[0010] or

[0011] , wherein processing the event-based data (EVD) further comprises: processing (S202), by the processing unit (320), the event-based data (EVD) by analyzing features in the event-based data (EVD), the features comprising an edge of the object (340) and an edge of the shadow (360) cast by the object (340), and deriving, by the processing unit, a motion vector of the object (340) and a motion vector of the shadow (360) cast by the object (340) depending on the processed event-based data.

[0260]

[0013] The method according to any one of

[0010] to

[0012] , wherein evaluating how a shadow (360) changes in the dynamic scene (DYS) further comprises: calculating (S203), by the processing unit (320), a distance (dist) between the object (340) and the shadow (360) cast by the object (340) as a function of time to predict a position and time of collision (t-coll) with the surface and / or the other object depending on the processed event-based data.

Claims

7378527CLAIMS1. An imaging system comprising: an event-based vision sensor, EVS, configured to obtain event-based data of a dynamic scene, wherein the EVS detects as an event the occurrence of a change in intensities of light; and a processing unit configured to process the event-based data, wherein, depending on the processed event-based data, the imaging system is configured to evaluate how a shadow changes in the dynamic scene to determine whether an object casting the shadow collides with a surface and / or another object in the dynamic scene.

2. The imaging system according to claim 1, wherein the EVS is further configured to monitor the object and the shadow cast by the object in the dynamic scene.

3. The imaging system according to claim 1 , wherein the EV S is further configured to detect and track an edge of the object and an edge of the shadow cast by the object in the dynamic scene.

4. The imaging system according to claim 1, wherein the processing unit is further configured to: process the event-based data by analyzing features in the event-based data, the features comprising an edge of the object and an edge of the shadow cast by the object, and predict, depending on the processed event-based data, a motion vector of the object and a motion vector of the shadow cast by the object.

5. The imaging system according to claim 1, wherein the processing unit is further configured to calculate, depending on the processed event-based data, a distance between the object and the shadow cast by the object as a function of time to predict a position and time of collision with the surface and / or the other object.

6. The imaging system according to claim 1, further comprising a light source which provides artificial light.

7. The imaging system according to claim 6, wherein the light source comprises a light emitting diode, LED.

8. The imaging system according to claim 1, wherein the processing unit is further configured to apply a software framework for developing and providing machine learning tools to process the event-based data.

9. A head mounted display comprising the imaging system according to claim 1.

10. A method for an imaging system, the method comprising: obtaining, by an event-based vision sensor, EVS, event-based data of a dynamic scene, wherein the EVS detects as an event the occurrence of a change in intensities of light; processing, by a processing unit, the event-based data; and7378528 evaluating, by the imaging system, how a shadow changes in the dynamic scene to determine whether an object casting the shadow collides with a surface and / or another object in the dynamic scene depending on the processed event-based data.

11. The method according to claim 10, wherein obtaining the event-based data of the dynamic scene further comprises detecting and tracking, by the EVS, an edge of the object and an edge of the shadow cast by the object in the dynamic scene.

12. The method according to claim 10, wherein processing the event-based data further comprises: processing, by the processing unit, the event-based data by analyzing features in the event-based data, the features comprising an edge of the object and an edge of the shadow cast by the object, and deriving, by the processing unit, a motion vector of the object and a motion vector of the shadow cast by the object depending on the processed event-based data.

13. The method according to claim 10, wherein evaluating how a shadow changes in the dynamic scene further comprises: calculating, by the processing unit, a distance between the object and the shadow cast by the object as a function of time to predict a position and time of collision with the surface and / or the other object depending on the processed event-based data.