Method of tracking objects in a scene
By combining a two-process approach with an event-based sensor, object tracking with high temporal resolution and low computational complexity is achieved. This solves the robustness and stability problems of existing object tracking algorithms under sudden changes and occlusion conditions, and provides smooth and accurate trajectory estimation.
Patent Information
- Application Number
- CN201980078147.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-13
- Filing Date
- 2019-12-13
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2039-12-13
AI Technical Summary
Existing frame-based object tracking algorithms struggle to achieve smooth and accurate trajectory estimation with high temporal resolution and low computational complexity when faced with sudden changes, occlusion, and deocclusion. Furthermore, event-based algorithms lack robustness and stability.
A two-process approach is adopted. The first process detects objects in the scene at a low temporal resolution, while the second process tracks and updates the object's position at a high temporal resolution. By leveraging the high temporal resolution and sparse sampling of event-based sensors, combined with global information and a large field of view, small errors are corrected and assumptions about the object's shape are avoided.
It achieves high temporal resolution object tracking, improves robustness to drift and occlusion/de-occlusion, maintains low overall computational complexity, and provides smooth and accurate trajectory estimation.
Smart Images

Figure CN113168520B_ABST
Abstract
Description
BACKGROUND
[0001] The present invention relates to the field of object tracking, where objects are localized and their trajectories determined over time. In particular, the present invention relates to a method for tracking objects using event-based sensors.
[0002] Tracking objects in a scene is an important topic in computer vision in the sense of estimating the object's trajectory as it moves through the scene. Object tracking plays a fundamental role in a wide range of applications such as visual surveillance, video compression, driver assistance, odometry, simultaneous localization and mapping (SLAM), traffic monitoring, automated valet parking systems, etc.
[0003] For these applications, it is essential to have the ability to deal with sudden changes such as changes in object orientation, appearance of new objects in the scene, occlusions and de-occlusions. For example, in driver assistance, the system must react very quickly to a sudden change in car orientation or to a pedestrian entering the field of view.
[0004] In these applications, the tracked objects can suddenly change orientation, new objects can appear, occlusions / de-occlusions can occur and the tracking system should react very quickly to these situations.
[0005] In conventional algorithms, tracking is based on image sequences acquired at a fixed frame rate. Therefore, when a sudden change occurs, the tracking accuracy is limited by the inter-frame period, which can be much larger than the dynamics of the change. Moreover, the computational cost of running such a detection / tracking system at high frame rates would be very high and sometimes even unfeasible due to the latency of the system. For example, if a detector needs 100 ms to detect a car in a picture, it cannot run at more than 10 frames per second, while the whole system needs a much faster reaction time. Moreover, running the detector on an embedded platform can reduce the detection rate to about 1.5 frames per second, which can be too slow for the above-mentioned applications.
[0006] Alternatively, event-based light sensors can also be used to generate asynchronous signals, instead of regular frame-based sensors. Event-based light sensors deliver compressed digital data in the form of events. A presentation of such sensors can be found in "Activity-Driven, Event-Based Vision Sensors", T. Delbruck et al., Proceedings of 2010 IEEE International Symposium on Circuits and Systems, pp. 2426-2429. Compared to regular cameras, event-based vision sensors have the advantage of eliminating redundancy, reducing latency, and increasing dynamic range.
[0007] For each pixel address, the output of such event-based light sensors can exist in a series of asynchronous events representing a change in the light parameter (e.g. brightness, light intensity, reflectance) of the scene as it happens. Each pixel of the sensor is independent and detects a change in intensity greater than a threshold since the last event emission (e.g. 15% contrast in intensity log). When the intensity change exceeds a set threshold, the pixel can generate an ON or OFF event depending on whether the intensity increased or decreased. Such events can be referred to as "change detection events" or CD events hereinafter.
[0008] Generally, an asynchronous sensor is placed in the image plane of the optics for the acquisition. The asynchronous sensor comprises an array of sensing elements, like light-sensitive elements, organized in a matrix of pixels. Each sensing element corresponding to a pixel p generates a succession of events e(p, t) at times t depending on the variations of light in the scene.
[0009] For each sensing element, the asynchronous sensor generates a sequence of event-based signals from the variations of light received by the sensing element from the scene present in the field of view of the sensor.
[0010] The asynchronous sensor performs an acquisition to output a signal that, for each pixel, can take the form of successive instants t k (k = 0, 1, 2,... ) reaching an activation threshold Q. Each time this brightness increases by an amount equal to the activation threshold Q from its amount at time t k a new instant t k+1 is identified and a spike is emitted at this instant t k+1 Symmetrically, each time the brightness observed by the sensing element decreases by an amount Q from its amount at time t k a new instant t k+1 is identified and a spike is emitted at this instant tk+1 A spike is emitted. The signal sequence of the sensing element contains a series of spikes at times t k The spikes are positioned over time, depending on the light profile of the sensing element. The output of the sensor 10 is then in the form of an address event representation (AER). In addition, the signal sequence can contain a luminance attribute corresponding to the change in incident light.
[0011] The activation threshold Q can be fixed or can be adapted as a function of the luminance. For example, when the threshold is exceeded, the threshold can be compared to a change in the logarithm of the luminance used to generate the event.
[0012] By way of example, the sensor can be a dynamic vision sensor (DVS) of the type described in P. Lichtsteiner et al., “An 128x128 120dB 15μs Latency Asynchronous Temporal Contrast Vision Sensor”, IEEE Journal of Solid-State Circuits, vol. 43, no. 2, February 2008, pages 566-576, or in US patent application US 2008 / 0135731 Al. The dynamics of the retina (minimum duration between action potentials) can be processed with this type of DVS. The dynamic behavior exceeds that of a conventional video camera with a real sampling frequency. When the DVS is used as an event-based sensor 10, the data relating to an event originating from a sensing element contain the address of the sensing element, the time of occurrence of the event and a luminance attribute corresponding to the polarity of the event, for example +1 if the luminance increases and -1 if the luminance decreases.
[0013] Figure 1 represents, on an arbitrary scale, an example of a light profile received by such an asynchronous sensor and of a signal generated by this sensor in response to the light profile.
[0014] The light profile P seen by the sensing element contains an intensity pattern, with alternating rising and falling edges in the simplified example shown for the sake of explanation. At t0, an event is generated by the sensing element when the luminance increases by an amount equal to the activation threshold Q. The event has a time of occurrence t0 and a luminance attribute, i.e. a polarity in the case of a DVS (level +1 at t0 in Figure 1). As the luminance increases with the rising edge, a subsequent event with the same positive polarity is produced each time the luminance increases by the activation threshold. These events form a burst, denoted B0 in Figure 1. When the falling edge begins, the luminance decreases and another burst Bl with the opposite polarity, i.e. negative (level -1), is produced from time tl.
[0015] Some event-based light sensors can associate detected events with measurements of light intensity (e.g., grayscale), such as the Asynchronous Time-Based Image Sensor (ATIS) described in the article “A QVGA 143 dB Dynamic Range Frame-Free PWM Image Sensor With Lossless Pixel-Level Video Compression and Time-Domain CDS,” C. Posch et al., IEEE Journal of Solid-State Circuits, vol. 46, no. 1, January 2011, pp. 259-275. Such events can be referred to as “exposure measurement events” or EM events hereinafter.
[0016] Since an asynchronous sensor does not sample on a clock like a regular camera, it can consider the ordering of events with very high temporal precision (e.g., on the order of 1 microsecond temporal precision). If an image sequence is reconstructed using such a sensor, image frame rates of several kilohertz can be achieved, while regular cameras have frame rates of only a few tens of hertz.
[0017] Some event-based algorithms have been developed to track objects in a scene. However, the robustness and stability of these algorithms are far weaker than that of frame-based algorithms. Specifically, these algorithms are very sensitive to outliers and affected by drift. Furthermore, most of these algorithms (e.g., “Feature detection and tracking with the dynamic and active-pixel vision sensor,” D. Tedaldi et al., 2016 Second International Conference on Event-based Control, Communication, and Signal Processing) rely on assumptions about the shape of an object, which can greatly reduce the accuracy of object detection and tracking. In addition, these algorithms show poor results in the case of occlusions.
[0018] Accordingly, there is a need for object tracking methods that can be performed with high temporal resolution without suffering from the robustness and stability issues described above. SUMMARY
[0019] The invention relates to a method of tracking an object in a scene observed by an event-based sensor producing events asynchronously. The event-based sensor has a matrix of sensing elements. A respective event is produced by each sensing element in the matrix from a change in light incident on the sensing element. The method comprises:
[0020] detecting, by a first process, the object in the scene from the asynchronous events produced by the event-based sensor; and
[0021] determining, by a second process, a trace associated with the object detected from the asynchronous events produced by the event-based sensor;
[0022] wherein a time resolution of the first process is lower than a time resolution of the second process.
[0023] By "trace" is meant a time series of positions of an object in a scene over time, i.e. a time series corresponding to a trajectory of the object over time. For example, a trace can comprise a series of dated time coordinates (x i ,y i ,t i ): {(x1,y1,t1),..., (x k ,y k ,t k )}, where each (x i ,y i ) pair corresponds to the position of a point representing the object at time t i . When the object is represented by its shape (e.g. in case segmentation masks are used, as detailed below), this point can be the center of gravity, the center of gravity of a figure comprising the object, a corner of this figure, etc. In the following, "tracking an object" corresponds to determining a trajectory of an object in a scene.
[0024] By "object" is meant not limited to a feature point or a simple local pattern, like a corner or a line. It can be a more complex pattern or a substantial thing, which can be recognized in part or in whole, e.g. a car, a human body or a body part, an animal, a road sign, a bicycle, a red flare, a ball, etc. By "time resolution" is meant the time elapsed between two iterations of a process, which can be fixed or variable.
[0025] For example, the event-based sensor can produce EM events, CD events, or both EM events and CD events. In other words, the detection and / or tracking can be performed based on EM events and / or CD events.
[0026] The first process allows to accurately determine the objects present in the scene, for example by applying to the frame a detection algorithm based on the events received by the event-based sensor. However, such an algorithm is computationally costly and can not be suitable for very high temporal resolution. The second process compensates for this weakness by tracking the objects detected during the first process with a higher temporal resolution and updating the position of the objects when new events are received from the sensor.
[0027] This two-process approach can on the one hand take advantage of the high temporal resolution and sparse sampling of event-based cameras and on the other hand of the global information and large field of view accumulated over a larger time period. This solution provides robustness to drifts and occlusions / de-occlusions compared to event-based algorithms (e.g. the algorithm provided in “An asynchronous neuromorphic event-driven visual part-based shape tracking”, D.R. Valeiras et al., IEEE transactions on neural networks and learning systems, vol. 26, no. 12, December 2015, pages 3045-3059), specifically because the first process can be used to correct small errors accumulated by the second process. Moreover, no assumption on the shape of the object to track is needed.
[0028] This solution enables high temporal resolution compared to frame-based algorithms, thus providing smooth and accurate trajectory estimation while keeping a low overall computational complexity.
[0029] It should be noted that the two processes have complementary properties and there is no strict hierarchy between the two.
[0030] In one embodiment, the first process can comprise determining a class of the detected object. For example, the class can be determined in a class set, the set being defined according to the application. For example, in the context of driving assistance, the class set can be {car, pedestrian, road sign, bicycle, red flare}. The determination of the class can be performed by machine learning, for example a supervised learning algorithm.
[0031] In one embodiment, a new iteration of the first process is run according to a predefined time sequence.
[0032] This makes it possible to periodically adjust the results obtained by the second process, which can drift. The interval between two instants of the predefined time sequence can be fixed or variable. The order of magnitude of the time interval can depend on the application. For example, for a monitoring / surveillance system installed on a non-moving stand, the first process can run at a very low time resolution. The time between two iterations can be about 1 or 2 seconds. For a driving assistance system installed in the cockpit of a vehicle, the time between two iterations of the first process can be shorter, for example between 0.1 and 0.5 seconds.
[0033] It should be noted that, in addition to iterations run according to the predefined time sequence, new iterations of the first process can also be run under certain conditions, for example when the number of events received in a zone of the scene is below a threshold, as detailed below.
[0034] In one embodiment, the first process can comprise determining a zone of interest in the matrix of sensing elements, the zone of interest containing the detected object; and the second process can comprise determining the track based on asynchronous events generated by sensing elements in a zone containing the zone of interest.
[0035] The zone of interest can be a zone whose shape corresponds to the shape of the object (for example, a segmentation mask, as detailed below), or a zone larger than the object and including the object (for example, a bounding box, which can be a rectangle or another predefined figure). The bounding box allows simpler and less computationally costly data processing, but does not provide very precise information about the exact positioning and shape of the object. The segmentation mask can be used when precise information is needed.
[0036] When the object is moving, events are generated inside the object, at and around its edges. Using events generated inside and around the zone of interest allows better tracking of the moving object. For example, this zone can correspond to a predefined neighborhood of the zone of interest (i.e. a zone containing all the neighboring pixels of all the pixels of the zone of interest). The neighborhood can be determined, for example, based on a given pixel connectivity (for example, 8-pixel connectivity).
[0037] In such an embodiment, the second process can further comprise updating the zone of interest based on asynchronous events generated in a zone containing the zone of interest.
[0038] In addition, the second process can further comprise running a new iteration of the first process based on a parameter of the zone of interest, the parameter being one of a size of the zone of interest and a shape of the zone of interest.
[0039] Indeed, a parameter can indicate a strong change in the scene, and the second process can accumulate errors on said strong change. Running a new iteration of the first process allows to correct these errors and accurately determine the change in the scene.
[0040] Alternatively or additionally, the second process can further comprise running a new iteration of the first process if the number of events generated in the zone of interest during a time interval exceeds a threshold.
[0041] Similarly, a large number of events can indicate a strong change in the scene.
[0042] Furthermore, the method of the application can further comprise:
[0043] determining a relevance between the detected objects and the determined tracks; and
[0044] updating the set of detected objects and / or the set of determined tracks based on the relevance.
[0045] Thus, the results of the first process can be correlated with the results of the second process and the object localization can be more accurately determined. Indeed, the trajectories of the objects can be validated or corrected based on the determined relevance. This improves the robustness and stability of the tracking and prevents drifts.
[0046] The method can further comprise:
[0047] if a given object among the detected objects is associated with a given track among the determined tracks, updating the given track based on the position of the given object.
[0048] Additionally or alternatively, the method can further comprise:
[0049] if a given object among the detected objects is not associated with any track among the determined tracks, initializing a new track corresponding to the given object;
[0050] if a given track among the determined tracks is not associated with any object among the detected objects, deleting the given track; and
[0051] if a first object and a second object among the detected objects are both associated with a first track among the determined tracks, initializing a second track from the first track, each of the first and second tracks corresponding to a respective object among the first and second objects;
[0052] if a given object among the detected objects is associated with a first track and a second track among the determined tracks, merging the first and second tracks.
[0053] In one embodiment, the second process can further comprise:
[0054] For each detected object, determining a respective set of zones in the matrix of sensing elements, each zone of the set of zones being associated with a respective time;
[0055] receiving events produced by the event-based sensor during a predefined time interval;
[0056] determining groups of events, each group comprising at least one event of the received events;
[0057] For each group of events, determining a respective region in the matrix of sensing elements, the region containing all events of the group of events;
[0058] determining a correlation between the determined regions and the determined zones; and
[0059] updating the set of detected objects and / or the set of determined tracks based on the correlation.
[0060] Such a procedure corrects errors and improves tracking accuracy during the second process. The determination of regions in the matrix of sensing elements can be performed, for example, based on a clustering algorithm.
[0061] In one embodiment, the first process can further comprise determining a segmentation mask of the detected object. The second process can also comprise updating the segmentation mask based on asynchronous events produced in a neighborhood of the segmentation mask.
[0062] “Segmentation” means any process that divides an image into a plurality of sets of pixels, each pixel of the sets being assigned a label, such that pixels having the same label share certain properties. For example, the same label can be assigned to all pixels corresponding to a car, a bicycle, a pedestrian, etc.
[0063] “Segmentation mask” means a set of pixels having the same label and corresponding to the same object. Thus, the edges of the segmentation mask can correspond to the edges of the object, and the interior of the mask is the pixels belonging to the object.
[0064] Using a segmentation mask allows to more accurately determine the position of an object in a scene. This is possible because running the first process and the second process does not require a predefined model of the shape of the object.
[0065] Furthermore, the method can comprise:
[0066] extracting a first descriptive parameter associated with events received from the event-based sensor;
[0067] extracting a second descriptive parameter associated with events corresponding to the segmentation mask; and
[0068] updating the segmentation mask by comparing a distance between one of the first descriptive parameters and one of the second descriptive parameters with a predefined value;
[0069] wherein the first and second descriptive parameters associated with an event comprise at least one of a polarity of the event and a velocity of the event.
[0070] Thus, the segmentation mask can be accurately updated at each iteration of the first and / or second process.
[0071] Another aspect of the invention relates to a signal processing unit comprising:
[0072] an interface for connecting to an event-based sensor having a matrix of sensing elements that produce events asynchronously from light received from a scene, wherein a respective event is produced by each sensing element in the matrix from a change in light incident on the sensing element; and
[0073] a processor configured to:
[0074] detect, by a first process, objects in the scene from the asynchronous events produced by the event-based sensor; and
[0075] determine, by a second process, a track associated with an object detected from the asynchronous events produced by the event-based sensor;
[0076] wherein a temporal resolution of the first process is lower than a temporal resolution of the second process.
[0077] Yet another object of the invention relates to a computer program product comprising stored instructions for execution in a processor associated with an event-based sensor having a matrix of sensing elements that produce events asynchronously from light received from a scene, wherein a respective event is produced by each sensing element in the matrix from a change in light incident on the sensing element, wherein execution of the instructions causes the processor to:
[0078] detect, by a first process, objects in the scene from the asynchronous events produced by the event-based sensor; and
[0079] determine, by a second process, a track associated with an object detected from the asynchronous events produced by the event-based sensor;
[0080] wherein a temporal resolution of the first process is lower than a temporal resolution of the second process.
[0081] Other features and advantages of the methods and devices disclosed herein will become apparent in the following description in view of the attached drawings. BRIEF DESCRIPTION OF DRAWINGS
[0082] The application is illustrated in the accompanying drawings, which are included by way of example and not limitation, and in which like reference numerals refer to similar elements and in which:
[0083] - figure 1 represents an example of light distribution received by an asynchronous sensor and of a signal generated by the asynchronous sensor in response to the light distribution;
[0084] - Figure 2 is a flowchart describing a possible embodiment of the application;
[0085] - Figure 3a and 3b represent an example of detection and tracking of objects in a possible embodiment of the application;
[0086] - Figure 4 is a flowchart describing an update of a set of tracks and of a set of objects in a possible embodiment of the application;
[0087] - Figure 5 represents an update of a set of tracks and of a set of objects in a possible embodiment of the application;
[0088] - Figure 6 represents a segmentation mask associated with an object in a scene in a possible embodiment of the application;
[0089] - Figure 7 is a possible embodiment of a device implementing the application. DETAILED DESCRIPTION
[0090] In the context of the application, it is assumed that an event-based asynchronous vision sensor is placed facing a scene and receives a light flow of the scene through acquisition optics comprising one or more lenses. For example, the sensor can be placed in the image plane of the acquisition optics. The sensor comprises a set of sensing elements organized in a matrix of pixels. Each pixel corresponding to a sensing element generates successive events as a function of variations of light in the scene. A sequence of events e(p, t) received asynchronously from the light-sensitive elements p is processed to detect and track objects in the scene.
[0091] Figure 2 is a flowchart describing a possible embodiment of the application;
[0092] According to the application, the method of tracking objects can comprise a first iteration process 201 and a second iteration process 202. The time resolution of the first process 201 is lower than the time resolution of the second process 202. Thus, the first process 201 can also be referred to as the "slow process" and the second process 202 can also be referred to as the "fast process".
[0093] In one embodiment, during the iteration of the first process 201 at time t n The set of objects in the scene is detected based on information received for the whole field of view or a substantial part thereof (e.g. a part representing more than 70% of the field of view). The information can be asynchronous events produced by the event-based sensor during a certain time interval. For example, the information can correspond to asynchronous events produced by the event-based sensor between the previous iteration of the slow process 201 (at time t n-1 and the current iteration of the slow process 201 (at time t n and t n to At (At being a predetermined time interval).
[0094] During the iteration of the slow process 201, some objects in the scene are detected. For example, a frame can be generated from the received information (i.e. the received asynchronous events) and a state-of-the-art object detection algorithm can be run on this frame. When the sensor is configured to provide grey level information associated with the asynchronous events, for example in the case of an ATIS type of sensor, the frame can be built from the grey level measurements by attributing to each pixel of the image the grey level of the last (i.e. the most recent) event received at this pixel.
[0095] When the grey level measurements are not available, for example in the case of a DVS type of sensor, the frame can be obtained by accumulating the events received during a certain time interval. Alternatively, for each pixel of the image, the time of the last event (i.e. the most recent event) received at this pixel can be stored and the frame can be generated by using these stored times. Another alternative consists in reconstructing the grey levels from the events as proposed for example in "Real-time intensity-image reconstruction for event cameras using manifold regularization", Reinbacher et al., 2016.
[0096] Other detection algorithms that do not necessarily generate frames can also be used. For example, if the number of events is greater than a predefined threshold, the received events can be counted in a given region of the matrix of sensing elements and an event-based (or "event-driven") classifier can be applied. An alternative is to run such a classifier on all possible locations and scales of the field of view.
[0097] It should be noted that the time between two successive iterations of the slow process 201 is not necessarily fixed. For example, a new iteration of the slow process 201 can be run when the number of events produced in the entire field of view or in one or more regions of the field of view during a certain time interval exceeds a predefined threshold. Additionally or alternatively, the iterations of the slow process 201 can be performed according to a predefined time sequence (using constant or variable time steps).
[0098] The slow process 201 outputs a set of detected objects 203. This set of objects 203 can be used to update tracks or to initialize new tracks, as explained below. Advantageously, the slow process is configured to determine, for each detected object, a relative class of the object (e.g. "car", "pedestrian", "road sign", etc.). This can be performed, for example, by using a machine learning based algorithm.
[0099] During the fast process 202, the tracks associated with respective objects are updated based on asynchronous events produced by the event-based sensor. In one embodiment, such an update is performed based on asynchronous events produced in a small region of the matrix of sensing elements (e.g. in a zone located around the object and including the object). Of course, it can be several zones of interest (e.g. each zone of interest being associated with an object to be tracked). For example, such a zone can be defined by considering a neighborhood of the object edge (e.g. a 4-, 8- or 16-pixel connectivity neighborhood). Thus, the fast process 202 is performed based on local information around the object detected during the slow process 201 or on the position of the object given by the track at the previous iteration of the fast process 202 (which is performed based on global information of the entire field of view or a large portion thereof compared to the slow process 201).
[0100] New tracks can also be initialized during the fast process 202. For example, when a tracked object is split into two or more objects, a new track can be generated for each of the split objects, respectively. New tracks can also be initialized if a fast detector is run during the fast process 202.
[0101] A new iteration of the fast process can be run at each new event generated in a region of the matrix of sensing elements or when the number of events generated in a region of the matrix of sensing elements exceeds a predefined threshold. The time between two successive iterations of the fast process 202 is therefore generally variable. Since the information to be processed is local and only contains "relevant" data (no redundancy and related to changes in the scene), the fast process 202 can be run at a very high temporal resolution.
[0102] Moreover, a new iteration of the slow process 201 can be run if the number of events generated in one or several regions of interest exceeds another predefined threshold. Indeed, a large number of events can indicate that strong changes occur in the scene and running a new iteration of the slow process can provide more reliable information about these changes.
[0103] The region of the matrix of sensing elements can be updated during the iterations of the fast process. Indeed, when the object to be tracked moves in the scene, its shape can change over time (for example, a turning car can be seen from the back at a certain time and from the side at a later time). Since events are received on the edges of the object, the region of the matrix of sensing elements can be updated based on the received events (by suppressing / adding pixels corresponding to the edges). A new iteration of the slow process 201 can be run based on the parameters of the region of the matrix of sensing elements (for example, the shape or the size). For example, if the size of the region of the matrix of sensing elements significantly increases / decreases between two iterations of the fast process 202 (for example, the difference in size at two different times exceeds a fixed percentage) or if the shape of the region of the matrix of sensing elements significantly changes between two iterations of the fast process 202, this can indicate that the region of the matrix of sensing elements now includes two different objects or another object in addition to the object associated with the track. Running a new iteration of the first process allows to accurately determine the correct scene.
[0104] The slow process 202 outputs a set of tracks 204. This set of tracks 204 can be obtained by any event-based tracking algorithm in the state of the art. As mentioned above, pure event-based tracking algorithms can be very sensitive to outliers and can be affected by drifts. In one embodiment, a determination of the relevance 205 between the tracks determined by the fast process 202 and the objects detected by the slow process 201 can be made to avoid these drawbacks and improve the robustness and stability of the tracking. Based on these relevancies, the set of tracks 204 and / or the set of detected objects 203 can be updated 206.
[0105] Figure 3a and 3b represents an example of detection and tracking of objects in a possible embodiment of the application.
[0106] Figure 3arepresent objects (e.g., cars) detected during iterations of the slow process 201. In Figure 3a In embodiments of the application, an object is represented by a rectangular bounding box 301 comprising said object. In this case, a bounding box can be represented by a tuple (x,y,w,h), where (x,y) are the coordinates of the top-left corner of the rectangle, w is the width of the rectangle and h is the height of the rectangle. The "position" of a bounding box can correspond to the position (x,y) of the top-left corner (or any other point of the box, such as another corner, the barycenter, etc.). Of course, other representations (e.g., shapes of objects or another predefined geometric object other than a rectangle) are possible. More generally, such a representation can be referred to as a "region of interest". A region 302 comprising a region of interest can be determined for running the fast process 202: tracking can then be performed based on events occurring in this region 302.
[0107] Figure 3b represent a detection and tracking sequence of an object obtained at times tl, t2, t3, t4 in a possible embodiment, with tl < t2 < t3 < t4. The solid boxes 303a, 303b correspond to regions of interest output by the slow process at respective times tl and t4, and the dashed boxes 304a, 304b correspond to regions of interest output by the fast process at respective times t2 and t3. As represented in this figure, a region of interest can be updated during the fast process.
[0108] Figure 4 is a flowchart describing the updating of the set of tracks and the set of objects in a possible embodiment of the application.
[0109] As represented in Figure 4 The determination 205 of the association can be based on a predefined association function 401 : f(t, o) of an object o and a track t, as represented in
[0110] In case an object is represented by a respective region of interest, the association function 401 can be defined based on the area of the region of interest. For example, such a function 401 can be defined as:
[0111]
[0112] where area(r) is the area of a figure r (in case of a rectangle, area(r) = width x length), and tn o is the intersection of t and o, as represented in the left part of Figure 5 In this case, the association between a track t and an object o can be performed by maximizing the function 401 f(t, o).
[0113] Referring again to Figure 4 , the set of tracks and / or the set of detected objects can then be updated as follows. For each object o in the set of detected objects, determine one or more tracks t in the set of determined tracks for which the value of f(t, o) is maximal:
[0114] - if this value is not zero, associate the object o with the determined track t and update 407 the size and / or the position of t based on the size and / or the position of o and / or t (e.g. the position of t can be equal to the position of o or equal to a combination of the positions of t and o - e.g. the center or the face center);
[0115] - if there are two tracks t1 and t2 such that for a given object o,
[0116]
[0117] and f(o, t1) - f(o, t2) < ε, where ε is a predefined parameter (preferably chosen small enough), then merge 405 t1 and t2. The size and / or the position of the merged track can be determined based on the size and / or the position of o;
[0118] - if there are two objects o1 and o2 in the set of detected objects such that
[0119]
[0120] then split 406 t into two tracks t1 and t2, setting the position and / or the size of said two tracks respectively for the position and / or the size of o1 and o2; m
[0121] - if a track t is not associated with any object o, then delete 404 said track t; and
[0122] - if an object o is not associated with any track t, then create 403 a new track of size and / or position equal to the size and / or the position of o.
[0123] Figure 5 An example of such an update of the set of tracks is depicted in the right part of Fig. 5. Box 501 corresponds to the case where an object is associated with a single track (the size / position of the box is a combination of the size / position of o and t). Boxes 502 and 503 correspond to the case where a track is split into two tracks because it is associated with one object. Box 504 corresponds to the case where a new track is created because a new object is detected. Track 505 is deleted because it is not associated with any object.
[0124] Of course, other relevance functions can be used. For example, f(t, o) can correspond to a mathematical distance (e.g., Euclidean distance, Manhattan distance, etc.) between two predefined corners of t and o (e.g., between the top-left corner of t and the top-left corner of o). In this case, the association between the track t and the object o can be performed by maximizing the function f(t, o).
[0125] More complex distance metrics can be used to perform the association between objects and tracks, for example by computing descriptors from the events falling into the bounding box during a given time period (as defined for example in “HATS: Histograms of Averaged Time Surfaces for Robust Event-based Object Classification”, Sironi et al., CVPR 2018) and comparing the norm of the difference of the descriptors.
[0126] In one embodiment, the fast pass can include an update of the tracks (independent of the detection results of the slow pass). For example, events can be accumulated during the fast pass over a predefined time period T (e.g., 5 to 10 milliseconds). The event-based tracking pass can perform several iterations during T. The AE is instructed to include the set of events accumulated during T. Advantageously, the AE includes all events received in the entire field of view or a substantial part thereof during T. Alternatively, the AE can include all events received in one or more small parts of the field of view during T.
[0127] The events of the AE can then be clustered into at least one group by using a clustering algorithm (e.g., the Medoid shift clustering algorithm, but other algorithms can be used). Such clustering is typically done by grouping events together if they are “close” to each other (according to a mathematical distance) in time and space, and far from the events in other groups. A cluster can be represented by a minimal rectangle containing all the events of the cluster (other representations can be used as well). For simplicity, in the following, “cluster” refers to both the group of events and its representation (e.g., a rectangle).
[0128] For each trace t in the set of traces and each cluster c in the set of clusters, f(t, c) can be computed, f being a relevance function as defined above. For a given trace t, the cluster c corresponding to the maximum (or minimum, depending on the chosen relevance function) f(t, c) can be associated with t, and the position and size of t can be updated according to the position and size of c. In an embodiment, clusters c that are not associated with any trace t can be deleted (this allows to suppress outliers). In an embodiment, a cluster c can be decided to be suppressed only if the number of events in c is below a predetermined threshold.
[0129] In an embodiment, segmentation can be performed to label all pixels of the region of interest and to distinguish pixels belonging to the tracked object from other pixels. For example, all pixels of the object can be associated with label 1 and all other pixels can be associated with label 0. Such segmentation gives the position accuracy of the object within the region of interest. The set of labeled pixels corresponding to the object (e.g. all pixels associated with label 1) is referred to as a “segmentation mask”. After the object is detected during the slow step, segmentation can be performed by using a further segmentation algorithm, or simultaneously by using a detection and segmentation algorithm on the frames generated by the event-based sensor. The traces associated with the object can then comprise a time series of segmentation masks, where each mask can correspond to a position (e.g. the barycenter of the mask).
[0130] The segmentation mask can be generated from the detected region of interest (e.g. the detected bounding box). For example, the segmentation mask can be determined by running a convolutional neural network (CNN) algorithm on the frames generated by the sensor. Alternatively, the mask boundaries can be determined using the set of events received in the most recent time interval. For example, this can be done by computing the convex hull of the set of pixels for which an event was received in the most recent time interval. Figure 6 An example of a region of interest 601 and a segmentation mask 602 obtained in a possible embodiment of the invention.
[0131] The segmentation mask can be updated during the fast loop. In an embodiment, this update is performed event by event. For example, if an event is close (e.g. in a predefined neighborhood of a mask pixel), it can be assumed that this event was generated by a moving object. Therefore, if this event is located outside the mask, the corresponding pixel can be added to the mask, and if the event is located inside the mask, the corresponding pixel can be removed from the mask.
[0132] Additionally or alternatively, the segmentation mask can be updated based on a mathematical distance between one or more descriptors associated with the events and one or more descriptors associated with the mask. More specifically, the events can be described by descriptors which are vectors containing information on polarity and / or velocity (velocity can be computed for example by using an event-based optical flow algorithm). By taking the mean of the descriptors of all events assigned to the mask, a similar descriptor can be associated with the mask. Then, if the mathematical distance between the event descriptor and the mask descriptor is lower than a predefined threshold, the event can be added to the mask and if this distance is higher than this threshold, the event can be removed from the mask.
[0133] In another embodiment, updating the mask (during the fast progression) is performed by using event groups and probabilistic inference. The idea is to use the prior knowledge of the scene (result of the detection / segmentation algorithm of the slow progression) and the new information (a group of incoming events) to compute the likelihood of a scene change. The segmentation mask can then be updated based on one or more scenes associated with one or more maximum likelihoods.
[0134] For example, the prior knowledge of the scene can be represented by a probability matrix (knowledge matrix) which is a grid of values representing the probability of a pixel to belong to an object, for example a car (for example, there is a car inside the mask (100% probability of car) and no car outside the mask (0% probability of car)). This matrix can be obtained from the mask of the active trace by using state of the art (for example, Conditional Random Fields (CRF)) or optical flow.
[0135] The received events can then be used to update the image, for example based on their gray level (in the case of ATIS type of sensors). The received events are also used to update the knowledge of the scene. For example, events against the background can increase the probability of a car to be present (the events can mean that a car just moved into the corresponding pixel) and events against the car mask can decrease the probability of a car to still be present (the events can mean that the car left the pixel).
[0136] After those updates, there can be for example zones with very low probability, for example below 5% (zones outside the mask where no events are received), zones with medium probability, for example between 5% and 95% (zones where events are received) and zones with very high probability, for example above 95% (zones inside the mask where no events are received).
[0137] Once all the updates are performed, a probabilistic model can be built using the knowledge matrix and the gray level image by statistical inference (finding the most numerically likely configuration). The mask can then be updated based on this probabilistic model.
[0138] Optionally, a regularization operation can be applied on the updated mask (for example to avoid leaks in the segmentation).
[0139] If the mask is bisected, a new track can be created; if the mask collapses or the probability is too low, a new detection iteration can be performed (if no object is detected, the track can be deleted). Otherwise, the mask can simply be added to the track.
[0140] Figure 7 is a possible embodiment of a device implementing the application.
[0141] In this embodiment, the device 700 comprises a computer comprising a memory 705 for storing program instructions loadable into the circuit and adapted to make the circuit 704 perform the steps of the application when the program instructions are run by the circuit 704.
[0142] The memory 705 can also store data and useful information for performing the steps of the application as described above.
[0143] The circuit 704 can be, for example:
[0144] - a processor or processing unit adapted to interpret instructions in a computer language, which can comprise a memory including said instructions, which can be associated with or attached to said memory; or
[0145] - an association of a processor / processing unit with a memory, which is adapted to interpret instructions in a computer language, which comprises said instructions; or
[0146] - an electronic card, in which the steps of the application are described within the silicon; or
[0147] - a programmable electronic chip, such as an FPGA chip (for "Field-Programmable Gate Array").
[0148] This computer comprises an input interface 703 for receiving asynchronous events according to the application and an output interface 706 for providing the results of the tracking method, for example a series of temporal positions of the object and the category of the object.
[0149] To facilitate the interaction with the computer, a screen 701 and a keyboard 702 can be provided and connected to the computer circuit 704.
[0150] Figure 2 is a flowchart describing a possible embodiment of the application. Part of this flowchart can represent the steps of an example of a computer program that can be executed by the device 700.
[0151] In interpreting this specification and the associated claims, the expressions "comprise," "contain," "include," "have," "is," and "possess," and conjugations thereof, are each used to express the inclusion of one or more elements, components, or steps, but not the exclusion of any other element, component, or step. Similarly, the expression "an element or component of one list is not an element or component of another list" does not mean that the element or component is excluded from the list, but rather that the element or component is not included in the other list.
[0152] Those skilled in the art will readily understand that the various parameters disclosed in this specification can be modified and that various embodiments disclosed can be combined without departing from the scope of the present invention.
Claims
1. A method of tracking objects in a scene observed by an event-based sensor, the event-based sensor having a matrix of sensing elements, wherein a respective event being produced asynchronously by each sensing element of the matrix from a change in light incident on the sensing element, the method comprising: detecting, by a first process (201), objects in the scene from the asynchronous events produced by the event-based sensor; and determining, by a second process (202), tracks associated with the objects detected from the asynchronous events produced by the event-based sensor; wherein a temporal resolution of the first process (201) is lower than a temporal resolution of the second process (202); the method further comprising: determining a relevance between the detected objects and the determined tracks (205); and updating the set of detected objects (203) and / or the set of determined tracks (204) based on the relevance (206), the updating comprising: if a given object of the detected objects (203) is not associated with any of the determined tracks (204), initializing a new track corresponding to the given object (403); if a given track of the determined tracks (204) is not associated with any of the detected objects (203), deleting the given track; and if a first object and a second object of the detected objects (203) are both associated with a first track of the determined tracks (204), initializing a second track from the first track, each of the first track and the second track corresponding to a respective one of the first object and the second object; if a given object of the detected objects (203) is associated with a first track and a second track of the determined tracks (204), merging the first track and the second track.
2. The method of claim 1, wherein, the first process (201) comprises determining a class of the detected objects.
3. The method according to any of the preceding claims, characterized in that, new iterations of the first process (201) are run according to a predefined temporal sequence.
4. The method of claim 1, wherein, the first process (201) comprises determining a region of interest (301) in the matrix of sensing elements, the region of interest (301) containing the detected objects; and wherein the second process (202) comprises determining the tracks based on the asynchronous events produced by the sensing elements in a region (302) containing the region of interest (301).
5. The method of claim 4, wherein, the second process (202) comprises updating the region of interest (301) based on the asynchronous events produced in a region (302) containing the region of interest (301).
6. The method according to any one of claims 4 and 5, characterized in that, the second process (202) further comprises running new iterations of the first process (201) based on a parameter of the region of interest (301), wherein the parameter is one of a size of the region of interest (301) and a shape of the region of interest (301).
7. The method of claim 4, wherein, the second process (202) further comprises running new iterations of the first process (201) if a number of events produced in the region of interest (301) during a certain time interval exceeds a threshold value.
8. The method of claim 1, further comprising: if a given object among the detected objects is associated with a given track among the determined tracks (204), updating the given track based on a position of the given object.
9. The method of claim 1, wherein, The second process (202) further comprises: for each detected object, determining a respective set of zones in the matrix of sensing elements, each zone of the set of zones being associated with a respective time; receiving events produced by the event-based sensor during a predefined time interval; determining groups of events, each group comprising at least one event among the received events; for each group of events, determining a respective region in the matrix of sensing elements, the region containing all events of the group of events; determining a correlation between the determined regions and the determined zones; and updating the set of detected objects and / or the set of determined tracks based on the correlation.
10. The method of claim 1, wherein, The first process (201) further comprises determining a segmentation mask (602) of the detected objects, and wherein the second process (202) comprises updating the segmentation mask (602) based on asynchronous events produced in a neighborhood of the segmentation mask (602).
11. The method of claim 10, further comprising: extracting first descriptive parameters associated with events received from the event-based sensor; extracting second descriptive parameters associated with events corresponding to the segmentation mask (602); and updating the segmentation mask (602) by comparing a distance between one of the first descriptive parameters and one of the second descriptive parameters with a predefined value; wherein the first and second descriptive parameters associated with an event comprise at least one of a polarity of the event and a velocity of the event.
12. A signal processing unit, comprising: an interface for connecting to an event-based sensor having a matrix of sensing elements that asynchronously produce events from light received from a scene, wherein a respective event is produced by each sensing element of the matrix from a change in light incident on the sensing element; and a processor configured to: detect, by a first process (201), objects in the scene from asynchronous events produced by the event-based sensor; and determine, by a second process (202), tracks associated with objects detected from asynchronous events produced by the event-based sensor; wherein a time resolution of the first process (201) is lower than a time resolution of the second process (202); the processes are further configured to: determine a correlation (205) between detected objects and determined tracks; and update a set of detected objects (203) and / or a set of determined tracks (204) (206) based on the correlation, the updating comprising: if a given object of the detected objects (203) is not associated with any of the determined tracks (204), initializing a new track corresponding to the given object (403); if a given track of the determined tracks (204) is not associated with any of the detected objects (203), deleting the given track; and if a first object and a second object of the detected objects (203) are both associated with a first track of the determined tracks (204), initializing a second track from the first track, each of the first and second tracks corresponding to a respective one of the first and second objects; if a given object of the detected objects (203) is associated with a first track and a second track of the determined tracks (204), merging the first and second tracks.
13. A computer readable medium having stored thereon computer code to be executed in a processor associated with an event-based sensor having a matrix of sensing elements that produce events asynchronously from light received from a scene, wherein, a respective event is generated by each sensing element of the matrix from changes in light incident on the sensing element, wherein the instructions cause the processor to: detect, by a first process (201), objects in the scene from the asynchronous events generated by the event-based sensor; and determine, by a second process (202), tracks associated with the objects detected from the asynchronous events generated by the event-based sensor; wherein a time resolution of the first process (201) is lower than a time resolution of the second process (202); determine associations between the detected objects and the determined tracks (205); and update the set of detected objects (203) and / or the set of determined tracks (204) based on the associations (206), the updating comprising: if a given object of the detected objects (203) is not associated with any of the determined tracks (204), initializing a new track corresponding to the given object (403); if a given track of the determined tracks (204) is not associated with any of the detected objects (203), deleting the given track; and if a first object and a second object of the detected objects (203) are both associated with a first track of the determined tracks (204), initializing a second track from the first track, each of the first and second tracks corresponding to a respective one of the first and second objects; if a given object of the detected objects (203) is associated with a first track and a second track of the determined tracks (204), merging the first and second tracks (405).
Citation Information
Patent Citations
Photoarray for Detecting Time-Dependent Image Data
US20080135731A1
Check-row corn-planter
US566576A