Multi-object tracking method
The object monitoring process addresses the challenge of initializing tracking for new objects by using convolutive neuron networks for detection and a coherence error index for efficient association, resulting in improved performance and robustness.
Patent Information
- Application Number
- FR2023002975
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-03-28
AI Technical Summary
Existing object monitoring processes struggle with efficient initialization of monitoring identifiers for new objects, especially when image acquisition frequency is low relative to the dynamics of mobile objects, leading to delays or failures in tracking.
A process for monitoring multiple objects in temporal image acquisitions, which includes object detection using convolutive neuron networks, delimitation framework determination, association of unique monitoring identifiers, and calculation of a coherence error index to efficiently initialize tracking for new objects.
The process improves monitoring performance by rapidly initializing tracking for new objects and associating non-consecutive but coherent detections, reducing calculation time and enhancing robustness during the initialization phase.
Smart Images

Figure 00000017_0000 
Figure 00000017_0001 
Figure 00000018_0000
Abstract
Description
Title of the invention: Method for tracking multiple objects
[0001] The invention relates to a method for tracking multiple objects in temporal acquisitions of images, in particular of vehicles in a road context, the images originating in particular from a fixed shooting device whose field of vision includes a traffic lane.
[0002] The invention applies, in particular, to technical fields such as security for the identification of offending vehicles traveling at high speed.
[0003] The objective of multiple object tracking is to spatially and temporally identify the objects of interest in a sequence of images, that is, for each object of interest, to place a bounding box around it and to give it a unique tracking identifier that is consistent over time. In other words, each tracked object is traced and its tracer is represented by a temporal sequence of bounding boxes, for example in the form of bounding boxes, each tracer having a unique tracking identifier.
[0004] Multiple object tracking also requires that the tracking algorithm be real-time (also called "online tracking"), meaning that the bounding boxes and associated tracking identifiers must be determined as soon as possible after an image is captured (e.g., with a latency below a determined threshold, such as five frames (or images)).
[0005] Finally, in general, tracking algorithms allow to predict the future position of multiple moving objects based on the history of individual positions.
[0006] A critical point therefore corresponds to the appearance of a new object in the field of vision of the image acquisition device because the object is then unknown, without determined speed, that is to say without means of estimating the location of the next frame in the following image. This phase is also called the initialization phase, it is completed when the new object has been able to be associated with a unique tracking identifier and its current speed determined.
[0007] Such methods for tracking objects are known from the state of the art, using methods for geometrically initializing tracking identifiers between two consecutively acquired images, in particular by means of determining the intersection score on the union (or IOU), such as the algorithm called SORT described in the article Simple Online and Realtime Tracking, by Bewley Alex, Ge Zongyuan, Ott, Lionel, Ramos Fabio and Upcroft Ben in 2016 IEEE International Conference on Image Processing (ICIP), pages 3464-3468, but these methods do not always allow the tracking identifiers to be initialized sufficiently quickly, in particular when the image acquisition frequency is low with regard to the dynamics of the moving objects tracked.
[0008] Improvements to these methods have been developed, including the hybrid geometric initialization method for tracking identifiers in which matching the tracker to its unique tracking identifier involves a comparative analysis of the objects in the images by deep learning using neural networks, such as the algorithm called DEEPSORT described in the article Simple Online and Realtime Tracking with a Deep Association Metric by Nicolai Wojke, Alex Bewley, and Dietrich Paulus in 2017 IEEE International Conference on Image Processing (ICIP), pages 3645—3649, however this method is computationally expensive and makes tracking slower.
[0009] The present invention aims to at least partially overcome these drawbacks, possibly leading to other advantages.
[0010] To this end, there is proposed, according to a first aspect, a method for tracking multiple objects in temporal acquisitions of images, in particular monocular videos, composed of a temporal sequence of images acquired consecutively, said method comprising: - a step of receiving n consecutively acquired images from said sequence, n being greater than or equal to 3; - a step of detecting an object in each of said images, in particular by means of a convolutional neural network; - a step of delimiting the detected object, in particular by determining a delimiting frame encompassing said at least one detected object, in a frame of reference of the image in which said object was detected; - a step of association under a unique tracking identifier specific to an object followed by a delimiting frame of a first image of the sequence of images to a delimiting frame of a second image of said sequence, said first and second images being in particular consecutive, according to a minimum overlap threshold; - a step of determining a current speed of said tracked object; - a step of determining the bounding box of the detected object without associated identifier among the n images; - a step of determining a set of candidate sequential multiplets of k detected object delimitation frames without identifier, with k greater than or equal to 3 and k less than or equal to n, each multiplet comprising as first element a frame without identifier of an image h among the n images, as second element a frame without identifier of an image i among the n images, with i greater than h and as third element a frame without identifier of an image j among the n images, with j greater than i; - a step of calculating a consistency error index for each candidate sequential multiplet; - a step of assigning to the candidate byte having the lowest consistency error index, and in particular lower than a predetermined consistency error threshold, another unique tracking identifier specific to another tracked object and association under said other unique tracking identifier of said k frames of the byte; and - a step of determining a current speed of said other tracked object.
[0011] The multiple object tracking method according to the invention makes it possible to improve object tracking performance without a significant increase in computation time, and contributes to robustness during the object tracking initialization phase by quickly matching the bounding box encompassing each new detected object appearing in the field of view of the image acquisition system to its unique tracking identifier.
[0012] Indeed, the object tracking method according to the invention allows rapid initialization even in the event of non-detection of an object in an image of the sequence and to associate under a unique tracking identifier detections from non-consecutive but mutually consistent images. Indeed, said first and second images are notably consecutive but are not necessarily so.
[0013] The object tracking method according to the invention applies equally well to sequences of images in the form of monocular videos acquired by a fixed acquisition device, as well as to sequences of images acquired by lidar in the form of point clouds, whether by a fixed or mobile device (in an autonomous vehicle for example, provided its instantaneous speed is known), or to a sequence of merged images originating from a radar and lidar acquisition device for example, fixed or mobile.
[0014] Furthermore, the object tracking method according to the invention allows the use of a two-dimensional or three-dimensional reference frame of the image, i.e. to work either in pixel coordinates of the image or in three dimensions, the calculation in three dimensions allowing simpler geometric calculations than in a pixel reference frame. In the case of the acquisition of a monocular image, in two dimensions, the transition to three dimensions is allowed by the knowledge of context data, i.e. here of the geometric environment in the case of a fixed device acquired during the installation and calibration of the fixed device. In the case of an acquisition initially in three dimensions (for example by means of a lidar, a time-of-flight sensor (ToF for Time of Fly in English), stereoscopic sensor or lidar coupled to a radar), the received image is already in three dimensions.In the case of a two-dimensional reference frame, the bounding boxes are notably rectangles and in the case of a three-dimensional reference frame, the bounding boxes are notably parallelepipeds.
[0015] Furthermore, the object tracking method according to the invention refers to a coherence error index, so as to reason in terms of error minimization, nevertheless a coherence index could also be used and to be retained the multiplet should have the highest coherence index.
[0016] The iterative loop repetition of the steps intended to assign a unique tracking identifier to the candidate multiplet having the lowest error consistency index, follows the logic of a greedy algorithm and makes it possible to associate a frame with only one identifier since once the frame is associated with an identifier it is no longer determined as a bounding box of a detected object without an associated identifier.
[0017] Advantageously, the attribution and association step is conditioned by the lowering of said lowest consistency error index to a predetermined consistency error threshold, which makes it possible not to match unrelated objects under a unique tracking identifier.
[0018] For example, the predetermined consistency error threshold is set based on the applications.
[0019] Advantageously, for each candidate multiplet the coherence error index is calculated as a function of a relative positioning of said frames without identifier of said multiplet, which allows a coherence calculation based on geometric considerations.
[0020] Advantageously, the relative positioning of said frames without identifier of said multiplet is determined relative to the positioning of a single reference point specific to each of said frames without identifier, which facilitates the calculation of the relative positioning.
[0021] Advantageously, in the case of a fixed acquisition device, the positioning reference is the frame of reference of the image in which said object was detected, this frame of reference being unique and identical for all the images.
[0022] For example, the single reference point of a frame is a point of the frame, in particular a midpoint of an upper segment of the frame.
[0023] Advantageously, the single reference point of each of said frames without identifier is the center of each of said frames without identifier, which makes it easy to determine, and less sensitive to acquisition hazards than a corner of the frame in the case of an image comprising an incomplete view of the vehicle, whether during the partial incursion of the vehicle into the field of vision of the acquisition device (at the edge of the image) or during overtaking by another vehicle partially hiding it.
[0024] Advantageously, said unique reference points specific to each of said frames without identifier are determined by a convolutional neural network having learned on a base of images comprising truncated objects, which gives more latitude on the definition of the single reference point of the frame because it can then be defined for example as being the center of the object, without necessarily being geometrically at the center of the frame.
[0025] Advantageously, the coherence error index is calculated as a function of k-1 vectors each connecting the reference point of an element to the reference point of the following element, said vectors being in particular a function of a time difference between the acquisitions of the two images, by calling - first vector, the vector connecting the reference point of the first element and the reference point of the second element of the multiplet, said first vector being in particular a function of a time difference between the two acquisitions of the images h and i, and; - second vector, the vector connecting the reference point of the second element and the reference point of the third element of the multiplet, said second vector being in particular a function of a time difference between the two acquisitions of images i and j, which makes it possible to take into consideration pairwise alignments, taking into account the temporal difference making it possible to dimension a speed and in particular to adapt to the case of a multiplet of three elements whose elements do not belong to directly consecutive images, or in the case of image acquisition with irregular temporal sampling for example.
[0026] Advantageously, each image received includes the timestamp of its acquisition time in the time reference of the acquisition device by means, for example, of metadata integrated into said image.
[0027] According to one embodiment, the consistency error index is calculated as a function
[0028]
[0029] of an angle between said first and second vectors, and in particular of a ratio between the first and second vector, such that — ^1 - cos^zjj| 1 - -^4-1 ' 'anê'c between the two vectors making it possible to assess their collinearity, and the ratio of the vectors making it possible to assess their relative magnitude. According to another embodiment, the consistency error index is calculated as a function of the average v of said k-1 vectors (vi), such that vk - 1) — ? ' This C'U' Allows a symmetrical approach and an ap simple plication to multiplets of more than three elements. Advantageously, said detected objects are vehicles or pedestrians, which allows an application to the detection of traffic offenses as well as an application to the tracking of pedestrians, for which it is common for the image acquisition frequency to be low with regard to the dynamics of movement of the object.
[0030] The invention also provides a computer program product comprising the program instructions implementing the steps of the determination method according to the invention, when the program instructions are executed by a computer, having the same advantages as the method according to the invention.
[0031] The invention also relates to a multiple object tracking system comprising an image acquisition device and a computer program product according to the invention, having the same advantages as those of the method according to the invention. Preferably, said computer program product is stored in a memory of the image acquisition device, the latter comprising in particular a monocular camera.
[0032] Advantageously, said multiple object tracking system is fixed.
[0033] The invention will be better understood and its advantages will appear better on reading the detailed description which follows, given for information purposes only and in no way limiting, with reference to the appended drawings in which:
[0034] - [Fig.l] [Fig.l] schematically shows the steps of the method of the invention according to an example.
[0035] - [Fig.2] [Fig.2] shows an example of an image without overlap between the frames of four consecutively acquired images.
[0036] - [Fig.3] [Fig.3] shows an example of a triplet.
[0037] - [Fig.4] [Fig.4] is a schematic block diagram of a processing device information for implementing one or more embodiments of the invention.
[0038] Identical elements shown in the above-mentioned figures are identified by identical reference numerals.
[0039] [Fig.l] schematically presents the steps of a method for tracking multiple objects in temporal image acquisitions according to the invention.
[0040] Temporal image acquisitions are composed of a temporal sequence of consecutively acquired images.
[0041] These images may in particular come from monocular videos, or from image acquisition systems (for example in the form of point clouds) in 3 dimensions (for example lidar, time-of-flight sensor, stereoscopic sensor, combined radar and lidar).
[0042] The method comprises the following steps: - a step E1 of receiving n consecutively acquired images from said sequence, n being greater than or equal to 3; - a step E2 of object detection in each of said images, in particular by means of a convolutional neural network; - a step E3 of delimitation of the detected object, in particular by determining a bounding box encompassing said at least one detected object, in a frame of reference of the image in which said object was detected; - a step E4 of association under a unique tracking identifier specific to the object followed by a delimiting frame of a first image of the sequence of images to a delimiting frame of a second image of said sequence, said first and second images being in particular consecutive (but not necessarily), according to a minimum overlap threshold; - a step E5 of determining a current speed of said tracked object; - a step E6 of determining the delimitation frame of the detected object without associated identifier among the n images, including in particular the images called here h, ietj,; - a step E7 of determining a set of candidate sequential multiplets of k frames (dl, d2, d3) for delimiting a detected object without identifier, each multiplet comprising k elements, with k greater than or equal to 3 and k less than or equal to n, each multiplet comprising as first element a frame without identifier of an image h, as second element a frame without identifier of an image i, with i>h and as third element a frame without identifier of an image j with j>i; - a step E8 of calculating a consistency error index for each candidate sequential multiplet; - a step E9 of assigning to the candidate multiplet having the lowest consistency error index, and in particular lower than a predetermined threshold, another unique tracking identifier and associating under said other unique tracking identifier said k frames of the n images of the multiplet; and - a step E10 of determining a current speed of said other tracked object.
[0043] Steps E2 of object detection and E3 of delimitation of each detected object are executed on each received image, the detection of the objects in each image being carried out in particular in a known manner by artificial intelligence, by means of convolutional neural networks, as described in the article You Only Look Once: Unified, Real-Thne Object Detection, by Joseph Redmon, Santosh Divvala, Ross Girshick and Ali Farhadi, in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 2016, pages 779-788.
[0044] These two steps E2 of object detection and E3 of delimitation of each detected object can be executed simultaneously, in particular by a single neural network, the order of these two steps can also be reversed.
[0045] The subsequent steps, from E4 to E9, are more precisely dedicated to the initialization phase of object tracking, i.e. when a new object appears in the field of vision of the acquisition device and does not yet have a tracking identifier or known speed.
[0046] Step E4 of association by sequential matching is known and is based on a comparison of overlap of delimiting frames between said two images, and then on association of a delimiting frame of the first image with a delimiting frame of the second image as a function of a minimum overlap threshold.
[0047] The evaluation of the overlap is based on an analysis of proximity and / or overlap rate, for example by means of the evaluation of the ratio of the area of the intersection of the frames divided by the union of the frames (also called loU) between the two frames each belonging to one of said two images, in particular with a threshold of 0.3.
[0048] If the new object appearing in the field of vision of the acquisition device is in particular too fast or is for example not detected in the second acquired image, because the object is for example hidden by another object, the object will then not be able to be associated with any frame during this association step E4. At the end of step E4, there will then remain unpaired frames in the images.
[0049] The following steps E6 to E9 aim to reduce, in this initialization phase, the number of unpaired frames remaining in the images, and thus improve the tracking of the objects in the sequence. To do this, steps E6 to E9 are carried out iteratively, so that once the multiplet with the lowest coherence error index is preferably lower than a predetermined coherence error threshold, the frames of this multiplet are associated under a unique tracking identifier specific to this object, the speed of this object is determined, then iteratively the next most promising multiplet is determined on the basis of the frames without remaining identifiers and so on until no more multiplets satisfy the condition linked to the coherence error threshold.
[0050] This method makes it possible to better initialize the tracking of new objects, which is of great importance when starting up the system but also continuously when a new object appears in the field of vision of the acquisition device, because without initialization of the tracer the object cannot be tracked.
[0051] Indeed, this initialization by multiplet (k-tuples with k greater than or equal to 3) makes it possible to better track fast objects for which an initialization phase based on an overlap analysis was not successful, while guaranteeing robustness and avoiding erroneous pairings since if it would be probable to have two consecutive coherent detections to initialize a tracer while in reality, these detections do not correspond to the same object, it is much less probable that three or more consecutive detections (even if the images comprising each of these detections are not consecutive) are coherent by chance. Preferably, k is chosen close to n, for example for n=6, k could be equal to 4
[0052] [Fig.2] illustrates the application of the invention to the tracking of moving objects and more precisely here of vehicles in the context of road control nevertheless other applications are conceivable, in particular for other types of vehicles than automobiles. The figure shows an image taken at time t3 on which have been superimposed: - the bounding box dl of the object detected in the image tl, which is here the first image in which an object is detected, - the frame d2 of delimitation of the object detected in the image t2, - the frame d3 of delimitation of the object detected in the image t3.
[0053] By applying the classic method of initializing a tracer, stopping at step E4, no association could have been made here since no overlap exists between said frames dlet d2 (nor d2 and d3) of the three consecutively acquired images. At time t3 the object would therefore still not have been able to be tracked.
[0054] In the example illustrated here, ultimately, after a few additional time steps the perspective could have allowed initialization of the tracking of the object by applying step E4 because an overlap would have ended up existing but this is not always the case and the initialization is then potentially carried out too late to draw up a report with the required plate information.
[0055] The method then continues with step E6 (since here no object could be tracked, step E5 is therefore irrelevant) of determining the delimitation frame of the detected object without associated identifier among the three images; that is to say here the frames d1, d2 and d3.
[0056] Here the application of step E7 of determining a set of candidate sequential multiplets of said detected object delimitation frames without identifier, only concerns a triplet (d1,d2,d3). A detection having been able to be made on each image, there are here as many frames as there are images.
[0057] According to the embodiment illustrated in [Fig.3], in step E8 the coherence error index for this candidate sequential triplet is here calculated as a function of 2 vectors vl, v2 represented in the form of arrows and each connecting the reference point of an element to the reference point of the following element, said vectors vl, v2 being in particular a function of a time difference between the acquisitions of the two images, by calling - first vector vl, the vector connecting the reference point of the first element dl and the reference point of the second element d2 of the triplet, said first vector being in particular a function of a time difference between the two acquisitions of the images h and i, and; - second vector v2, the vector connecting the reference point of the second element d2 and the reference point of the third element d3 of the multiplet, said second vector being in particular a function of a time difference between the two acquisitions images i and j.
[0058] More precisely, the calculation of the consistency error index IEC is here a function of an angle a between said first vl and second v2 vectors, and in particular of a ratio between the first vl and the second v2 vector, such that j pç = f 1 - co> / a^! 1 - i • \ \ / / ] h'2| i
[0059] The angle a between the two vectors vl, v2 allows us to assess their collinearity, it is represented between the vector v2 and the dotted line collinear with vl. The first factor 1 - cos(a) expresses that the two vectors must be aligned, and the second factor 11 Mi | expresses that they must have the same norm, the ratio of the vectors I ' I | allowing us to appreciate their relative size.
[0060] Taking into account the time difference by defining each vector as the ratio of a distance between two reference points to the time difference between the acquisitions of the two related images makes it possible to dimension a speed vector, and thus to be robust to cases in which the frames of the multiplet come from images not acquired regularly or non-consecutive acquired with regularity (case of non-detection on one of the acquired images).
[0061] In the example illustrated in [Fig.3] the chosen reference point of the frame is the center of said frame.
[0062] Then the execution of step E9 assigns to the candidate multiplet having the lowest consistency error index IEC and lower than a predetermined consistency error threshold another unique tracking identifier and associates under said other unique tracking identifier specific to this new tracked object said frames d1, d2 and d3 of the n images of the triplet.
[0063] In the illustrated example, the coherent triplet is a triplet such that the three detection frames d1, d2, d3 are aligned and uniformly distant, with a predetermined coherence error threshold. For the application of the tracking method to vehicles in the context of road control, the predetermined coherence error threshold is for example here set at 0.01 for k=3 and an acquisition speed of 15 images per second, since in this application case if we consider a short time of the order of 0.1s, at high speed it is not possible for the vehicle to significantly change direction or speed.
[0064] According to another embodiment, the consistency error index IEC is calculated according to _ 1 V b'M with v the average of said k-1 vectors vi, this which would have been written in the case of a triplet: jgçy2) = - [ ILiïU 4. 1112211 j
[0065] This hypothesis of consistency can also be expressed by the fact that the vectors vl and v2 must be approximately the same, and therefore be equal to their average. Another advantage of this method of calculating the IEC consistency error index is that it is symmetrical with respect to vl and v2, and only cancels out when vl=v2.
[0066] For reasons of clarity, the method of tracking objects according to the invention has been illustrated here with a triplet but is easily generalized to quadruplets, quintuplets, etc.
[0067] The interest of considering k-uplets with k greater than 3 is twofold: - the higher the number of elements in the multiplet, the higher the level of confidence in the tracers initialized in this way, because they are less likely to be the result of chance; - they allow a detection to be missed (which can happen frequently, because the detector is not perfect): for example, for a quadruplet initialization, among the last four sets of detections, we can look at the consecutive quadruplets, but also the sets of triplets in which the frames come from three out of four acquisitions and therefore do not belong to images acquired all consecutively. Indeed, the vi being defined by the ratio between a distance and a time difference, their estimation is not disturbed by calculating them on non-consecutive frames.
[0068] Preferably, k is less than or equal to 7 so as not to significantly increase the calculation time, nor delay the initialization time too much and so as to maintain consistency with the alignment hypothesis which could no longer be appropriate for the application if k is too high (for example a vehicle could turn).
[0069] In the case where the acquisition device is mobile, knowledge of its instantaneous speed is required (estimated or measured) so as to compensate the positioning reference by the movement of the device.
[0070] [Fig.4] is an example of a schematic block diagram of a multiple object tracking system 10 comprising an image acquisition device 100 comprising an information processing device 106 in which a computer program product according to the invention is stored. The information processing device 106 is capable of implementing one or more embodiments of the invention.
[0071] The acquisition system 100 is controlled by an information processing device 106 which makes it possible, for example, to control the monocular camera 102. The information processing device 106 receives and processes the images received from the camera 102. The information processing device 106 is typically a peripheral such as a microcomputer, a mobile telecommunications terminal or any other device allowing the execution of a computer program responsible for controlling the camera, acquiring the images and the various steps of the method according to the invention.
[0072] In certain embodiments, the monocular camera 102 may be supplemented or replaced by additional sensors contributing to the acquisition of three-dimensional images, such as, for example, a time-of-flight sensor, a radar, a lidar or any other sensor.
[0073] The device 106 comprises a communication bus connected to: - a central processing unit 601, such as a microprocessor, denoted CPU; - a random access memory 602, denoted RAM, for storing the executable code of the method for implementing the invention as well as the registers adapted to record variables and parameters necessary for implementing the method according to embodiments of the invention; the memory capacity of the device can be supplemented by an optional RAM memory connected to an expansion port, for example; - a read-only memory 603, denoted ROM, for storing computer programs for implementing the embodiments of the invention; - a network interface 604, denoted NET, is normally connected to a communication network on which digital data to be processed are transmitted or received. The network interface 604 may be a single network interface, or composed of a set of different network interfaces (e.g. wired and wireless interfaces or different types of wired or wireless interfaces). Data packets are sent over the network interface for transmission or are read from the network interface for reception under the control of the software application running in the processor 601; - a user interface 605 for receiving input from a user or for displaying information to a user; - a storage device 606 as described in the invention and noted HD; - an input / output module 607 for receiving / sending data from / to external devices such as hard disk, removable storage medium or others.
[0074] The executable code may be stored in a non-volatile memory 603, for example a flash memory or a read-only memory, on the storage device 606 or on a removable digital medium such as for example a disk. According to a variant, the executable code of the programs may be received by means of a communication network, via the network interface 604, in order to be stored in one of the storage means of the communication device 600, such as the storage device 606, before being executed.
[0075] The central processing unit 601 is adapted to control and direct the execution of the instructions or portions of software code of the program or programs according to one of the embodiments of the invention, instructions which are stored in one of the aforementioned storage means. After power-up, the CPU 601 is capable of executing instructions from the main RAM memory 602, relating to a software application. Such software, when executed by the processor 601, causes the described methods to be executed.
[0076] In this embodiment, the information processing device 106 is a programmable apparatus that uses software to implement the invention. However, alternatively, the present invention may be implemented in hardware (e.g., in the form of an application-specific integrated circuit (ASIC) or in the form of a field-programmable gate array (FPGA)).
[0077] Although the present invention has been described above with reference to specific embodiments, the present invention is not limited to the specific embodiments, and modifications that fall within the scope of the present invention will be apparent to a person skilled in the art.
Claims
Claims
1. Method for tracking multiple objects in temporal image acquisitions, in particular monocular videos, composed of a temporal sequence of consecutively acquired images, said method comprising: - a step (El) of receiving n consecutively acquired images from said sequence, n being greater than or equal to 3; - a step (E2) of object detection in each of said images, in particular by means of a convolutional neural network; - a step (E3) of delimiting the detected object, in particular by determining a delimiting frame encompassing said at least one detected object, in a frame of reference of the image in which said object was detected; - a step (E4) of association under a unique tracking identifier specific to an object followed by a delimiting frame of a first image of the sequence of images to a delimiting frame of a second image of said sequence, said first and second images being in particular consecutive, according to a minimum overlap threshold; - a step (E5) of determining a current speed of said tracked object; said method being characterized in that it further comprises: - a step (E6) of determining the delimitation frame of the detected object without associated identifier among the n images; - a step (E7) of determining a set of candidate sequential multiplets of k frames (dl, d2, d3) of delimitation of detected object without identifier, with k greater than or equal to 3 and k less than or equal to n, each multiplet comprising as first element a frame (dl) without identifier of an image h among the n images, as second element a frame (d2) without identifier of an image i among the n images, with i greater than h and as third element a frame (d3) without identifier of an image j among the n images, with j greater than i; - a step (E8) of calculating a consistency error index (CEI) for each candidate sequential multiplet; - a step (E9) of assigning to the candidate multiplet having the lowest consistency error index (CEI), and in particular lower than a predetermined consistency error threshold, another unique tracking identifier specific to another tracked object and association under said other unique tracking identifier of said k frames (dl,d2,d3) of the multiplet; and - a step (E10) of determining a current speed of said other tracked object.
2. Method for tracking multiple objects according to the preceding claim, characterized in that for each candidate multiplet the consistency error index (CEI) is calculated as a function of a relative positioning of said frames (dl, d2, d3) without identifier of said multiplet.
3. Method for tracking multiple objects according to the preceding claim, characterized in that the relative positioning of said frames (dl, d2, d3) without identifier of said multiplet is determined relative to the positioning of a single reference point specific to each of said frames (dl, d2, d3) without identifier.
4. Method for tracking multiple objects according to the preceding claim, characterized in that the single reference point of each of said frames (dl, d2, d3) without identifier is the center of each of said frames (dl, d2, d3) without identifier.
5. Method for tracking multiple objects according to any one of claims 3, 4 characterized in that said unique reference points specific to each of said frames (dl, d2, d3) without identifier are determined by a convolutional neural network having learned on a base of images comprising truncated objects.
6. Method for tracking multiple objects according to any one of claims 3 to 5 characterized in that the coherence error index (CEI) is calculated as a function of k-1 vectors (vl, v2) each connecting the reference point of an element to the reference point of the following element, said vectors (vl, v2) being in particular a function of a time difference between the acquisitions of the two images, by calling - first vector (vl), the vector connecting the reference point of the first element (dl) and the reference point of the second element (d2) of the multiplet, said first vector being in particular a function of a time difference between the two acquisitions of images h and i, and; - second vector (v2), the vector connecting the reference point of the second element (d2) and the reference point of the third element (d3) of the multiplet, said second vector being in particular a function of a time difference between the two acquisitions of images i and j.
7. Method for tracking multiple objects according to the preceding claim, characterized in that the consistency error index (CEI) is calculated as a function of an angle (a) between said first (vl) and second (v2) vectors, and in particular of a ratio between the first and the second vector, such that IEC = _ JW ]. MI ■ \ \ / / 1 1v -1 1
8. Method for tracking multiple objects according to claim 6 characterized in that the consistency error index (CEI) is calculated as a function of the average T' of said k-1 vectors (vi), such that / £c(vl, =
9. A method of tracking multiple objects according to any one of the preceding claims, characterized in that said detected objects are vehicles or pedestrians.
10. A computer program product comprising the program instructions implementing the steps of the determination method according to any one of claims 1 to 9, when the program instructions are executed by a computer.