Improved person identification using classification fusion

WO2026104888A1PCT designated stage Publication Date: 2026-05-21ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
ORANGE SA
Filing Date
2025-10-16
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing multi-view identification systems in video surveillance fail to accurately identify individuals due to occlusion, leading to incorrect identifications in crowded or complex environments.

Method used

A method and system that utilize a plurality of cameras with synchronized views to calculate feature distances across multiple images, employing pre-trained supervised machine learning engines to identify individuals based on vectors of features, even in cases of partial occlusion.

Benefits of technology

Ensures accurate person identification by correlating features across multiple views, enhancing reliability and reducing errors even when individuals are partially obscured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025000546_21052026_PF_FP_ABST
    Figure IB2025000546_21052026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a multi-view classification of a person, involving a method (P4) comprising : - obtaining (S43) a plurality of query images of a target person (Persf ) from a plurality of views captured by a plurality of cameras respectively at the same capture time; - for each query image of said plurality of query images, calculating (S45) a plurality of distances between a vector of features associated with the query image and vectors of features associated with images of a gallery of images representing a person in a plurality of views of the area captured respectively at a plurality of times, distinct of said capture time, by the camera which has captured the view from which the query image is obtained; and - identifying (S47) said target person based on said plurality of distances calculated for each query image.
Need to check novelty before this filing date? Find Prior Art

Description

IMPROVED PERSON IDENTIFICATION USING CLASSIFICATION FUSIONTECHNICAL FIELD

[0001] This disclosure pertains to the field of image analysis. It more specifically pertains to the field of identification of persons in images captured by cameras.BACKGROUND

[0002] Video surveillance systems use cameras to monitor and record activities within a specific area to enhance security and safety. These systems typically consist of a network of cameras strategically placed to capture video footage of various locations. The captured footage can be analyzed in real-time or stored for later review.

[0003] Components of surveillance systems include in general one or more of:Cameras: These can be fixed or pan-tilt-zoom (PTZ) cameras, capturing images and videos in different resolutions and fields of view.Recording Devices: recording devices such as Digital Video Recorders (DVRs) or Network Video Recorders (NVRs) store the footage for future access and analysis.Monitoring Stations: Screens and control panels where security personnel can view live feeds and manage camera operations.Software: Advanced software solutions provide features like motion detection, facial recognition, and automated alerts and are used to analyze the images and videos captured by the cameras.

[0004] Video surveillance systems are widely used in various settings, including public spaces, businesses, and private properties, to deter crime, ensure public safety, and provide valuable evidence in investigations.

[0005] Video surveillance systems are typically used to identify the persons (for example the pedestrians) that are located in the monitored area. In general, video surveillance systems identify persons in the recorded images and video through a series of steps involving image processing and machine learning techniques. Such steps may for example comprise:• Preprocessing: The captured images can be processed to enhance quality, adjust lighting, and remove noise, preparing them for analysis;• Bounding box Detection: Algorithms detect and isolate human figures within the images, often using techniques like background subtraction or motion detection. “Bounding boxes” that each comprise an image of a person can thus be extracted. The bounding boxes generally present themselves as rectangles that comprise a subset of an image that represents a person;• Feature Extraction: Key features such as facial characteristics, body shape, and clothing are extracted from the detected figures. This step is usually performed by a trained supervised machine learning engine, which in the context of image analysis often makes use of convolutional neural networks (CNNs) to identify distinguishing features;• Classification / Recognition : The extracted features are compared against a database of known individuals. Machine learning models, such as deep learning classifiers, are used to match features and identify individuals. In practice, the identification of a person is performed by a classification task of the features into classes, wherein each class corresponds to a defined person.

[0006] Multi-view identification corresponds to a situation where a person is identified from a plurality of images of a same area captured from different cameras / field of views rather than a single camera.

[0007] An accurate person identification is crucial for video surveillance systems. However, the accuracy of the person identification can be significantly reduced by occlusion. Occlusion occurs when an object or person is partially or fully obscured by another object or person in an image or video. This can significantly reduce the accuracy of person identification in several ways.

[0008] In particular, key identifying features, such as facial characteristics or body shape, may be hidden, making it difficult for algorithms to extract necessary information. With parts of the person not visible, the system may also not have enough data to make an accurate identification. Occlusion can therefore lead to confusion between individuals, resulting in incorrect identifications. The likelihood of errors increases, with the video surveillance system either failing to identify a person (false negative) or incorrectly identifying someone else as the target (false positive).

[0009] Occlusion is very often found in crowded environments and / or complex urban environments, because the likelihood of a person being partially occluded by another person or a part of the urban environment is particularly high. However, these urban and / or crowded environments are typically environments wherein the video surveillance systems are the most useful.

[0010] Multi-view identification could be used to increase the accuracy of the video surveillance systems in crowded or complex environments, because a person that is occluded in one view may not be occluded in another view. Furthermore, the features extracted from one view can be matched with features extracted with other views to improve the person identification.

[0011] The article “MCTR: Multi Camera Tracking Transformer”, by Alexandru Niculescu-Mizil et al. (arXiv.org, 11 September 2024, XP093262006) proposes a multi-camera person tracking architecture relying on a DETection Transformer (DETR) model, where persons may be detected for each image using this DETR model.

[0012] However, such existing multi-view identification systems fail to efficiently match the information obtained from different cameras to obtain an accurate person identification.

[0013] There is therefore the need for a method and system to perform multi-view identification, that is able to accurately identify the persons, when at least one view suffers from occlusion.SUMMARY

[0014] This disclosure improves the situation.

[0015] It is proposed a method implemented by one or more processing unit of a computing system, said computing system comprising a plurality of cameras having different fields of views of a same area, each camera of said plurality being respectively configured to capture a view of the area at a same capture time, said method comprising:obtaining a plurality of query images of a target person from a plurality of views captured by said plurality of cameras respectively at the same capture time;for each query image of said plurality of query images, calculating a plurality of distances between a vector of features associated with the query image and vectors of features associated with images of a gallery of images representing a person in a plurality of views of the area captured respectively at a plurality of times, distinct of said capture time, by the camera which has captured the view from which the query image is obtained; andidentifying said target person based on said plurality of distances calculated for each query image.

[0016] In the context of the present invention, it should be understood that the expressions "at least one of ..., ... or ..." and "one or more of ..., ... or ..." are used interchangeably unless otherwise specified. Specifically, the use of the expressions "at least one of A, B or C." and "one or more of A, B or C" does not limit the invention to a single element chosen among A, B and C (A, B and C being selected by means of non limitative example only, the expressions “one or more of ... or” and “at least one of ... or” being applicable to any number of elements equal to or higher than 2) but should be interpreted as including the possibility of selecting multiple elements among A, B and C, unless clearly indicated otherwise. Therefore, the terms "at least one" and "one or more" also encompasses the situation where only one element is present. The expressions "at least one of A, B or C." and "one or more of A, B or C" encompasses all the possibilities listed below: "A or B or C or (A and B) or (A and C) or (B and C) or (A and B and C)". It shall also be noted that each of the expressions A, B and C may itself refer to more than one element, in particular if the wording "a," "at least one" or "one or more" is used.

[0017] By “a view”, we designate an image of an area captured by a camera.

[0018] By “a same capture time”, we designate a capture of a view by the camera that ensures that the views are captured by the different cameras in a synchronized manner. A same capture time may be defined for example by a same time in a discrete time series, or a predefined time interval in a continuous time series. The predefined time interval may be defined for example to be low enough to ensure that persons / pedestrians did not perform significant moves during the predefined time interval.

[0019] The synchronization of the cameras to capture the views at the same capture time may be for example performed by:Synchronizing the internal clocks of the cameras, and making them capture the views at predefined I regular times;Causing the cameras to capture the views upon the reception of a capture instruction broadcasted to the plurality of cameras.

[0020] This method allows an accurate prediction of the person represented by the query images even in case of occlusion, because the determination of the person is done based upon a plurality of views and may be achieved even if the person is partially occluded in some of the images.

[0021] In another aspect of the invention, it is proposed a computing system comprising:a plurality of cameras having different fields of views of a same area, each camera of said plurality being respectively configured to capture a view of the area at a same capture time;one or more computing device comprising one or more processing unit configured to :obtain a plurality of query images of a target person from a plurality of views captured by said plurality of cameras respectively at the same capture time;for each query image of said plurality of query images, calculate a plurality of distances between a vector of features associated with the query image and vectors of features associated with images of a gallery of images representing a person in a plurality of views of the area captured respectively at a plurality of times, distinct of said capture time, by the camera which has captured the view from which the query image is obtained; andidentify said target person based on said plurality of distances calculated for each query image.

[0022] In another aspect, it is proposed a computing system comprising: a plurality of cameras having different fields of views a same area, each camera of said plurality being respectively configured to capture a view of the area at a same capture time; one or more computing device comprising: one or more communication endpoints configured to receive, from said plurality of cameras respectively, the views captured by said plurality of cameras at said capture time; one or more memory storing, for each camera belonging said plurality of cameras respectively: a gallery of images associated to the camera, each image of said gallery representing a person, and being obtained by detecting bounding boxes in a plurality of views of the area captured by the camera at a plurality of times respectively, each of said plurality of times being distinct of said capture time; a pre-trained supervised machine learning engine associated with said camera, said supervised machine learning engine being previously trained to identify persons using a labelled trained dataset comprising training images that represent a person, said training images being obtained by detecting bounding boxes in a plurality of views of the area captured by the camera; one or more processing unit configured to execute a method as defined here.

[0023] We designate by “processing unit” an electronic component capable of performing electronic or computer calculations for a function. A processing unit can designate any type of processor or electronic component capable of performing digital calculations. For example, a processing unit can be an integrated circuit, an ASIC (from the English acronym "Application-Specific Integrated Circuit", literally in French "integrated circuit specific to an application", a microcontroller, a microprocessor, a Digital Signal Processor (DSP), a processor, a Graphical Processing Unit (GPU). A processing unit according to the invention is not limited to a particular type of calculation architecture. For example, a processor can implement a Harvard or Von Neumann type architecture.

[0024] We designate by “Memory”, or "Memory storage" a digital electronic device or digital electronic device component used to store data. Different types of memories can be used in the invention, such as read only memory (ROM), random access memory (RAM), volatile memory or flash memory. A computing device according to the invention can be equipped with one or more non-volatile memories which can be of different types such as mass memories, flash memory, read only memories or SSD memories. A computing device according to the invention may also comprise one or more random access memories such as RAM, DRAM, SRAM, DPRREAM, VRAM, eDRAM or 1T-SRAM.

[0025] We designate by “communication endpoint” a specific interface or point within a computing system that facilitates the exchange of data between devices or networks. It serves as a connection point for sending and receiving information.

[0026] The computing system is thus able, thanks to the execution of a method according to an embodiment, to use the images captured by the plurality of camera to accurately identify the persons in the area. In particular,the identification of a person remains reliable even if the person is occluded in one or more of the views that are used for the identification.

[0027] In another aspect, it is proposed a computer software comprising instructions to implement at least a part of a method as defined here when the software is executed by a processor.

[0028] In another aspect, it is proposed a computer-readable non-transient recording medium on which a software is registered to implement the method as defined here when the software is executed by a processor.

[0029] In another aspect, it is proposed a method implemented by one or more processing unit of a computing system, said computing system comprising: a plurality of cameras having different fields of views of a same area, each camera of said plurality being respectively configured to capture a view of the area at a same capture time; one or more memory storing, for each camera belonging said plurality of cameras respectively: a gallery of images associated to the camera, each image of said gallery representing a person, and being obtained by detecting bounding boxes in a plurality of views of the area captured by the camera at a plurality of times respectively, each of said plurality of times being distinct of said capture time; a pre-trained supervised machine learning engine associated with said camera, said supervised machine learning engine being previously trained to identify persons using a labelled trained dataset comprising training images that represent a person, said training images being obtained by detecting bounding boxes in a plurality of views of the area captured by the camera; said method comprising: receiving the plurality of views captured by said plurality of cameras respectively at said capture time; detecting in said plurality of views bounding boxes, each bounding box encompassing an image that represents a person; obtaining a plurality of query images of a target person, said plurality of query images comprising for each view a query image associated to the view, said query image associated to the view being one of the images encompassed in the bounding boxes detected in the view; for each query image of said plurality of query images, said query image being associated to a view captured by a camera: using the supervised machine learning engine associated to the camera to determine: a vector of features associated with the query image; and, for each image of the gallery associated with the camera: a vector of features associated with the image of the gallery; a unique identifier, among a predefined set of unique identifiers, of a person represented by the image of the gallery; calculating a plurality of distances between the vector of features associated with the query image, and the vectors of features associated with the images of the gallery; identifying said target person based on said plurality of distances for each query image.

[0030] By “a bounding box”, we designate rectangular frame used in image processing to define the position and size of an object within an image. It is typically used to isolate and identify specific objects, such as people, within a scene.

[0031] This method allows an accurate prediction of the person represented by the query images even in case of occlusion, because the determination of the person is done based upon a plurality of views, and because the use of a model specifically trained for each view / camera allows correlating the features of query images and query images even if the person is partially occluded in some of the images.

[0032] The following features, can be optionally implemented, separately or in combination one with the others:

[0033] Said computing system further comprises at least one memory which stores, for at least one camera of said plurality of cameras, a machine learning engine associated with said camera, the method further comprises using the machine learning engine associated to said camera to determine the vector of features associated with the query image obtained from the view captured by said camera and, for each image of thegallery of images associated with said camera, the vector of features associated with the image of the gallery. By “gallery of images associated with said camera”, it is meant a plurality of images representing a person in a plurality of views of the area respectively captured by this camera at a plurality of times, distinct from the capture time of the query image by said camera.

[0034] This machine learning engine may be in particular a supervised a machine learning engine, such as a supervised machine learning engine which has been previously trained to identify persons using a labelled trained dataset comprising training images that represent a person.

[0035] Said machine learning engine comprises a convolutional neural network. Convolutional neural networks are particularly efficient for performing image classification tasks, and thus to allow an accurate person identification.

[0036] Said method further comprises using the machine learning engine associated to said camera to determine, for each image of the gallery of images associated with said camera, a unique identifier, among a predefined set of unique identifiers, of a person represented by the image of the gallery, the identification of said target person being further based on said unique identifiers.

[0037] Identifying said target person based on said plurality of distances for each query image comprises: for each candidate unique identifier belonging to said predefined set of unique identifiers : for each query image associated with a view captured by a camera belonging to said plurality of query images: determining a confidence score that said query image represents the person associated to said candidate unique identifier, based on: the distances with the images of the gallery associated to said camera, and the unique identifiers of the persons represented by the images of the gallery determined by the supervised machine learning engine; obtaining an average confidence score that the plurality of query image represents the person associated to the candidate unique identifier by averaging the confidence scores for each image of said plurality of query images; identifying said target person as the person associated to the unique identifier that has the highest average confidence score. These additional steps provide an accurate identification of the person, because the person associated with the highest confidence among all the images of the plurality of images is identified.

[0038] Identifying said target person based on said plurality of distances for each query image comprises: for each query image associated with a view captured by a camera belonging to said plurality of query images: sorting the images of the gallery associated to the camera by increasing distance; determining a reference distance as the lowest distance of order m, m being a predefined integer; identifying the target person as the person identified in the image of the gallery having the lowest reference distance of order m. This identification method provides accurate results, because the identified person will be a person represented in a gallery image having a lowest distance with one of the query images.

[0039] Identifying said target person based on said plurality of distances for each query image comprises: for each query image associated with a view captured by a camera: sorting the images of the gallery associated to the camera by increasing distance; selecting a predefined number of the images of the gallery that have the lowest distance; counting the number of occurrences of each candidate unique identifier belonging to said predefined set of unique identifiers among the selected images of the galleries associated with the query images; identifying the target person as the person associated to the candidate unique identifier that has the highest number of occurrences. These additional steps provide accurate results, because the selected ID corresponds to an ID that has been detected for a large number of gallery images similar to the different queryimages. The accuracy is thus increased by both the use of multiple views, and correlation of features with the gallery images.

[0040] The training images are obtained by detecting bounding boxes in a plurality of views of the area captured by the camera, the method further comprising a step of detecting, in said plurality of views captured by said plurality of cameras, bounding boxes each encompassing an image that represents a person, the plurality of query images of a target person being images encompassed in said bounding boxes.

[0041] Said one or more memory further stores, for each camera belonging to said plurality of cameras respectively, the gallery of images associated to this camera, each image of said gallery of images representing a person. Such images can in particular be obtained by detecting bounding boxes in a plurality of views of the area captured by this camera at a plurality of times respectively, each of said plurality of times being distinct of said capture time.BRIEF DESCRIPTION OF DRAWINGS

[0042] Other features, details and advantages will be shown in the following detailed description and on the figures, on which:Fig. 1

[0043] [Fig. 1] is computing system according to an embodiment.Fig.2

[0044] [Fig. 2] is an example of detection of a bounding box according to an embodiment.Fig.3

[0045] [Fig. 3] is an example of an architecture of a supervised machine learning engine according to an embodiment.Fig.4

[0046] [Fig. 4] is an example of a method according to an embodiment.Fig.5

[0047] [Fig. 5] is an example of a method of identification of a person based upon an average confidence score according to an embodiment.Fig.6

[0048] [Fig. 6] is an example of a method of identification of a person based upon a distance according to an embodiment.Fig.7

[0049] [Fig. 7] is an example of a method of identification of a person based upon the number of identifications of persons in gallery images according to an embodiment.Fig.8

[0050] [Fig. 8] is a first example of query images and predictions according to an embodiment.Fig.9

[0051] [Fig. 9] is a second example of query images and predictions according to an embodiment.Fig. 10

[0052] [Fig. 10] is a third example of query images and predictions according to an embodiment.DETAILED DESCRIPTION OF EMBODIMENTS

[0053] It is now referred to figure 1 .

[0054] Figure 1 represents an example of a computing system Sys1 in an embodiment.

[0055] The computing system Sys1 aims at performing person identification using multiple cameras.

[0056] In the example of figure 1 , the system aims at monitoring an area Areal . The area Areal may be an area located in a public or private space. It may for example be a place, a portion of street, a hall of an airport, a hall of a shopping mall, a shop, etc. In order to ease the presentation of figure 1 , a single person Persl to identify is represented. Of course, in real examples a larger number of persons to identify may be found in the monitored area.

[0057] A plurality of cameras (cameras Cam1 .1 , Cam1 .2, Cam1 .3 in the example of figure 1) are placed around the area Areal . The cameras are typically digital cameras that capture a digital video. Each camera Cam1 .1 , Cam1 .2, Cam1 .3 is placed in order to cover a different field of view. Thus, each camera is associated with a different field of view, respectively the fields of view FOV1.1 , FOV1.2, FOV1 .3 for the cameras Cam1.1 , Cam1 .2, Cam1 .3. These cameras are configured to capture the views of the area at the same capture time, i.e perform a synchronized capture. In practice, a view is an image of the area at a given time. Thus, the person Persl to identify is represented at the same time in each of the views. The person Persl may be occluded in one of the fields of view, but not in the others.

[0058] In the remainder of the description, the term “the capture time” will designate the time at which the views of the area are captured to perform the person identification in the method according to the invention. The person identification will thus be based upon query images captured at the capture time, but, as will be explained in more details hereinafter, the identification method will also make use of images captured at times that are different from the capture time.

[0059] The cameras are connected to one or more computing device Compl and are configured to send the captured images and videos to this one or more computing device Compl , so that these one or more computing device Compl can analyze the images and videos to identify the persons in the area. The cameras and computing devices can be connected using different kinds of wired or wireless connections. The connection between the cameras and computing device may for example involve optical fiber, Ethernet cables, or radiofrequency wireless connections. Various technologies such as Wi-Fi, 4G, 5G, etc. may be used for the purpose of the connection between the cameras and computing device Compl .

[0060] In order to ease the reading of figure 1 , a single computing device Compl is represented. However, it is worth noting that a system such as the system Sys1 may comprise a plurality of computing devices, for example to perform distributed computing. The processing modules and data storage represented in figure 1 may therefore be either found in a single computing device, or in a plurality of interconnected computing devices, for example in a server farm.

[0061] A computing device Compl may be any kind of computing devices, for example a personal computer or a server.

[0062] The one or more computing device Compl comprises one or more communication endpoint Comml to receive the image data from the plurality of cameras.

[0063] The one or more computing device Compl comprises one or more memory Mem1 . Although a single memory Mem1 is represented in figure 1 , the one or more computing device Compl may in reality comprise a plurality of memories, stored either in single or a plurality of computing devices.

[0064] The one or more memories store a plurality of galleries of images Gall .1 , Gall .2, Gall .3, and a plurality of pre-trained supervised machine learning engines ML1 .1 , ML1 .2 and ML1 .3. Each camera is associated to one gallery of images, and one pre-trained supervised machine learning engines. Thus, the one or more memory Mem1 store one pre-trained supervised machine learning engines, and one gallery of images for each camera respectively. For example:the camera Cam1.1 is associated to the gallery of images Gall .1 and the pre-trained supervised machine learning engine ML1 .1 ;the camera Cam1.2 is associated to the gallery of images Gall .2 and the pre-trained supervised machine learning engine ML1 .2;the camera Cam1.3 is associated to the gallery of images Gall .3 and the pre-trained supervised machine learning engine ML1 .3;etc.

[0065] In the remainder of the disclosure, an image of a gallery can be referred to as “a gallery image”.

[0066] A gallery of images associated to a camera comprises a plurality of gallery images. Each gallery image represents a person and has been obtained by detecting bounding boxes in a plurality of views of the area captured by the camera at a plurality of times distinct from the capture time.

[0067] A pre-trained supervised machine learning engine associated to a camera is a supervised machine learning engine previously trained to identify persons using a labelled training dataset comprising training images that represent a person, said training images being obtained by detecting bounding boxes in a plurality of views of the area captured by the camera. In the labelled training dataset, each person to classify is typically represented by a unique identifier that belongs to a predefined set of unique identifiers (of persons to classify), and the images are labelled with an identifier that unambiguously designates the person represented in the image. The term “unique identifier” indicates that each person to detect is identified by a single identifier. Of course, there are a plurality of unique identifiers associated respectively with a plurality of different persons. In order to take into account that some of the persons will not be associated with a unique identifier, a unique identifier representative of “other / unknown persons” may also be used.

[0068] Stated otherwise, the supervised machine learning engine associated to a given camera is trained using a labelled training dataset comprising only images obtained from that camera. The labelled training dataset can be for example obtained by :receiving views from the camera at different times;detecting bounding boxes in each view, each bounding box being a rectangle that encompasses an image of a person;forming the training dataset by associating each image of a person to a label. The label corresponds to an identifier of a given person. For example, each person to detect can be associated to a label, for example an index of the person. The labels may also comprise a label representative of unknown or other persons (e.g all the persons that are not to be detected). The labels can be for example entered by a user, and / or provided by another pre-trained model.

[0069] Test and validation dataset can be obtained in the same way for the supervised machine learning engine.

[0070] The supervised machine learning engine can thus be trained to classify the images into the labels, i.e identify the persons represented in the images, since the labels correspond to the persons to identify. The classification also involves the determination of features of the images. Each camera is thus associated to a supervised machine learning engines that is trained to classify persons from views captured by the camera.

[0071] It is worth noting that, even if the different supervised machine learning engines are trained using images from different cameras, the classes used for the training are the same. Stated otherwise, a same label or identifier correspond to a same person for each of the supervised machine learning engines ML1 .1 , ML1 .2, ML1.3, etc. Thus, the classifications provided by the different cameras can be fused, as will be explained in more details below.

[0072] It is now referred to figure 2.

[0073] Figure 2 is an example of detection of a bounding box according to an embodiment.

[0074] In the example of figure 2, a view Vw2 of an area has been captured by a camera.

[0075] A plurality of persons are represented in the view. A detection of bounding boxes can be performed to detect the persons in the view. Each bounding box is detected as a rectangle that encompasses a person in the view. For example, the bounding box Bbox2 is a rectangle that encompasses a person in the view. The bounding boxes can be detected in different ways. For example, one of the methods described by Chavdarova, T., Baque, P., Bouquet, S., Maksai, A., Jose, C., Bagautdinov, T., ... & Fleuret, F. (2018). Wildtrack: A multicamera hd dataset for dense unscripted pedestrian detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5030-5039).

[0076] Each bounding box defines an image that represents a person. For example, the image Img2 represents a person. The image that represents a person Img2 is therefore a subset of the view Vw2 encompassed by the bounding box Bbox2.

[0077] As will be explained in more details below, the “images that represent a person” can be used:In the training phase, to train a supervised machine learning engine. In this case, they are associated with a label that represents the person in the image. For example, the image Img2 is associated with a label indicating the person that is represented, for example an index of the person;In the inference phase, to determine the persons that are represented in a view. In this case, the image that represents a person is fed to the trained machine learning engines to determine which is the person represented in the image.

[0078] It is now referred to figure 3.

[0079] Figure 3 is an example of an architecture of a supervised machine learning engine according to an embodiment.

[0080] The pre-trained supervised machine learning engines (such as for example the supervised machine learning engines ML1.1 , ML1.2, ML1.3) are image classifiers, in the sense that they are able to classify an image in a class, i.e attribute a label to an image. More specifically, they are trained to identify persons, i.e associated to an input image an identifier that is representative of a person represented in the image. An identifier may for example be a number representative of a person. For example, the persons to identify may be identified by identifiers labeled as ID 0001 , ID 0002, ID 0003, etc. or more generally using any suitable identifier. The supervised machine learning engines are furthermore configured, in order to perform the classification, to extract features of the image.

[0081] In some embodiments of the invention, the supervised machine learning may be further configured to provide a classification confidence to the classification.

[0082] It is thus apparent that any supervised machine learning engine capable to perform image classification and extract features from an image may be used in an embodiment. In particular, supervised machine learning engines that rely on convolutional neural networks (CNNs) are well suited for performing image classification and features extraction.

[0083] Figure 3 presents, by means of non limitative example only, a classifier called “MGN” (standing for Multiple Granularity Network), which was presented by Wang, G., Yuan, Y., Chen, X., Li, J., & Zhou, X. (2018, October). Learning discriminative features with multiple granularities for person re-identification. In Proceedings of the 26th ACM international conference on Multimedia (pp. 274-282) and may be used in an embodiment. The architecture of this classifier is briefly presented with reference to figure 3. This architecture uses CNNs, is especially well suited to identify pedestrians and may thus be used within the context of the present disclosure.

[0084] This MGN framework focuses on improving person re-identification by integrating global and local feature learning. It employs a multi-branch deep network architecture, consisting of one global branch and two local branches, to capture discriminative features at various granularities. The key aspects of this MGN framework are:• Global and Local Features: The network combines global feature representations with local features obtained by partitioning images into stripes, enhancing the ability to capture fine details.• End-to-End Learning: MGN is designed as an end-to-end learning process, simplifying implementation and improving robustness.• Loss Functions: The framework uses softmax and triplet loss functions to enhance classification and metric learning, improving ranking performance.• State-of-the-Art Performance: MGN achieves superior results on several mainstream person reidentification datasets, demonstrating its effectiveness in handling challenges like occlusion and pose variation.

[0085] The MGN architecture is spit into 3 branches (namely Global branch, Part-2 branch and Part-3 branch), takes as input an input image Inlmg3 that represents a person (in the image, a plurality of input images are represented, and corresponds to a mini-batch of images for training, where the MGN architecture takes each of the images as input to calculate the loss function and updates the parameters of the neural network), and comprises the following elements:• The first layers ResNet3 of a residual neural network. For example, the first layers of the ResNet50 residual neural network (described by He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 770-778).) may be used. The output of the first layers Res3 is then forwarded independently to three branches BrGlob3, BrPRT23 and BrPrt33;• The Global Branch BrGlob3 captures global features of the entire image, focusing on the overall appearance of the person. The features are first extracted from the intermediary layers IntLayGlobS of the residual network;• The Part-2 Branch BrPrt23: divides the image into two horizontal stripes, extracting local features from each part to capture finer details. The features are first extracted based upon a division of the image lntLayPart23 into 2 slices;• The Part-3 Branch BrPrt33: further divides the image into three horizontal stripes, capturing even finer local features. The features are first extracted based upon a division of the image lntLayPart33 into 3 slices.

[0086] Then, global max pooling operations GbMaxPool3, and 1x1 convolutional reduction ConvRed3 are applied. In order to train the supervised machine learning engine, different losses are computed. To this effect, the figure 3 represents different types of arrows that represent the data flow within the architecture of the supervised machine learning engine. In particular:Black arrows represent a direct forward of data;Dark gray arrows represent a Softwax_256 loss;Dashed dark gray arrows represent a Softwax_2048 loss;Light gray arrow represent a triplet loss.

[0087] The loss functions are used during the training phase to adjust the weights of the layers of the residual neural network.

[0088] During the inference phase, the supervised machine learning engine is pre-trained, and only the features and classification are determined. The features obtained for all the branches are used in inference.

[0089] The architecture effectively combines global and local features, improving the network's ability to identify individuals across different views and conditions.

[0090] It is now referred to figure 4.

[0091] Figure 4 represents an example of a method P4 implemented by one or more processing unit, such as for example the processing unit Prod , of a computing system, for example the system Sys1 , according to an embodiment.

[0092] The method P4 comprises a first step S41 of receiving a plurality of views of a same area captured by a plurality of cameras of the system respectively, at a same capture time.

[0093] In the example of figure 1 , the step S41 consists in receiving 3 views captured respectively by the 3 cameras Cam1 .1 , Card .2, Card .3 at a same capture time. Each camera captures a view of the area Areal from its field of view in a synchronized manner, so that the views represent the same moment in time.

[0094] The method P4 further comprises a step S42 of detecting bounding box(es) in said plurality of views, each bounding box encompassing an image that represents a person.

[0095] As shown in figure 2, the detection of bounding boxes involves identifying human figures within the views, and may be performed for example using techniques such as background subtraction or motion detection. The bounding boxes are extracted as rectangular images that contain the image of a person. At the output of step S42, the human beings who are represented in the views are extracted as independent rectangular images, thereby allowing for further analysis and identification.

[0096] The method P4 further comprises a step S43 of obtaining a plurality of query images of a target person, said plurality of query images comprising for each view a query image associated to the view, said query image associated to the view being one of the images encompassed in the bounding boxes detected in the view.

[0097] The step S43 corresponds in practice to a determination, among the images that represent a person encompassed within the bounding boxes detected for each view, of the images that represent the same person. Typically, in the example of figure 1 , the step S43 would lead to the obtention of a plurality of query images that contain:a first query image of the person Persl that was encompassed in one of the bounding boxes detected in the view captured by the camera Cam1 .1 at the capture time;a second query image of the person Persl that was encompassed in one of the bounding boxes detected in the view captured by the camera Cam1 .2 at the capture time;A third query image of the person Persl that was encompassed in one of the bounding boxes detected in the view captured by the camera Cam1 .3 at the capture time.

[0098] Thus, at the output of the step S43, one query image is obtained for each of the views of the area captured by each of the cameras respectively, all the query images representing the same person. The plurality of query images thus show the same person from different angles of view.

[0099] The images that represent the same person can be determined in different ways. For example, the paper Hou, Y., Zheng, L., & Gould, S. (2020). Multiview detection with feature perspective transformation. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VII 16 (pp. 1-18). Springer International Publishing, introduces the innovative MultiView dataset and proposes a method that projects all views onto the ground plane to create a bird's-eye view. Utilizing this perspective, the position of an individual in the ground plane can be determined, which then allows for the calculation of their position in each respective view by applying the parameters of the camera.

[0100] The next steps S44 and S45 are performed, sequentially or in parallel, for each of the query images of the plurality of query images. As a reminder, each of the query images is associated to a view captured by a camera. Thus, the steps S44 and S45 are performed for each combination of a query image, a view and a camera.

[0101] In the fourth step S44, the supervised machine learning engine associated to the camera (for example, in fig. 1 , the supervised machine learning engine ML1 .1 if the query image is extracted from the view captured by camera Cam1 .1 , the supervised machine learning engine ML1 .2 if the query image is extracted from the view captured by camera Cam1.2, the supervised machine learning engine ML1.3 if the query image is extracted from the view captured by camera Cam1 .3) takes as input the query image to determine a vector of features associated with the query image.

[0102] Additionally, the supervised machine learning engine takes as input each image of the gallery associated with the camera (e.g, in the example of figure 1 , each image of the gallery Gall .1 for the supervised machine learning engine ML1 .1 , each image of the gallery Gall .2 for the supervised machine learning engine ML1 .2, each image of the gallery Gall .3 for the supervised machine learning engine ML1 .3), to determine:a vector of features associated with the image of the gallery and ;an identifier, among the predefined set of identifiers, of a person represented by the image of the gallery.

[0103] Stated otherwise, at the output of step S44, the trained supervised machine learning engine is allowed to obtain:The features of the query image;The features, and classification of each of the images of the gallery associated with the same view / camera than the query image.

[0104] These features may be organized in vectors of features which may be represented in different ways. For example, the vector of features associated with the query image may be represented as a standalone vector, while the vectors of features associated with the images of the gallery may be grouped within a matrix of features.

[0105] It is worth noting that the order of the steps represented in the figure 4 is not limitative. In particular, the order of execution of some of the steps may be modified when it does not modify the output of the method. In particular, the step S44 may be executed at once to calculate in the same time the vectors of features associated with the query image, and the image of the gallery, or in two separate sub-steps to calculate the vector of features associated with the query image on one hand, and the vectors of features associated with the images of the gallery and the identifiers of the person represented by the images of the gallery on the other hand. For example, the vectors of features associated with the images of the gallery and the identifiers of the person represented by the images of the gallery may be calculated during a preliminary step, for example before step S41 .

[0106] The method P4 then comprises, for each query image, a further step S45 of calculating a plurality of distances between the vector of features associated with the query image and the vectors of features associated with the images of the gallery.

[0107] Stated otherwise, the step S45 consists in calculating, for each image of the gallery associated with the camera, a distance between the vector of features associated with the image and the vector of features associated with the query image. The distance may be for example a cosine distance.

[0108] Thus, at the output of step S45, a plurality distances are obtained, between the query image and each of the images of the gallery respectively. The plurality of distances provides an indication of the level of similarity between the query image and each of the images of the gallery respectively. The lower the distance between the vectors of features representative of the image of the gallery and the query image (and thus the more similar an image from the gallery is from the query image), the more likely the image of the gallery represents the same image than the query image.

[0109] Then, in a further step S46, it is verified whether there is a subsequent query image to process. If there are more query images to process, the method loops back to the previous steps S44 and S45 to process the next image.

[0110] Once the distances are calculated for all the query images of the plurality of query images, the method comprises a further step S47 of identifying the target person based on said plurality of distances obtained for each query image.

[0111] As explained before, at the output of step S45, the distances are known between each query image and each image of the associated gallery. Furthermore, as the predictions of the persons represented in each image of the gallery are also known, it is possible provide an accurate estimation of the person represented by the query image as there is a high likelihood that the person represented in the query image is a person who is predicted for gallery images that have a low distance with the query image. The accuracy is further enhanced by using a model trained specifically for each camera.

[0112] The method allows an accurate prediction of the person represented by the query images even in case of occlusion, because the determination of the person is done based upon a plurality of views, and because the use of a model specifically trained for each view / camera allows correlating the features of query images and query images even if the person is partially occluded in some of the images.

[0113] Now a plurality of examples of determination of a person will be presented in a plurality of embodiments.

[0114] It is now referred to figure 5.

[0115] Figure 5 is an example of a method of identification of a person based upon an average confidence score according to an embodiment.

[0116] The figure 5 represents an example of embodiment of step S47, that relies on the calculation of confidence scores that each query image represent each candidate person, then an averaging of the confidences for all the query images.

[0117] In the example of figure 5, the first three steps S51 to S53 allow obtaining a confidence score that the plurality of query images represent a person that is identified by given candidate unique identifier that belongs to the predefined set of unique identifiers.

[0118] To this effect, a plurality of confidence score, that each of the plurality of images respectively represents the person associated to said candidate unique identifier, are calculated.

[0119] To this effect, a first step S51 consists in determining a confidence score that a query image associated with a view captured by a camera represents the person associated to a candidate unique identifier. The determination of the confidence score that a query image represents the person associated with a given candidate identifier is determined based on its distances with the images of the gallery associated to said camera, and the unique identifiers of the persons represented by the images of the gallery as determined by the supervised machine learning engine. Stated otherwise, each of the supervised machine learning engine has determined for each image of the gallery a unique identifier of the person represented by the image of the gallery. A short distance between the query image and an image of the gallery that represents the person associated with a given candidate identifier provides an insight that the query image likely represents that person associated with the given candidate identifier. Conversely, a long distance between the query image and an image of the gallery that represents the person associated with a given candidate identifier provides an insight that the query image does not likely represent that person associated with the given candidateidentifier. A confidence score that the query image represents a given candidate identifier can thus be determined based upon the distances between the query image and the gallery image, and the identifiers of the persons detected by the supervised machine learning engine in the gallery images.

[0120] In substance, this confidence score can be calculated based upon the distances with the images of the gallery and the unique identifiers of the persons represented by the images of the gallery, because the lower the distance between a query image and a gallery image that represents a person identified by the unique identifier, the higher the confidence that the query image also represents the person identified by the unique identifier.

[0121] The confidence score may for example be one of the following metrics:a “rank-i” metrics that refers to the probability of correctly identifying the target person within the top i positions of the retrieval results (e.g the probability that the unique identifier of the target person is found the i images of the gallery that have the lowest distance with the query image). For example, rank-1 indicates whether the gallery image that is the closest to the query image represents the target person. This metric intuitively reflects the accuracy of the retrieval system in returning the most likely result;a “mAP” metrics, that is a more comprehensive evaluation metric that takes into account the ranking of all relevant results. Specifically, mAP involves calculating the Average Precision (AP) for each query image, and then averaging the AP values across all query images. AP represents the area under the Precision-Recall (P-R) curve for a query in information retrieval, where precision is the proportion of true positive samples in the retrieval results (e.g in the gallery images that have a close distance to the query image), and recall is the proportion of all true positive samples that are retrieved. mAP provides an assessment of the overall performance of the retrieval system;one of the metrics defined by: Ye, M., Shen, J., Lin, G., Xiang, T., Shao, L., & Hoi, S. C. (2021 ). Deep learning for person re-identification: A survey and outlook. IEEE transactions on pattern analysis and machine intelligence, 44(6), 2872-2893.

[0122] At step S52, it is verified if there is a subsequent query image to process. If yes, the step S51 is executed for the subsequent query image. Otherwise (i.e if a confidence score that the query image represents the person associated to said candidate unique identifier has been calculated for each of the plurality of the query image for a given unique identifier), the step S53 is executed.

[0123] The step S53 consists in obtaining an average confidence score that the plurality of query images represents the person associated to the candidate unique identifier by averaging the confidence scores for each image of said plurality of query images.

[0124] Stated otherwise, a confidence score that the plurality of query images represents, as a whole, a person associated to the unique identifier is obtained by averaging the confidence scores that each image of the plurality represent the person associated to the unique identifier.

[0125] Thus, at the output of steps S51 to S53, an average confidence score that the plurality if images represents the person associated with the unique identifier is obtained.

[0126] Then, at step S54, it is checked whether there are subsequent candidate unique identifiers to evaluate. If yes, the process loops back to the steps S51 to S53 with the next candidate unique identifier.

[0127] Otherwise (i.e if all the unique identifiers have been processed, and thus if an average confidence score that the plurality of images represent a person associated with a unique identifier has been calculated for each unique identifier), the step S55 is executed.

[0128] The step S55 identifies the target person as the one associated with the unique identifier for which the highest average confidence score has been obtained. This step S55 thus concludes the identification process by selecting the most likely match based on the average confidence scores.

[0129] The method steps represented in figure 5 thus provides an accurate identification of the person, because the person associated with the highest confidence among all the images of the plurality of images is identified.

[0130] It is now referred to figure 6.

[0131] Figure 6 is an example of a method of identification of a person based upon a distance according to an embodiment.

[0132] More specifically, figure 6 represents an embodiment of the step S47, wherein the identification of the target person is based on distances. The identification method represented in figure 6 will thus be called “by distance”.

[0133] Firstly, a loop of steps S61 to S63 is performed for each query image associated with a view captured by a camera belonging to the plurality of query images.

[0134] In step S61 , the images of the gallery associated to the camera are sorted by increasing distance. More specifically, the distances calculated at step S45, between the vector of features associated with the query image and the vectors of features associated with the images of the gallery, are used to sort the gallery images.

[0135] At the output of S61 , the query image is thus associated with a list of gallery images sorted by increasing distance, where each of the gallery image is also associated with a unique identifier of a person (calculated at step S44).

[0136] In an illustrative example, where two views (View 1 and View 2 corresponding to images of the same area captured by two distinct cameras respectively) are considered, the output of step S61 may be:For View 1 :o A list of distances [0.3, 0.35, 0.4, 0.45, ...] paired witho A list of unique identifiers (IDs) [1 , 2, 5, 2, ...].For View 2:o A list of distances [0.32, 0.33, 0.46, 0.57, ...] paired witho A list of unique IDs [2, 1 , 1 , 4, ...].

[0137] In step S62, a reference distance is determined for the query image as the lowest distance of order m, where m is a predefined integer.

[0138] The reference distance of order m, or “Rank-m distance”, is the distance of the m-th gallery image in the list. For example, in the illustrative example above:the Rank-1 distance for View 1 is 0,3, corresponding to the unique ID 1 ;the Rank-1 distance for View 2 is 0,32, corresponding to the unique ID 2;the Rank-2 distance for View 1 is 0,35, corresponding to the unique ID 2;the Rank-2 distance for View 2 is 0,33, corresponding to the unique ID 1 ;etc.

[0139] At step S63, it verified if there is a subsequent query image to process. If yes, the steps S61 and S62 are executed for the subsequent query image. Otherwise (i.e if a reference distance of order m has been calculated for each of the query image), the step S64 is executed.

[0140] The step S64 identifies the target person as the person identified in the image of the gallery having the lowest reference distance of order m. Stated otherwise, the lowest distance of order m is selected among the distances calculated in step S62, then the target person is identified as the person identified in the gallery image corresponding to the lowest reference distance of order m.

[0141] For example, in the illustrative example provided above:if a reference distance of order 1 is considered, the person will be identified as ID 1 , because this corresponds to the lowest distance of order 1 (distance 0.3 with view 1 , lower than distance 0.32 with view 2);if a reference distance of order 2 is considered, the person will be identified as ID 1 , because this corresponds to the lowest distance of order 2 (distance 0.33 with view 2);if a reference distance of order 3 is considered, the person will be identified as ID 5, because this corresponds to the lowest distance of order 3 (distance 0.4 with view 1 ) ;etc.

[0142] This identification method provides accurate results, because the identified person will be a person represented in a gallery image having a lowest distance with one of the query images.

[0143] It is now referred to figure 7.

[0144] Figure 7 is an example of a method of identification of a person based upon the number of identifications of persons in gallery images according to an embodiment. The identification method represented in figure 7 will thus be called “by vote”.

[0145] More specifically, figure 7 represents an embodiment of the step S47, wherein the identification of the target person is based on distances.

[0146] Firstly, a loop of steps S71 to S73 is performed for each query image associated with a view captured by a camera belonging to the plurality of query images.

[0147] In step S71 , the images of the gallery associated to the camera are sorted by increasing distance. More specifically, the distances calculated at step S45, between the vector of features associated with the query image and the vectors of features associated with the images of the gallery, are used to sort the gallery images.

[0148] At the output of S71 , the query image is thus associated with a list of gallery images sorted by increasing distance, where each of the gallery image is also associated with a unique identifier of a person (calculated at step S44).

[0149] In step S72, a predefined number of the closest gallery images are selected from the sorted list. For example, the 2nd, 3rd, 5th, 10th, etc. closest gallery images may be selected. This step ensures that only the most relevant images are considered for identification.

[0150] Taking the same illustrative example as in figure 6, selecting the 3 closest images would result in selecting :For view 1 (query image captured by a first camera), two gallery images respectively associated with the unique IDs, 1 , 2 and 5 (respectively corresponding to the distances 0.3, 0.35 and 0.4);For view 2 (query image captured by a second camera), two gallery images respectively associated with the unique IDs, 2, 1 and 1 (respectively corresponding to the distances 0.32, 0.33 and 0.46).

[0151] Therefore, at the output of step S72, the query image that has been processed is associated with a predefined number of selected closest gallery images, each associated with an identifier of a person that has been detected in the gallery image by the supervised machine learning engine associated with the same camera as the gallery and the query image.

[0152] At step S73, it verified if there is a subsequent query image to process. If yes, the steps S71 and S72 are executed for the subsequent query image. Otherwise (i.e if a reference distance of order m has been calculated for each of the query image), the step S74 is executed.

[0153] At step S74, the number of occurrences of each candidate unique identifier among the selected images is determined. Stated otherwise, all the gallery images selected at step S72 for each query images are considered, and the number of times each unique identifier is detected is one of the gallery images is counted. It is worth noting that the numbers of occurrences counted at step S74 are the number of occurrences for the gallery images selected for all the query images. The number of occurrences of each unique identifier are thus counted over a number of gallery images equal to the number of query images multiplied by the predefined number of gallery images (i.e the number of selected gallery images for each query image).

[0154] In the same illustrative example of figure 6, using two query images (for view 1 and view 2 respectively) and a predefined number of 3 selected gallery images for each query image, there are 6 (2*3) gallery images to consider, and the step S74 would count, for view 1 and view 2:3 occurrences of ID 1 (1 gallery image for view 1 , 2 gallery images for view 2);2 occurrences of ID 2 (1 gallery image for view 1 , 1 gallery image for view 2);1 occurrence of ID 5 (1 gallery image for view 1 ).

[0155] The embodiment represented in figure 7 further comprises a step S75 of identifying the target person as the one associated with the candidate unique identifier that has the highest number of occurrences. In the illustrative example, this would therefore result in selecting the person associated with the ID 1 , as it has the most occurrences.

[0156] The embodiment represented in figure 7 provides accurate results, because the selected ID corresponds to an ID that has been detected for a large number of gallery images similar to the different query images. The accuracy is thus increased by both the use of multiple views, and correlation of features with the gallery images.

[0157] It is now referred to figures 8, 9 and 10.

[0158] Each of the figures 8, 9 and 10 represents an example of a plurality of query images of a target person, and the corresponding image of gallery that has the lowest distance to the query image (which will be referred to as “Prediction Rank 0”). These three figures are organized in the same way:on the upper part of the figure, the plurality of query images. In each case, the plurality of query images corresponds to images of a same target person captured at a same capture time for different views, i.e by a plurality of different cameras having a different field of view of the same area where the target person is located at the capture time. The ID of the target person is represented under each query image. As the query image represents the same person, the same ID is found under each query image; on the lower part of the figure, the corresponding predictions rank 0, which are the gallery images that are the closest to the query image for each view. Under each gallery image is represented the ID of the person identified in the gallery image, as well as the distance between the closest gallery image and the query image.

[0159] In the example of figure 8, four views are considered:For view 1 , the query image Quer8.1 has been captured by a first camera, and represents the target person associated with the ID 0046. Among the gallery images associated with the first camera I view 1 , the closest gallery image is the gallery image Gal8.1 , wherein the person having the ID 0046 has been identified. Thus, the Rank 0 prediction is correct for view 1. The gallery image Gal8.1 and the query image Quer8.1 have a distance equal to 0.225. Stated otherwise, the supervised machine learning engine associated with the first camera / view 1 has classified the gallery image Gal8.1 into ID 0046 at step S44, and a distance of 0.225 between the vector of features of the query image Quer8.1 and the gallery image Gal8.1 has been calculated at step S45;For view 2, the query image Quer8.2 has been captured by a second camera, and also represents the target person associated with the ID 0046. Among the gallery images associated with the second camera I view 2, the closest gallery image is the gallery image Gal8.2, wherein the person having the ID 0046 has been identified. Thus, the Rank 0 prediction is correct for view 2. The gallery image Gal8.2 and the query image Quer8.2 have a distance equal to 0.194;For view 3, the query image Quer8.3 has been captured by a third camera, and also represents the target person associated with the ID 0046. Among the gallery images associated with the third camera I view 3, the closest gallery image is the gallery image Gal8.3, wherein the person having the ID 0019 has been identified. Thus, the Rank 0 prediction is incorrect for view 3. The gallery image Gal8.3 and the query image Quer8.3 have a distance equal to 0.415;For view 4, the query image Quer8.4 has been captured by a fourth camera, and also represents the target person associated with the ID 0046. Among the gallery images associated with the fourth camera I view 4, the closest gallery image is the gallery image Gal8.4, wherein the person having the ID 0047 has been identified. Thus, the Rank 0 prediction is incorrect for view 4. The gallery image Gal8.4 and the query image Quer8.4 have a distance equal to 0.512.

[0160] In the example of figure 8, based on the query images, the identification method “by distance” presented in figure 6 would correctly identify the target person as the one having ID 0046, because the lowest distance is the distance between the query image Quer8.2 and gallery image Gal8.2 for view 2 (0.194), and the personhaving the ID 0046 has been identified in the gallery image Gal8.2 by the supervised machine learning engine associated with the second camera.

[0161] In the example of figure 8, based on the query images, the identification method “by vote” presented in figure 7 using a single gallery image (predefined number of the images of the gallery equal to 1) would also correctly identify the target person as the one having ID 0046, because ID 0046 has the highest number of occurrences of identifications among the gallery images Gal8.1 , Gal8.2, Gal8.3 and Gal8.4 (2 occurrences, versus a single occurrence for ID 0019 and ID 0047).

[0162] In the example of the figure 8, the first and second cameras produced, for views 1 and 2 respectively, images that were both neat and unobstructed, thereby leading to a correct identification and low distances between query images and gallery images representing the right target person. On the contrary, the third camera produced occluded images for view 3, and the fourth camera produced blurred images for view 4, thereby leading to both incorrect predictions and high distances between query images and gallery images.

[0163] The example of figure 8 therefore demonstrates how the invention allows obtaining overall accurate identifications, even when some of the views are impaired by low quality or obstructed images, thereby leading to incorrect identifications.

[0164] In the example of figure 9, five views are considered:For view 1 , the query image Quer9.1 has been captured by a first camera, and represents the target person associated with the ID 0048. Among the gallery images associated with the first camera I view 1 , the closest gallery image is the gallery image Gal9.1 , wherein the person having the ID 0048 has been identified. Thus, the Rank 0 prediction is correct for view 1. The gallery image Gal9.1 and the query image Quer9.1 have a distance equal to 0.189.For view 2, the query image Quer9.2 has been captured by a second camera, and also represents the target person associated with the ID 0048. Among the gallery images associated with the second camera I view 2, the closest gallery image is the gallery image Gal9.2, wherein the person having the ID 0048 has been identified. Thus, the Rank 0 prediction is correct for view 2. The gallery image Gal9.2 and the query image Quer9.2 have a distance equal to 0.518;For view 3, the query image Quer9.3 has been captured by a third camera, and also represents the target person associated with the ID 0048. Among the gallery images associated with the third camera I view 3, the closest gallery image is the gallery image Gal9.3, wherein the person having the ID 0042 has been identified. Thus, the Rank 0 prediction is incorrect for view 3. The gallery image Gal9.3 and the query image Quer9.3 have a distance equal to 0.620;For view 4, the query image Quer9.4 has been captured by a fourth camera, and also represents the target person associated with the ID 0048. Among the gallery images associated with the fourth camera I view 4, the closest gallery image is the gallery image Gal9.4, wherein the person having the ID 0048 has been identified. Thus, the Rank 0 prediction is correct for view 4. The gallery image Gal9.4 and the query image Quer9.4 have a distance equal to 0.287;For view 5, the query image Quer9.5 has been captured by a fifth camera, and also represents the target person associated with the ID 0048. Among the gallery images associated with the fifth camera I view 5, the closest gallery image is the gallery image Gal9.5, wherein the person having the ID 0048has been identified. Thus, the Rank 0 prediction is correct for view 5. The gallery image Gal9.5 and the query image Quer9.5 have a distance equal to 0.311.

[0165] In the example of figure 9, based on the query images, the identification method “by distance” presented in figure 6 would correctly identify the target person as the one having ID 0048, because the lowest distance is the distance between the query image Quer9.1 and gallery image Gal9.2 for view 1 (0.189), and the person having the ID 0048 has been identified in the gallery image Gal9.1 by the supervised machine learning engine associated with the second camera.

[0166] In the example of figure 9, based on the query images, the identification method “by vote” presented in figure 7 using a single gallery image (predefined number of the images of the gallery equal to 1) would also correctly identify the target person as the one having ID 0048, because ID 0048 has the highest number of occurrences of identifications among the gallery images Gal9.1 , Gal9.2, Gal9.3, Gal9.4 and Gal9.5 (4 occurrences, versus a single occurrence for ID 0042).

[0167] In the example of the figure 9, the first, second and fourth cameras produced for views 1 , 2 and 4 respectively images that were both neat and unobstructed, thereby leading to a correct identification between query images and gallery images representing the right target person. The distance was also low for view 1 , wherein the query image Quer9.1 and gallery image Gal9.1 are very similar. The fifth camera produced partially occluded images for view 5, wherein the target person was nevertheless correctly identified. On the contrary, the third camera produced even more occluded images for view 3, thereby leading to both incorrect predictions and high distances between query images and gallery images.

[0168] The example of figure 9 therefore demonstrates how the invention allows obtaining overall accurate identifications, even when some of the views are impaired by low quality or obstructed images, thereby leading to incorrect identifications.

[0169] In the example of figure 10, six views are considered:For view 1 , the query image Querl 0.1 has been captured by a first camera and represents the target person associated with the ID 0054. Among the gallery images associated with the first camera I view 1 , the closest gallery image is the gallery image Gall 0.1 , wherein the person having the ID 0054 has been identified. Thus, the Rank 0 prediction is correct for view 1 . The gallery image Gall 0.1 and the query image Querl 0.1 have a distance equal to 0.126.For view 2, the query image Querl 0.2 has been captured by a second camera, and also represents the target person associated with the ID 0054. Among the gallery images associated with the second camera I view 2, the closest gallery image is the gallery image Gall 0.2, wherein the person having the ID 0054 has been identified. Thus, the Rank 0 prediction is correct for view 2. The gallery image Gall 0.2 and the query image Querl 0.2 have a distance equal to 0.554;For view 3, the query image Querl 0.3 has been captured by a third camera, and also represents the target person associated with the ID 0054. Among the gallery images associated with the third camera / view 3, the closest gallery image is the gallery image Gall 0.3, wherein the person having the ID 0054 has been identified. Thus, the Rank 0 prediction is correct for view 3. The gallery image Gall 0.3 and the query image Querl 0.3 have a distance equal to 0.404;For view 4, the query image Querl 0.4 has been captured by a fourth camera, and also represents the target person associated with the ID 0054. Among the gallery images associated with the fourthcamera I view 4, the closest gallery image is the gallery image Gall 0.4, wherein the person having the ID 0055 has been identified. Thus, the Rank 0 prediction is incorrect for view 4. The gallery image Gall 0.4 and the query image Querl 0.4 have a distance equal to 0.332;For view 5, the query image Querl 0.5 has been captured by a fifth camera, and also represents the target person associated with the ID 0054. Among the gallery images associated with the fifth camera / view 5, the closest gallery image is the gallery image Gall 0.5, wherein the person having the ID 0054 has been identified. Thus, the Rank 0 prediction is correct for view 5. The gallery image Gall 0.5 and the query image Querl 0.5 have a distance equal to 0.298;For view 6, the query image Querl 0.6 has been captured by a sixth camera, and also represents the target person associated with the ID 0054. Among the gallery images associated with the sixth camera / view 6, the closest gallery image is the gallery image Gall 0.6, wherein the person having the ID 0054 has been identified. Thus, the Rank 0 prediction is correct for view 6. The gallery image Gall 0.6 and the query image Querl 0.5 have a distance equal to 0.153.

[0170] In the example of figure 10, based on the query images, the identification method “by distance” presented in figure 6 would correctly identify the target person as the one having ID 0054, because the lowest distance is the distance between the query image Querl 0.1 and gallery image Gall 0.2 for view 1 (0.126), and the person having the ID 0054 has been identified in the gallery image Gall 0.1 by the supervised machine learning engine associated with the second camera.

[0171] In the example of figure 10, based on the query images, the identification method “by vote” presented in figure 7 using a single gallery image (predefined number of the images of the gallery equal to 1 ) would also correctly identify the target person as the one having ID 0054, because ID 0054 has the highest number of occurrences of identifications among the gallery images Gall 0.1 , Gall 0.2, Gall 0.3, Gall 0.4, Gall 0.5 and Gall 0.6 (5 occurrences, versus a single occurrence for ID 0055).

[0172] In the example of figure 10, the first, fifth and sixth cameras produced, for views 1 , 5, and 6 respectively, images that were both neat and unobstructed, thereby leading to a correct identification between query images and gallery images representing the right target person, and low distances. The second and third cameras produced, for views 2 and 3 respectively, low contrast I slightly occluded images, wherein the target person was nevertheless correctly identified but with a high distance. On the contrary, the fourth camera produced an image for view 4 wherein the target person was visible on the background, and largely occluded by two other pedestrians, thereby leading to both incorrect predictions despite a moderate distance between query imageQuer10.4 and gallery image Gall 0.4.

[0173] The example of figure 10 therefore demonstrates how the invention allows obtaining overall accurate identifications, even when some of the views are impaired by low quality or obstructed images, thereby leading to incorrect identifications. In particular, the correct predictions in most of the view caused a correct overall identification in the “by vote” method, while the very neat perspective in some of the views (views 1 and 6 in the example of figure 10) led to correct identifications and low distances, thereby leading to an overall correct identification in the “by distance” identification method.

[0174] This disclosure is not limited to the method, computer program and device described here, which are only examples. The invention encompasses every alternative that a person skilled in the art would envisage when reading this text.

Claims

CLAIMS1. A method (P4) implemented by one or more processing unit (Prod) of a computing system (Sys1), said computing system comprising a plurality of cameras (Cam1 .1 , Card .2, Card .3) having different fields of views (FOV1.1 , FOV 1.2, FOV 1.3) of a same area (Areal), each camera of said plurality being respectively configured to capture a view of the area at a same capture time;said method comprising:obtaining (S43) a plurality of query images of a target person (Persl ) from a plurality of views captured by said plurality of cameras respectively at the same capture time;for each query image of said plurality of query images, calculating (S45) a plurality of distances between a vector of features associated with the query image and vectors of features associated with images of a gallery of images representing a person in a plurality of views of the area captured respectively at a plurality of times, distinct of said capture time, by the camera which has captured the view from which the query image is obtained; andidentifying (S47) said target person based on said plurality of distances calculated for each query image.

2. The method of claim 1 , wherein said computing system further comprises at least one memory (Mem1) which stores, for at least one camera of said plurality of cameras, a machine learning engine (ML1 .1 , ML1 .2, ML1.3) associated with said camera, the method further comprises using (S44) the machine learning engine associated to said camera to determine the vector of features associated with the query image obtained from the view captured by said camera and, for each image of the gallery of images associated with said camera, the vector of features associated with the image of the gallery.

3. The method of claim 2, wherein said machine learning engine comprises a convolutional neural network.

4. The method of claim 2 or 3, wherein the method further comprises using (S44) the machine learning engine associated to said camera to determine, for each image of the gallery of images associated with said camera, a unique identifier, among a predefined set of unique identifiers, of a person represented by the image of the gallery, the identification (S75) of said target person being further based on said unique identifiers.

5. The method of claim 4, wherein identifying said target person based on said plurality of distances for each query image comprises:for each candidate unique identifier belonging to said predefined set of unique identifiers :o for each query image associated with a view captured by a camera belonging to said plurality of query images:determining (S51) a confidence score that said query image represents the person associated to said candidate unique identifier, based on:• the distances with the images of the gallery associated to said camera, and • the unique identifiers of the persons represented by the images of the gallery determined by the machine learning engine;o obtaining (S53) an average confidence score that the plurality of query image represents the person associated to the candidate unique identifier by averaging the confidence scores for each image of said plurality of query images;identifying (S55) said target person as the person associated to the unique identifier that has the highest average confidence score.

6. The method of any one of claims 1 to 5, wherein identifying said target person based on said plurality of distances for each query image comprises:for each query image associated with a view captured by a camera belonging to said plurality of query images:o sorting (S61 ) the images of the gallery associated to the camera by increasing distance; o determining (S62) a reference distance as the lowest distance of order m, m being a predefined integer;identifying (S64) the target person as the person identified in the image of the gallery having the lowest reference distance of order m.

7. The method of any one of claims 1 to 6, wherein identifying said target person based on said plurality of distances for each query image comprises:for each query image associated with a view captured by a camera:o sorting (S71 ) the images of the gallery associated to the camera by increasing distance; o selecting (S72) a predefined number of the images of the gallery that have the lowest distance; counting (S74) the number of occurrences of each candidate unique identifier belonging to said predefined set of unique identifiers among the selected images of the galleries associated with the query images;identifying (S75) the target person as the person associated to the candidate unique identifier that has the highest number of occurrences.

8. The method of any one of claims 1 to 7, wherein said training images are obtained by detecting bounding boxes in a plurality of views of the area captured by the camera, the method further comprising a step of detecting (S42), in said plurality of views captured by said plurality of cameras, bounding boxes eachencompassing an image that represents a person, the plurality of query images of a target person (Persl) being images encompassed in said bounding boxes.

9. A computing system (Sys1 ) comprising:a plurality of cameras (Cam1 .1 , Cam1 .2, Cam1 .3) having different fields of views (FOV1 .1 , FOV 1 .2, FOV 1.3) of a same area (Areal), each camera of said plurality being respectively configured to capture a view of the area at a same capture time;one or more computing device (Compl) comprising one or more processing unit (Prod) configured to :o obtain (S43) a plurality of query images of a target person (Persl) from a plurality of views captured by said plurality of cameras respectively at the same capture time;o for each query image of said plurality of query images, calculate (S45) a plurality of distances between a vector of features associated with the query image and vectors of features associated with images of a gallery of images representing a person in a plurality of views of the area captured respectively at a plurality of times, distinct of said capture time, by the camera which has captured the view from which the query image is obtained; and; and o identify (S47) said target person based on said plurality of distances calculated for each query image.

10. The computing system (Sys1) of claim 9, wherein said at least one or more computing device (Compl) further comprises at least one of said memory (Mem1 ) which stores, for at least one camera of said plurality of cameras, a machine learning engine (ML1.1 , ML1 .2, ML1.3) associated with said camera, configured to determine the vector of features associated with the query image obtained from the view captured by said camera and, for each image of the gallery of images associated with said camera, the vector of features associated with the image of the gallery.11 . The computing system (Sys1 ) of claim 10, wherein said machine learning engine comprises a convolutional neural network.

12. Computer software comprising instructions to implement at least a part of a method according to one of claims 1 to 8 when the software is executed by a processor.

13. Computer-readable non-transient recording medium on which a software is registered to implement at least a part of the method according to one of claims 1 to 8 when the software is executed by a processor.