Method of multi-camera person reidentification

The multi-camera person re-identification method addresses the challenge of tracking individuals across multiple cameras by using unique local identifiers and feature vectors to compare and associate detections, achieving accurate and efficient re-identification without the need for camera calibration.

WO2025124737A1PCT designated stage expired Publication Date: 2025-06-19STUDIO FILMOWE M5 SP ZOO SPK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/086221
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing multi-camera systems struggle to effectively re-identify individuals across multiple cameras, especially in dynamic environments like stores, due to occlusions, detector imperfections, and the inability to aggregate detection data from multiple cameras.

Method used

A method for multi-camera person re-identification that involves obtaining frames from multiple cameras, executing a person detection algorithm to generate unique local identifiers and feature vectors, reconstructing paths for each person, creating single representative feature vectors, and comparing these vectors to determine if they represent the same person, using techniques like cosine similarity and graph-based association algorithms.

Benefits of technology

This method enables accurate re-identification of individuals across multiple cameras without requiring camera calibration or synchronization, improving tracking accuracy and reducing false associations, thereby enhancing security and operational efficiency in store environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023086221_19062025_PF_FP_ABST
    Figure EP2023086221_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A subject of the present invention is a system for multi-camera person reidentification. The core of the invention is a methodology used for associating person detections from various cameras as representing a single person, preferably in the store environment. A method of mutli-camera person reidentification comprising steps of: a. Obtaining at plurality of frames from each camera; b. Executing, for a plurality of frames from each camera, a person detection algorithm such that a unique local identifier is assigned to each of the detected person, and at least one feature vector representing the visual aspects of the detected person is generated for each frame; c. Reconstructing a path for each person by the tracking algorithm such that the path is associated with the unique local identifier, wherein the path is recorded as subsequent coordinates in each of the frame; d. Creating a single representative feature vector for each of the unique local identifiers on a basis of the feature vector generated for each frame; e. Indicating that at least two of the unique local identifiers represent the same person by comparing the single representative feature vectors of each of the unique local identifiers and when one of the single representative feature vectors is within a defined range from any one of the single representative feature vectors it is determined that this the single representative feature vectors represent the same person.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method of multi-camera person reidentification

[0002] A subject of the present invention is a system for multi-camera person reidentification. The core of the invention is a methodology used for associating person detections from various cameras as representing a single person, preferably in the store environment.

[0003] Background

[0004] Hereinbelow known methods for detecting and reidentifying person are presented.

[0005] AVIGILION - only works with Avigilon cameras with self-learning video analytics and selected Avigilon network video recorders (NVRs). It allows you to search for people by selecting clothing color, gender and age categorization. Expanding the watchlist with alarms. If a match is found to the watchlist, operators can be notified via armed panels or an alarm view.

[0006] COSMOEYE - a system that quickly processes images from video cameras, allows you to find, recognize and classify events on the screen and take appropriate actions. The system sends security alerts. The system improves as the analyzed data increases and constantly adapts to the production environment, sometimes specific to certain industries. The system analyzes data and responds to critical events on an ongoing basis. Optimization and streamlining of processes allow for better management of the company and translate directly into cost reduction.

[0007] HIKVISION - They use a single camera with built-in intelligent business support functions. Precise analysis of customer traffic. Options to exclude employees and regular customers from analysis. Multicamera statistics and report export. Repeat shoplifter alarm provides early warning for protection. Using heat map technology, it shows the amount of time shoppers spend in specific areas of the store. The data can be accumulated over time, allowing retailers to work out the optimal placement of products in the store and how to design the best overall store layout. Offered only with Hikvision hardware and cameras.

[0008] VEESION (startup) - Offers an intelligent anti-theft system based on artificial intelligence that allows detecting suspicious gestures in real time. The server containing the Veesion solution is installed on the surveillance cameras already existing in the store. Less than 30 minutes after connecting the recorder to the Veesion server, the store starts receiving the first alerts. Alerts are sent to a tablet or phone provided by the supermarket. These notifications are short videos allowing for quick identification of the suspect and intervention in real time.

[0009] ANAVEO - Optimizes the monitoring of people and property, adapting to all needs and types of installations. The system works with cameras. ANAVEO also offers CurveTrack®, a mobile CCTV system. Operation requires a trolley equipped with a 360-degree dome that moves silently within a track suspended from the ceiling and concealed by flexible, touch-free mirror-effect windows. Thanks to realtime 3D image analysis, it informs about the presence of an intruder in a given area, the type of intrusion, and associates this alarm with specific instructions to be followed.

[0010] None of the above commercially available solutions offer re-identification between multiple cameras.

[0011] Re-identification of people based on detection from one or more cameras will make the tracking system more resistant to disturbances, but it will have to be properly defined. The simplest method of associating subsequent detections of a person in a video stream from a single camera will be to use the so-called trackers (e.g. DeepSORT or ByteTracker), which will be able to make connections using visual detection features (both analytically and using deep neural networks for feature extraction), as well as using the analysis of changes in the position of objects (using, e.g., Kalman filtration). In practice, however, due to occlusions, imperfections of the detector or disappearance of an object from the camera's field of view, there may be a failure to detect the same person or an incorrect association of two people.

[0012] Recently, trackers have been used to track the movement of objects by combining information from many cameras with overlapping fields of view. However, this approach requires the proper arrangement of cameras and their mutual calibration, their synchronization, as well as combining data in real time. This is an area where the results are promising and allow for improved object tracking accuracy.

[0013] There are currently methods that allow for re-identification in a Closed-Set, where all people in the recordings are known a priori, also at the time of training the neural network. An illustration of this approach is a paper by S. Xuan and S. Zhang, titled "Intra-lnter Domain Similarity for Unsupervised Person Re-Identification," in IEEE Transactions on Pattern Analysis and Machine Intelligence, doi: 10.1109 / TPAMI.2022.3163451 , where in a study on a data set of 4101 people, the best result achieved among the available methods was 56.4%. This is not an approach that can be used for re-identification within stores, because the customer list is not known in advance, and it can be expanded to include a new person at any time. The Open-Set approach to reidentification extends the problem to recognize people who have not been seen before, but is often limited to assigning them to the "different person" class. The generalization approach, where each new person is assigned a separate class, which allows for association detection of previously unseen people, is a new and open problem that has been addressed in the literature only for a few years, and there are few solutions to the problem, and those that are tested, their effectiveness is still very low, definitely below the needs of our target group (e.g. for a population of 30,000 people, the best re-identification result achieved is 36% - see https: / / github.com / wanggrun / SYSU-30k). Even if the single-camera system for person tracking would be able to correctly track a person in the camera frame, an image from just one camera is not enough to cover the entire store. Even if it would be theoretically possible to place the camera to cover the entire store, the persons in the camera frame (especially for bigger stores) would be too small to be detected, and there are many obstacles that would cover in part or entirely the person (the occlusion problem). On the other hand, the multicamera systems can not easily aggregate subsequent detection from multiple cameras and connect the paths from each camera to a single path covering the entire store. The proposed methodology addresses this problem, filling the gap and connecting the tracking data across multiple cameras, providing complete data about the path in the store for each customer.

[0014] The present invention relates to a method of mutli-camera person reidentification. The method comprises of the following steps: a. Obtaining at plurality of frames from each camera; b. Executing, for a plurality of frames from each camera, a person detection algorithm such that a unique local identifier is assigned to each of a detected person, and at least one feature vector representing the visual aspects of the detected person is generated for each frame; c. Reconstructing a path for each person by the tracking algorithm such that the path is associated with the unique local identifier, wherein the path is recorded as subsequent coordinates in each of the frame; d. Creating a single representative feature vector for each of the unique local identifiers on a basis of the feature vector generated for each frame; e. Indicating that at least two of the unique local identifiers represent the same person by comparing the single representative feature vectors of each of the unique local identifiers and when one of the single representative feature vectors is within a defined range from any one of the single representative feature vectors it is determined that this the single representative feature vectors represent the same person.

[0015] Preferably at least steps d and e are performed on a historical data, preferably the historical data contain frames from each camera from a single day.

[0016] Preferably step b is performed on every N-th frame from the camera, wherein N is a number from a range of 1-100.

[0017] Preferably the single representative feature vector is calculated by averaging feature vectors.

[0018] Preferably the feature vector represents a deep features. Preferably a cosine similarity is used for comparison, in step e, of the single representative feature vectors.

[0019] Preferably the range is predefined.

[0020] Preferably the defined range is defined with a lower threshold and an upper threshold, wherein preferably the upper threshold is narrower than the lower threshold.

[0021] Preferably the step e comprise a substeps of: el. Creating a similarity matrix which contains the similarity measures for each pair of the unique local identifier, preferably calculated as a distance between the single representative feature vectors for each pair of the unique local identifiers, within the range [0,1], wherein the similarity measure equals to the distance subtracted from 1; e2. Creating the Aggregated Maximal Precision Graph, AM PG, which is constructed that the unique local identifiers within a range of the upper threshold are connected; e3. Creating the Aggregated Maximal Recall Graph, AMRG, which is constructed that the unique local identifiers within a range of the lower threshold are connected e4. Marking the largest clique in the AMRG as a potential single person; e5. Associating at least two cliques from AM PG which are covered by the largest clique in the

[0022] AMRG if all the nodes from both cliques are connected to each other as the same person; e6. Determining the potential single person as the same person if there are no conflicts between two cliques and removing all the connections not included in the clique but going out of clique; e7. Analyzing the next biggest clique is set as a candidate

[0023] The invention will be presented with a reference to the drawings which show:

[0024] Fig. 1 - an ideal representation of a graph with 4 distinct persons (each person is denoted by a separate color). This graph is not known before processing, only the similarity measure between each pair of nodes.

[0025] Fig. 2 - aggregated Maximum Precision Graph with high-confidence connections indicated

[0026] Fig. 3 - aggregated Maximal Recall Graph with all the true connections included

[0027] The developed methodology allows for the re-identification of people based on pose, combining data from different cameras. In the proposed methodology, the mutual arrangement of cameras is not important, so mutual calibration or measurements of distances and angles between cameras are not required.

[0028] A method of mutli-camera person reidentification comprising the following steps. During step a) at plurality of frames from each camera in location, i.e. store, are obtained. Next, during step b), a person detection algorithm is executed for a plurality of frames from each camera. The person skilled in the art will know that detections from frames taken in a short time may not differ much and thus it may be beneficial to analyze only some of the recorded frames - according to one embodiment every N-th frame from the camera is analyzed, wherein N is a number from a range of 1-100. A unique local identifier is assigned during this step to each of a detected person, and at least one feature vector representing the visual aspects of the detected person is generated for each frame. During step c) a path for each person is reconstructed by the tracking algorithm such that the path is associated with the unique local identifier, wherein the path is recorded as subsequent coordinates in each of the frame. During step d) a single representative feature vector is created for each of the unique local identifiers on a basis of the feature vector generated for each frame. During step e it is indicated that at least two of the unique local identifiers represent the same person by comparing the single representative feature vectors of each of the unique local identifiers and when one of the single representative feature vectors is within a defined range from any one of the single representative feature vectors it is determined that these single representative feature vectors represent the same person. A cosine similarity may be used for comparison of the single representative feature vectors

[0029] It should be noted that the feature vector provides features which are abstract and are not associated, precisely, with any physical aspect of any observed person. Those parameters, in the preferred embodiment, are the results from the detection (part of the image with the detected person) to the feature extractor after detection.. Such parameters are presented, for example, in algorithms based on an Al such as LA Transformer (https: / / github.com / SiddhantKapil / LA-Transformer) or OSNet (https: / / github.com / KaiyangZhou / deep-person-reid). Such approach provides that a single person is recognized but after the data are extracted the privacy of the person is protected since parameters may not be associated with any physical feature of the person.

[0030] The defined range, in one embodiment, is predefined, either manually or during set-up. The defined range may also change and adapt, for example during reevaluation of the results the operator, human or automated, may decide, that the defined range should be changed. The defined ranged in yet another embodiment is defined with a lower threshold and an upper threshold, wherein preferably the upper threshold is narrower than the lower threshold. In the last step, in such embodiment, comparison is made for both the lower threshold and the upper threshold for the single representative feature vectors of each of the unique local identifiers and when one of the single representative feature vectors such that two groups are generated. One group shows only true connections, wherein the unique local identifiers connected show only one person detection, even if not all unique local identifiers are connected, and the second one shows all detections of a single person, wherein it is possible that some other of the unique local identifiers not associated with this person may also be connected. Such groups shows a certain group and plausible group of the unique local identifiers related to the single person.

[0031] It should be noted that that the present invention may be performed on a fresh set of data, that is when the people observed are still in the observed area. It is also possible to only store on site data and later on analyze the full data set, from e.g. one day, few days, a weak etc.

[0032] For each camera in the store, the person detection algorithm is executed and paths for each person in a single camera frame are being reconstructed by the tracking algorithm. As a result, each person has a unique local identifier in a camera frame, and with each local id there is a path (recorded as subsequent coordinates in the camera frame) associated. Also, for each detection, a feature vector representing the visual aspects of the detected person is extracted from a bounding box of a person (or, in the case of many detections, it can be collected every N-th detection). It should be however noted that the person detection and tracking algorithms as well as feature extraction are not part of this invention. It means that the invention applies to different detection and tracking algorithms as well as various feature extractors.

[0033] The data from all the cameras for a given day are collected in a single database, containing paths and features extracted for each unique local identifiers. The reidentification algorithm fetches all the data and for each unique local identifiers creates a single feature vector (aggregated from a set of feature vectors from subsequent detection, by, for instance, averaging them). Then, the distance between the representative feature vectors for each pair of unique local identifiers is calculated, within the range [0,1], The distance is subtracted from 1, resulting in the similarity measure.

[0034] The visual-based association algorithm proceeds. In general, the detections could be represented as a graph, where a line between two nodes denotes association (i.e., these two local ids represent the same person). An exemplary graph is visualized in Fig. 1, where four persons are recognized (nodes 0, 3, 8; 4; 6, 7; 1, 2, 5, 9). It requires two parameters: lower and upper association threshold. The upper association threshold is selected in a way that the precision is maximized. It means that the processed graph should have no false connections (at the expense of some connections that should be made). The lower threshold is selected so that we have all the true connections in the graph (accepting the possibility that some of the non-true connections are there as well). The algorithm works as follows: 1. The similarity matrix is built, containing the similarity measures for each pair of unique local identifiers. Each row is associated with a separate unique local identifier, and columns are associated with the same order of unique local identifiers.

[0035] 2. The Aggregated Maximal Precision Graph (AMPG) is constructed on the basis of the upper threshold. An exemplary AMPG is presented in Fig. 2 where two Cliques were created (nodes 3, 0, 8; 2, 5, 9). Since the upper threshold was selected in a way that assures that all the true connections are there, we are certain that those connections should persist in the final processed graph.

[0036] 3. Now the lower threshold is used to construct the Aggregated Maximal Recall Graph (AMRG), where all the true connections are present, but also some non-true connections are there. An exemplary AMRG is presented in Fig. 3 where each node is associated with at least one clique (of cardinality 1 or more).

[0037] 4. Now the results from AMPG and AMRG are connected together in order to extract the true connections. For this purpose, a mathematical definition of Clique is used. A clique is a set of nodes where all the nodes are connected to each other. As a result of the reidentification algorithms, a set of (distinct) cliques representing various persons, but each clique representing a single person, is to be calculated. The AMPG cliques are used as a primer, and cannot be removed. The connection algorithm works as follows: i. Look for the largest clique in the AMRG, and mark it as a candidate. If two cliques from AMPG are covered by the candidate, they can be connected together only when all the nodes from both cliques are connected to each other. Otherwise, they should be separated. The candidate is set as a person if there are no conflicts between two cliques. When the candidate becomes the detection, all the connections not included in the clique but going out of clique's nodes are removed. A conflict should be understood that there are two cliques where not all nodes from one clique are connected to nodes from other clique. ii. Then, the next biggest clique is set as a candidate and is analyzed from the viewpoint of AMPG cliques conflicts. If there are conflicts, the biggest clique spanning a single AMPG clique is selected. Then, it is set as connected clique and all the other connections from this clique are removed. Then this is procedure is repeated until we have all the cliques (with the size of at least one node).

[0038] The constructed cliques represent the associated persons from multiple cameras

Claims

Claims1. A method of mutli-camera person reidentification comprising steps of: a. Obtaining at plurality of frames from each camera; b. Executing, for a plurality of frames from each camera, a person detection algorithm such that a unique local identifier is assigned to each of the detected person, and at least one feature vector representing the visual aspects of the detected person is generated for each frame; c. Reconstructing a path for each person by the tracking algorithm such that the path is associated with the unique local identifier, wherein the path is recorded as subsequent coordinates in each of the frame; d. Creating a single representative feature vector for each of the unique local identifiers on a basis of the feature vector generated for each frame; e. Indicating that at least two of the unique local identifiers represent the same person by comparing the single representative feature vectors of each of the unique local identifiers and when one of the single representative feature vectors is within a defined range from any one of the single representative feature vectors it is determined that this the single representative feature vectors represent the same person.

2. The method according to claim 1, wherein at least steps d and e are performed on a historical data, preferably the historical data contain frames from each camera and / or the feature vectors and unique local identifiers from a single day.

3. The method according to claim 1 or 2, wherein step b is performed on every N-th frame from the camera, wherein N is a number from a range of 1-100.

4. The method according to anyone of the previous claims, wherein the single representative feature vector is calculated by averaging feature vectors.

5. The method according to anyone of the previous claims, wherein the feature vector represents a deep features.

6. The method according to anyone of the previous claims, wherein a cosine similarity is used for comparison, in step e, of the single representative feature vectors.

7. The method according to anyone of the previous claims, wherein the defined range is predefined.

8. The method according to anyone of the previous claims, wherein the defined range is defined with a lower threshold and an upper threshold, wherein preferably the upper threshold is narrower than the lower threshold.

9. The method according to claim 8, wherein the step e comprises a substeps of: el. Creating a similarity matrix which contains the similarity measures for each pair of the unique local identifier, preferably calculated as a distance between the single representative feature vectors for each pair of the unique local identifier, within the range [0,1], wherein the similarity measure equals to the distance subtracted from 1; e2. Creating the Aggregated Maximal Precision Graph, AMPG, which is constructed that the unique local identifier within a range of the upper threshold are connected; e3. Creating the Aggregated Maximal Recall Graph, AMRG, which is constructed that the unique local identifier within a range of the lower threshold are connected e4. Marking the largest clique in the AMRG as a potential single person; e5. Associating at least two cliques from AMPG which are covered by the largest clique in the AMRG if all the nodes from both cliques are connected to each other as the same person; e6. Determining the potential single person as the same person if there are no conflicts between two cliques Removing all the connections not included in the clique but going out of clique; e7. Analyzing the next biggest clique is set as a candidate and is analyzed from the viewpoint of AMPG cliques conflicts - if there are conflicts, the biggest clique spanning a single AMPG clique is selected and after that it is set as connected clique and all the other connections from this clique are removed e8. Repeating steps e4-e7 for all other cliques