Multi-camera object tracking system and multi-camera object tracking program
The multi-camera object tracking system addresses the challenge of inconsistent tracking by using cameras to capture front and rear images, generating a three-dimensional pose for stable re-identification and accurate tracking across multiple cameras.
Patent Information
- Application Number
- JP2025101669
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-06-17
AI Technical Summary
Conventional multi-camera object tracking systems struggle with stable person re-identification and accurate tracking across multiple cameras due to the use of images or image features captured from a single angle, leading to inconsistent matching and tracking failures.
A multi-camera object tracking system that utilizes two or more cameras positioned to capture images of a person's front and rear sides, generating a three-dimensional pose and assigning an identification ID based on these images, enabling stable re-identification and tracking across multiple cameras using a three-dimensional pose tracking mechanism.
Enables accurate and stable person re-identification and tracking across multiple cameras by utilizing images from different angles, improving tracking consistency and accuracy.
Smart Images

Figure 0007752458000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a multi-camera object tracking system and a multi-camera object tracking program. [Background technology]
[0002] Conventionally, a multi-camera object tracking technique has been known in which multiple cameras are installed in a certain area and objects (such as people) within the camera's field of view are tracked using images captured by these cameras. In this multi-camera object tracking, SORT (Simple Online and Realtime Tracking) and DeepSORT (Simple Online and Realtime Tracking with a deep association metric) (see Non-Patent Document 1) are used to track objects within each camera's field of view (single camera object tracking). SORT predicts object movement using a Kalman filter, calculates the IoU between the predicted object and the detected object using a Hungarian algorithm, and tracks these objects (actually, bounding boxes) by associating them. DeepSORT is an improved version of the above SORT. Multi-camera object tracking combines (single camera) object tracking techniques such as SORT and DeepSORT with person re-identification technology.
[0003] In the above-mentioned multi-camera object tracking (Person Re-Identification), for example, a camera is installed at the starting point (tracking start point) of multi-camera object tracking, such as an entrance, and an image of a person (person image) taken by this camera, or an image of a face (face image) cut out from the person image, or image features extracted from the person image or face image, is registered in a database for registering persons to be tracked, and this registered image (person image or face image) or image feature is matched with an image or image feature taken by a camera installed other than the above-mentioned tracking start point, thereby tracking a person (person to be tracked) across the shooting ranges of multiple cameras. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Nicolai Wojke, Alex Bewley, Dietrich Paulus, “SIMPLE ONLINE AND REALTIME TRACKING WITH A DEEP ASSOCIATION METRIC”, [online], March 21, 2017, University of Koblenz-Landau, Queensland University of Technology, [Retrieved January 6, 2023], Internet<URL:https: / / userpages.uni-koblenz.de / ~agas / Documents / Wojke2017SOA.pdf> Summary of the Invention [Problem to be solved by the invention]
[0005] However, in conventional multi-camera object tracking, the person images captured by the camera installed at the tracking start point, such as an entrance, are only person images captured at one (camera) angle of view, so the images (person images or facial images) or image features registered in the database for registering tracking targets are only images or image features of the person viewed from a specific direction. If only images or image features of such a person viewed from a specific direction are used to match images or image features captured by a camera installed other than the tracking start point, stable person re-identification cannot be performed, and therefore a person (tracking target) cannot be accurately tracked across the shooting ranges of multiple cameras.
[0006] The present invention is intended to solve the above-mentioned problems, and aims to provide a multi-camera object tracking system and a multi-camera object tracking program that enable stable person re-identification and enable accurate tracking of a tracking target across the shooting ranges of multiple cameras. [Means for solving the problem]
[0007] In order to solve the above problems, a multi-camera object tracking system according to a first aspect of the present invention is a multi-camera object tracking system that performs multi-camera object tracking to track a tracked person across the shooting ranges of a plurality of cameras, wherein two or more of the plurality of cameras are installed at a starting point of the multi-camera object tracking and are arranged so as to be able to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked person; a three-dimensional pose generation means for generating a three-dimensional pose of the person to be tracked, which is a pose in three-dimensional space, based on at least an image of a substantially front side and an image of a substantially back side of the person to be tracked, which are synchronously captured by the two or more cameras; and a three-dimensional pose tracking means for assigning an identification ID to the person to be tracked corresponding to the generated three-dimensional pose, and tracking each of the people to be tracked assigned this identification ID, using the generated three-dimensional pose, in a video frame composed of images captured by the two or more cameras installed at the starting point. a tracking target registration means for registering at least the image of the substantially front side and the image of the substantially back side of the tracking target photographed synchronously by the two or more cameras, or image feature quantities of these images, and an identification ID of the tracking target in a tracking target registration database; andThe tracking target is tracked across the imaging ranges of the multiple cameras based on at least an image of the approximately front side and an image of the approximately back side of the tracking target registered in the tracking target registration database, or image feature quantities of these images. Here, the "approximately front side image" of the tracking target is a captured image that captures more than half of the tracking target's face. Furthermore, the "approximately back side image" of the tracking target is a captured image that captures the back of the tracking target's head in the case of a "head image" of the tracking target, and is a captured image that captures the back of the tracking target's head, back, and buttocks in the case of a "person image" (image of the entire body) of the tracking target. The reason for the term "head image" here is that an image of the tracking target's head on the front side is a "face image," but an image of the tracking target's head on the back side is an image of the head and not a face image.
[0009] In this multi-camera object tracking system, it is desirable that the 3D pose generation means includes a 2D pose estimation means for estimating key points of the tracked person in at least an image of the approximately front side and an image of the approximately rear side of the tracked person captured synchronously by the two or more cameras, and that the 3D pose of the tracked person is generated from the key points of the tracked person estimated by the 2D pose estimation means and calibration parameters of the two or more cameras.
[0010] In this multi-camera object tracking system, the starting point of the multi-camera object tracking may be the entrance of a building or room, and two or more of the multiple cameras may be positioned so as to capture at least an image of the approximately front side and an image of the approximately back side of a tracked person passing through the entrance.
[0011] A multi-camera object tracking system according to a second aspect of the present invention is a multi-camera object tracking system that performs multi-camera object tracking to track a tracked person across the shooting ranges of a plurality of cameras, wherein two or more of the plurality of cameras are installed at a starting point of the multi-camera object tracking and are positioned so as to be able to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked person, and further comprises a tracked person registration means that registers the at least the image of a substantially front side and an image of a substantially rear side of the tracked person, which are synchronously captured by the two or more cameras, and an identification ID of the tracked person, in a tracked person registration database;The tracking target registration means further includes an interpolated image generation and registration means for registering a head image or a person image of the tracking target in the tracking target registration database, generating a head image or a person image of the tracking target taken from an intermediate direction between the two directions based on the head image or person image of the tracking target taken from two directions using a generative AI model, and registering the generated head image or person image of the tracking target as an interpolated image in the tracking target registration database by the tracking target registration means. The tracking target is tracked across the imaging ranges of the plurality of cameras based on at least a head image or a person image of the tracking target registered in the tracking target registration database on the substantially front side and the substantially back side, and a head image or a person image of the tracking target captured from a direction intermediate between the two generated directions. .
[0012] It is desirable that this multi-camera object tracking system further comprises a two-dimensional pose estimation means for estimating key points of the tracked person in the head image or person image of the tracked person taken from the two directions, and that the interpolated image generation and registration means generates, using a generative AI model, an image of the head or person image of the tracked person taken from a direction intermediate between the two directions, using the head image or person image of the tracked person taken from the two directions and the key points of the tracked person estimated by the two-dimensional pose estimation means.
[0014] A multi-camera object tracking program according to a third aspect of the present invention is a multi-camera object tracking program for performing multi-camera object tracking in which a tracking target person is tracked across the shooting ranges of a plurality of cameras, the program comprising: When two or more of the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked object, Computer, a three-dimensional pose generation means for generating a three-dimensional pose of the person to be tracked in three-dimensional space based on at least an image of the substantially front side and an image of the substantially back side of the person to be tracked that are synchronously captured by the two or more cameras; a three-dimensional pose tracking means for assigning an identification ID to the person to be tracked that corresponds to the generated three-dimensional pose, and tracking each of the people to be tracked that have been assigned this identification ID in a video frame that is made up of images captured by the two or more cameras installed at the starting point; The tracking target registration means functions as a tracking target registration means that registers, in a tracking target registration database, the image of at least the approximately front side and the image of the approximately back side of the tracking target, or the image feature values of these images, and an identification ID of the tracking target, which are taken synchronously by two or more cameras, and the tracking target is tracked across the shooting ranges of the multiple cameras based on the image of at least the approximately front side and the image of the approximately back side of the tracking target, or the image feature values of these images, registered in the tracking target registration database.
[0015] A multi-camera object tracking program according to a fourth aspect of the present invention is a multi-camera object tracking program for performing multi-camera object tracking in which a tracked person is tracked across the shooting ranges of a plurality of cameras, wherein when two or more of the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to be able to take images of at least a substantially front side and a substantially rear side of the tracked person, the program makes a computer function as tracked person registration means for registering, in a tracked person registration database, at least the images of the substantially front side and the substantially rear side of the tracked person taken in synchronization by the two or more cameras, and an identification ID of the tracked person, and the tracked person registration means registers, in a tracked person registration database, an image of the head of the tracked person or A person image is registered in the tracked person registration database, and the computer is further made to function as interpolated image generation and registration means that generates, using a generative AI model, a head image or person image of the tracked person taken from a direction intermediate between the two directions based on a head image or person image of the tracked person taken from the two directions, and registers the generated head image or person image of the tracked person as an interpolated image in the tracked person registration database by the tracked person registration means, so that the tracked person is tracked across the shooting ranges of the multiple cameras based on the head images or person images of at least the approximately front and approximately back sides of the tracked person registered in the tracked person registration database and the generated head image or person image of the tracked person taken from a direction intermediate between the two directions. [Effects of the Invention]
[0016] According to the multi-camera object tracking system of the first aspect of the present invention and the multi-camera object tracking program of the third aspect, the images captured by two or more cameras installed at the start point of multi-camera object tracking are images of at least a substantially front side and a substantially rear side of the person being tracked, and therefore the images or image features registered in the tracked person registration database are images of at least a substantially front side and a substantially rear side of the person being tracked, or the image features of these images. This makes it possible, unlike conventional multi-camera object tracking systems, to match images or image features captured by cameras installed other than the tracking start point using images of at least a substantially front side and a substantially rear side of the person being tracked, thereby enabling stable person re-identification. Therefore, the person being tracked can be accurately tracked across the capture ranges of multiple cameras. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a block diagram showing a schematic configuration of a multi-camera object tracking system according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing a schematic hardware configuration of an edge device of the multi-camera object tracking system. [Figure 3] FIG. 2 is a block diagram showing a schematic hardware configuration of an analysis server of the multi-camera object tracking system. [Figure 4] This diagram shows the input / output relationship between the CPU of the analysis server in Figure 3 above and the functional blocks of the SoC of the edge device in Figure 2, related to multi-camera object tracking processing. [Figure 5] FIG. 5 is an explanatory diagram of the arrangement of the two entrance cameras in FIG. 4. [Figure 6]A figure showing a person detection image of the approximately front side detected by the person detection unit 31b in Figure 4 from the image captured by the entrance camera 3b, and a person detection image of the approximately back side detected by the person detection unit 31a from the image captured by the entrance camera 3a. [Figure 7] An explanatory diagram of the 17 keypoints assigned as teacher labels in the COCO dataset. [Figure 8] FIG. 1 is an explanatory diagram of a camera calibration algorithm in a pinhole camera model. [Figure 9] FIG. 5 is an explanatory diagram of a two-dimensional pose association process performed by the edge device 2a in FIG. 4. [Figure 10] An explanatory diagram of epipolar distance. [Figure 11] An illustration of the triangulation required to generate a 3D pose. [Figure 12] (a) shows an example of a 2D pose estimated by the 2D pose estimation unit of Figure 4 from an image captured at time t0, and an example of a 2D pose generated from the next image captured (at time t1); (b) shows an example of a 3D pose generated by the 3D pose generation unit of Figure 4 at time t0, and an example of a 3D pose generated next (at time t1). [Figure 13] 5A and 5B are diagrams showing examples of 2D poses estimated from captured images by the 2D pose estimation unit of FIG. 4 and 3D poses generated by the 3D pose generation unit of FIG. 4, respectively. [Figure 14] 10 is a flowchart of a multi-camera object tracking process performed by the multi-camera object tracking system according to the first embodiment. [Figure 15] FIG. 10 is a diagram showing the input / output relationship of blocks related to multi-camera object tracking processing among the functional blocks of the CPU of the analysis server and the SoC of the edge device in the multi-camera object tracking system according to the second embodiment of the present invention. [Figure 16] FIG. 10 is a diagram showing the input / output relationship of blocks related to multi-camera object tracking processing among the functional blocks of the CPU of the analysis server and the SoC of the edge device in the multi-camera object tracking system according to the third embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0019] A multi-camera object tracking system and a multi-camera object tracking program according to embodiments of the present invention will be described below with reference to the drawings. FIG. 1 is a block diagram showing a schematic configuration of a multi-camera object tracking system 10 according to a first embodiment of the present invention. As shown in FIG. 1, this embodiment describes an example in which multiple fixed cameras 3 (a collective term for entrance camera 3a, entrance camera 3b, multiple in-store cameras 3c, and exit camera 3d shown in the figure) that are surveillance network cameras that capture a predetermined capture area, and an edge device 2 (a collective term for edge devices 2a, 2c, and 2d) that analyzes images from each fixed camera 3 are installed inside a building or room (corresponding to store S in FIG. 1). In this embodiment, two of the multiple fixed cameras 3 (entrance camera 3a and entrance camera 3b) are installed at the starting point of multi-camera object tracking, i.e., the entrance of store S (referred to as the "building or room" in the claims).
[0020] The multi-camera object tracking system 10 is mainly composed of the above-mentioned multiple fixed cameras 3 and edge device 2 installed in a store S, and an analysis server 1 arranged on a cloud C. The edge device 2 and analysis server 1 correspond to the "computer" in the claims. The edge device 2 and analysis server 1 are connected via the Internet.
[0021] 1, the multi-camera object tracking system 10 includes a hub 7 and a router 8 in a store S where the fixed cameras 3 and edge devices 2a, 2c, and 2d are located. Note that the following description will be given of an example in which the fixed cameras 3 and the edge device 2 are separate entities, but the cameras in the multi-camera object tracking system of the present invention may also be so-called AI cameras in which the fixed cameras 3 and the edge device 2 are integrated.
[0022] The fixed cameras 3 have IP addresses and can be directly connected to a network. As shown in FIG. 1, edge device 2a is connected to two fixed cameras 3 (entrance cameras 3a and 3b) at the entrance via a LAN (Local Area Network), and performs 3D pose tracking of each person (tracking target) in the captured images based on the captured images input from each of these fixed cameras 3. Details of 3D pose tracking will be described later. Edge device 2c is connected to in-store camera 3c, and performs single-camera tracking of each person (tracking target) in the captured images based on the captured images input from in-store camera 3c. Edge device 2d is connected to exit camera 3d, and performs single-camera tracking of each person (tracking target) in the captured images based on the captured images input from exit camera 3d.
[0023] The analysis server 1 is a server installed in the management department (head office, etc.) that manages and supervises each store S. Although details will be described later, the analysis server 1 performs multi-camera object tracking processing, which is processing for tracking a tracking target across the shooting ranges of multiple cameras, using the results of 3D pose tracking of the tracking target performed by the edge device 2a and the results of single-camera object tracking of the tracking target performed by the edge devices 2c and 2d.
[0024] Next, the hardware configuration of the edge device 2 will be described with reference to FIG. 2. The edge device 2 includes a System-on-a-Chip (SoC) 11, a hard disk 12 for storing various data and programs, a Random Access Memory (RAM) 13, and a communication control IC 15. The SoC 11 includes a CPU 11a for controlling the entire device and performing various calculations, and a GPU 11b for inference processing of trained Deep Neural Network (DNN) models for various inference processing, including the object detection processing and feature vector extraction processing. The data stored on the hard disk 12 includes video data obtained by decoding the video streams (data) input from each of the fixed cameras 3. The programs stored on the hard disk 12 include trained DNN models for various inference processing, including object (tracked person) detection processing and feature vector extraction processing, and a device control program (for the edge device 2). The edge device 2a of the edge device 2 includes a part of the multi-camera object tracking program claimed in the claims.
[0025] Next, the hardware configuration of the analysis server 1 will be described with reference to Fig. 3. The analysis server 1 comprises a CPU 21 that controls the entire device and performs various calculations, a hard disk 22 that stores various data and programs, a RAM (Random Access Memory) 23, a display 24, an operation unit 25, and a communication unit 26. The hard disk 22 stores a tracking target registration DB 27 (referred to as the "tracking target registration database" in the claims) and a multi-camera object tracking program 28. The multi-camera object tracking program 28 is part of the multi-camera object tracking program in the claims, and is a program that mainly corresponds to the multi-camera tracking unit 39 in Fig. 4.
[0026] Next, an overview of the multi-camera object tracking process in the multi-camera object tracking system 10 according to the first embodiment of the present invention will be described with reference to FIG. 4. FIG. 4 shows the input / output relationship of functional blocks related to the multi-camera object tracking process among the SoC 11 of the edge device 2 (a collective term for the edge devices 2a to 2c) and the CPU 21 of the analysis server 1. The edge device 2a includes, as functional blocks, person detection units 31a and 31b, 2D pose estimation units 32a and 32b, a 3D pose generation unit 33, a 3D pose tracking unit 34, an image feature extraction unit 35a, and a tracked person registration unit 36. The 2D pose estimation units 32a and 32b, the 3D pose generation unit 33, the 3D pose tracking unit 34, and the tracked person registration unit 36 in FIG. 5 correspond to the 2D pose estimation means, the 3D pose generation means, the 3D pose tracking means, and the tracked person registration means, respectively.
[0027] The person detection units 31a and 31b shown in Figure 4 extract frame images from video frames input from entrance cameras 3a and 3b, respectively, placed at the entrances of store S, detect people included in each frame image, and send frame images (person detection images) with bounding box information for each detected person added to two-dimensional pose estimation units 32a and 32b.
[0028] The arrangement of the entrance cameras 3a and 3b will now be described in detail with reference to FIG. 5. As shown in FIG. 5, of the two entrance cameras, entrance camera 3a is installed above an entrance ENT (starting point of multi-camera object tracking) such as a door (including an automatic door) and captures an image of the approximately rear side of the tracking target passing through the entrance (captured image 41 in FIG. 4). In contrast, the other entrance camera 3b is installed, for example, above a passageway (such as on the ceiling) diagonally in front of the entrance and captures an image of the approximately front side of the tracking target passing through the entrance (captured image 42 in FIG. 4). Here, in this embodiment, the "approximately front side image" of the tracking target refers to a captured image that captures at least the nose, left eye, and right eye, among the main facial features of the tracking target. However, in the present invention, the "approximately front side image" of the tracking target may be any captured image that captures more than half of the tracking target's face. Furthermore, "an image of the approximate rear side" of the person to be tracked means a captured image showing the back of the head of the person to be tracked if the image registered in the DB27 for registering people to be tracked is an image of the head of the person to be tracked (a facial image), and means a captured image showing the back of the head, back, and buttocks of the person to be tracked if the image registered in the DB27 for registering people to be tracked is a human image of the person to be tracked (an image of the entire body).
[0029] As shown in FIG. 5, entrance camera 3a and entrance camera 3b are preferably positioned so that the angle between their optical axes 43a and 43b is between 110° and 140°. The reason why it is better not to position the two entrance cameras (entrance camera 3a and entrance camera 3b) so that the angle between their optical axes 43a and 43b is close to 180° is as follows. That is, if entrance camera 3a is installed above entrance ENT as shown in FIG. 5, and entrance camera 3b is positioned so that its optical axis 43b forms an angle of 180° with the optical axis 43a of entrance camera 3a (the optical axes 43a and 43b overlap), the tracked person passing through the entrance will be photographed completely from both the front and back. Since people tend to gather densely at entrances, occlusion is likely to occur. When multiple people (tracking targets) passing through an entrance where occlusion is likely to occur are photographed completely from the front and the back, if the people overlap when viewed from the back, they will also overlap when viewed from the front. For this reason, it is difficult to accurately track multiple people passing through an entrance if the people passing through the entrance are photographed completely from the front and the back. Furthermore, if people passing through an entrance are photographed completely from the front and the back, it becomes difficult to generate (estimate) a three-dimensional pose, making it difficult to perform three-dimensional pose tracking based on the generated three-dimensional pose. For this reason, in this embodiment, as shown in FIG. 5, two entrance cameras (entrance camera 3a and entrance camera 3b) are positioned so that the angle between their optical axes 43a and 43b is 110° to 140°.
[0030] According to the results of tests conducted by our company, by arranging the two entrance cameras as described above, even if multiple people passing through the entrance overlap in front and behind in the image taken by entrance camera 3a, it is possible to increase the likelihood that multiple people will not overlap in front and behind in the image taken by entrance camera 3b in the image taken by entrance camera 3b in the front. Also, by arranging the two entrance cameras so that the angle formed by their optical axes 43a and 43b is 110° to 140°, it is possible to ensure that the images taken by these entrance cameras are not images taken completely from the front and back, making it easier to generate (estimate) 3D poses and to perform 3D pose tracking based on the generated 3D poses.
[0031] 5 indicates the portion of the shooting ranges of the two entrance cameras (entrance camera 3a and entrance camera 3b) where these shooting ranges overlap. This overlapping shooting area 44 is preferably large (wide), and is preferably at least 9 square meters or more.
[0032] Furthermore, the image 41 captured by the entrance camera 3a and the image 42 captured by the entrance camera 3b must be acquired (captured) in synchronization. Specifically, the image 41 captured by the entrance camera 3a and the image 42 captured by the entrance camera 3b are acquired simultaneously, or the images captured by the entrance cameras 3a and 3b are acquired with a predetermined time offset. The edge device 2a (SoC 11) shown in FIG. 4 acquires the image 41 captured by the entrance camera 3a and the image 42 captured by the entrance camera 3b in synchronization as described above, and sends the acquired image 41 captured by the entrance camera 3a and the image 42 captured by the entrance camera 3b to the person detection units 31a and 31b, respectively.
[0033] 6 shows a person detection image 48 (frame image with bounding box information of the detected person added) of the approximately front side detected by person detection unit 31b from image 42 captured by entrance camera 3b, and a person detection image 49 of the approximately rear side detected by person detection unit 31a from image 41 captured by entrance camera 3a. These person detection images may be original frame images to which bounding box information of each detected person has been added, or may be original frame images to which bounding box information has been added, as shown in FIG. 6, to which bounding box information has been added.
[0034] As shown in FIG. 4, a person detection image 49 of the approximately rear side detected by the person detection unit 31a and a person detection image 48 of the approximately front side detected by the person detection unit 31b are sent to two-dimensional pose estimation units 32a and 32b, respectively. The two-dimensional pose estimation unit 32a estimates key points of the detected person (tracked person) in the person detection image 49 of the approximately rear side detected by the person detection unit 31a, and the two-dimensional pose estimation unit 32b estimates key points of the detected person in the person detection image 48 of the approximately front side detected by the person detection unit 31b. In the estimation process by the two-dimensional pose estimation units 32a and 32b, 17 key points (see FIG. 7) of the detected person (tracked person) are estimated. FIG. 7 shows the 17 key points assigned as teacher labels in the COCO (Common Objects in Context) dataset.
[0035] The key points of the tracked person estimated by the above-mentioned two-dimensional pose estimation units 32a and 32b are sent to the three-dimensional pose generation unit 33 shown in Fig. 4 together with the calibration parameters (internal parameters and external parameters) of the entrance cameras 3a and 3b acquired in advance. The three-dimensional pose generation unit 33 generates a three-dimensional pose of the tracked person from the key points of the tracked person estimated by the two-dimensional pose estimation units 32a and 32b (in the person detection images 49 and 48 of the approximately rear and approximately front sides) and the calibration parameters of the entrance cameras 3a and 3b. The calibration parameters of the entrance cameras 3a and 3b are estimated by calibration of the entrance cameras 3a and 3b.
[0036] The above calibration parameters include intrinsic parameters, extrinsic parameters, and (lens) distortion coefficients. Generally, in camera calibration, 3D world points and their corresponding 2D image points are required to estimate the camera calibration parameters. This correspondence can be obtained using multiple images of a calibration pattern such as a checkerboard.
[0037] Figure 8 is an explanatory diagram of the camera calibration algorithm for the pinhole camera model. A pinhole camera is a simple camera with no lens and only one small opening. As shown in Figure 8, the extrinsic parameters are parameters for converting from the 3D world coordinate system to the 3D camera coordinate system, and the intrinsic parameters are parameters for projective transformation from the 3D camera coordinates to the 2D image coordinates (coordinates on the image plane). The extrinsic parameters consist of a rotation matrix R and a transformation matrix t, and are generally denoted as [Rt]. The intrinsic parameters include the focal length, optical center, and shear coefficient, and the intrinsic parameter matrix K is defined as follows:
[0038]
number
[0039] In the above matrix K, [c x c y ] is the optical center (principal point) in pixels, and (f x , f y ) is the focal length in pixels and s is the shear factor.
[0040] Next, the 3D pose generation process of the tracked person performed by the 3D pose generation unit 33 will be described in more detail. FIG. 9 is an explanatory diagram of the 2D pose association process performed by the edge device 2a (SoC 11 thereof) before the 3D pose generation process by the 3D pose generation unit 33. This diagram schematically shows people 53 to 56. In the image 51 captured by the entrance camera 3a shown in FIG. 9, the people (tracked people) on the substantially rear side detected by the person detection unit 31a are people 53 and 54, and in the image 52 captured by the entrance camera 3b, the people (tracked people) on the substantially front side detected by the person detection unit 31b are people 55 and 56. At this time, the edge device 2a (SoC11) associates the persons 53 and 54 (their poses) in the captured image 51 with the persons 55 and 56 (their poses) in the captured image 52 based on the average epipolar distance between each key point (each of the 17 key points) of the persons 53 and 54 in the captured image 51 by the entrance camera 3a and each key point of the persons 55 and 56 in the captured image 52 by the entrance camera 3b.
[0041] The "average epipolar distance" is the average value of the "epipolar distances" between corresponding keypoints of each person photographed by a different entrance camera. The "epipolar distance" is the distance described below. For example, as shown in FIG. 9, assume that the position of the left ear, which is one of the keypoints of person 54 estimated by 2D pose estimation unit 32a, is keypoint K1, and the positions of the left ears of persons 55 and 56 estimated by 2D pose estimation unit 32b are keypoints K1' and K1" respectively. In this case, when the above keypoints K1, K1', and K1" are plotted on a conceptual diagram of epipolar geometry as shown in FIG. 10, the distance between epipolar line l' on photographed image 52, which is the epipolar line of keypoint K1 of person 54, and keypoints K1' and K1" of persons 55 and 56 is the "epipolar distance." For example, in FIG. 10, the distance d' between the epipolar line l' and the key point K1' of person 55 is the "epipolar distance." In FIG. 10, the distance d' between the epipolar line l' and the key point K1" of person 56 is zero, and therefore is not shown. In FIGS. 9 and 10, it is assumed that person 54 in photographed image 51 and person 56 in photographed image 52 are the same person, and therefore the distance d' between epipolar line l', which is the epipolar line of key point K1 of person 54, and key point K1" of person 56 is set to zero. On the other hand, it is assumed that person 54 in photographed image 51 and person 55 in photographed image 52 are different people, and therefore the distance d' between epipolar line l', which is the epipolar line of key point K1 of person 54, and key point K1' of person 55 is not zero. Note that point x and point K1 on photographed image 51 in FIG. 10 are the same point.
[0042] Furthermore, the above-mentioned "average epipolar distance" refers to the average value of the "epipolar distance" between, for example, each of the 17 key points of person 54 in image 51 captured by entrance camera 3a (of which key points appear in both captured image 51 and captured image 52) and each of the 17 key points of person 56 in image 52 captured by entrance camera 3b (of which key points appear in both captured image 51 and captured image 52). This "average epipolar distance" approaches zero when the people being compared are the same person, and increases when the people being compared are different people. Therefore, by associating (the poses of) people 53 and 54 in captured image 51 with (the poses of) people 55 and 56 in captured image 52 based on the "average epipolar distance," it is possible to accurately associate the person (tracking target) detected in image 51 captured by entrance camera 3a with the person detected in image 52 captured by entrance camera 3b.
[0043] Once the above-mentioned two-dimensional pose association process is completed, the three-dimensional pose generation unit 33 uses the key points (positions on the two-dimensional image) of the person (tracked person) approximately at the rear side estimated by the two-dimensional pose estimation unit 32a and the key points (positions on the two-dimensional image) of the person approximately at the front side estimated by the two-dimensional pose estimation unit 32b to determine the position in three-dimensional space (world coordinates) of each key point of each person (tracked person) detected by the person detection units 31a and 31b by triangulation, and generates a three-dimensional pose of each detected person.
[0044] The above triangulation means that in stereo vision, once corresponding points (points x and x' in Figure 11) in images from different viewpoints (centers of the camera lenses) are determined, the three-dimensional position X of that point is determined as the intersection of the viewpoints (points O and O' in Figure 11) and the line of sight (W and W' in Figure 11) that passes through those points (points x and x') on the image plane.
[0045] Next, using FIG. 10 above, we will explain the principle of triangulation using images captured by cameras with different viewpoints, such as image 51 captured by entrance camera 3a and image 52 captured by entrance camera 3b. This figure shows a situation when a scene is captured by two cameras with different viewpoints. Points O and O' are the viewpoints (projection centers) of these cameras, and point X in three-dimensional space is projected onto the projection planes of the two cameras (corresponding to captured images 51 and 52). If only the left camera is used, it is unknown to which point in three-dimensional space point x in the image captured by the left camera corresponds. This is because any point on the line OX connecting point x to the viewpoint (projection center) O of the left camera is mapped to point x in the left image. However, in the image captured by the right camera, each point on the line OX is mapped to a different point. This shows that if we have two images of the same scene taken from different viewpoints, we can triangulate point X in three-dimensional space. Note that points on line OX form line l' in the image taken by the right camera. This line l' is the "epipolar line" of point x mentioned above.
[0046] Next, referring to FIG. 11 above, a calculation method will be described in which the 3D pose generation unit 33 determines the positions of the key points of this person in three-dimensional space (world coordinates) by triangulation using the key points (positions on the 2D image) of the person on the approximately rear side estimated by the 2D pose estimation unit 32a, the key points (positions on the 2D image) of the person on the approximately front side estimated by the 2D pose estimation unit 32b, and the calibration parameters (internal parameters and external parameters) of the entrance cameras 3a and 3b.
[0047] First, using the calibration parameters (internal parameters and external parameters) of the entrance cameras 3a and 3b acquired in advance, a 3-row, 4-column camera matrix P is calculated using the following formula: This camera matrix P is a matrix for mapping a scene in a 3D space onto an image plane. P = K[Rt] (2)
[0048] Using the above camera matrix P, the relationship between the coordinate position (x, y) of point x on the image plane (for example, the coordinate position of point x (key point K1) on image 51 captured by entrance camera 3a in FIG. 11 or the coordinate position of point x' (key point K1") on image 52 captured by entrance camera 3b) and the coordinate position (X, Y, Z) of point X in three-dimensional space (world coordinates) can be expressed as the following camera transformation formula.
[0049]
number
[0050] Here, λ is the scale factor (a coefficient representing the size (scale) of the object), (x, y) is the coordinate position of point x on the image plane, and (X, Y, Z) is the coordinate position of point X in three-dimensional space (world coordinates).
[0051] When the camera transformation formula (3) above is applied to points x1 and x2 on the images captured by two cameras (for example, entrance cameras 3a and 3b), such as points x and x' in FIG. 11, the following is obtained. λ1x1=P1X (4) λ2x2=P2X (5)
[0052] Here, x1 and x2 in the above equations (4) and (5) correspond to the matrix in the second term on the left side of equation (3), and X in equations (4) and (5) corresponds to the matrix in the second term on the right side of equation (3).
[0053] The above two equations (4) and (5) can be expressed as a matrix as follows:
[0054]
number
[0055] The above determinant takes the form Ax = 0, so by solving x (the second term on the left side of equation (6)) using SVD (Singular Value Decomposition), the coordinate position of point X in three-dimensional space (world coordinates) can be restored.
[0056] Using the above method, the three-dimensional pose generation unit 33 determines the positions in three-dimensional space (world coordinates) of all (17) key points for each person captured in both the image 51 captured by the entrance camera 3a and the image 52 captured by the entrance camera 3b. Then, by connecting all of the key points for each person captured in the captured images 51 and 52, the three-dimensional pose generation unit 33 generates three-dimensional poses 66 and 67, which are the poses of each person in three-dimensional space (world coordinates), as shown in Figure 12(b).
[0057] The 3D pose tracking unit 34 shown in Fig. 4 assigns an identification ID to the tracked person corresponding to each 3D pose generated by the 3D pose generation unit 33, and tracks the person (each of the tracked person) assigned this identification ID throughout the entire video frame consisting of images captured by the entrance cameras 3a and 3b using the person's 3D pose. Specifically, the 3D pose tracking unit 34 links the previous 3D pose (the 3D pose generated from the previous image (at time t0) captured by the entrance cameras 3a and 3b) (see Fig. 12(b)) for the same person generated by the 3D pose generation unit 33 with the next 3D pose (the 3D pose generated from the next image (at time t1) captured by the entrance cameras 3a and 3b) (see Fig. 12(b)). For each of the keypoints appearing in both of these two consecutive 3D poses, the 3D pose tracking unit 34 calculates the Euclidean distance between corresponding keypoints in the two 3D poses. If the average value of these Euclidean distances is less than a predetermined threshold, the 3D pose tracking unit 34 integrates these 3D poses into the same track. On the other hand, if the average value of the Euclidean distances is equal to or greater than the predetermined threshold, the 3D pose tracking unit 34 generates a new track based on the 3D pose generated from the latest captured image (at time t1) (and assigns a new identification ID to the person (tracking target) corresponding to the 3D pose).
[0058] 12(a) shows examples of 2D poses 61-64 estimated by 2D pose estimation units 32a and 32b from the immediately preceding captured images 51 and 52 (at time t0) and 2D poses 61'-64' generated from the immediately following captured images 51' and 52' (at time t1), while FIG. 12(b) shows examples of 3D poses 67 and 68 generated by 3D pose generation unit 33 (at time t0) and 3D poses 67' and 68' (at time t1). It can be seen from FIGS. 12(a) and 12(b) that the people (poses) that overlap in the 2D poses 61-64 based on the 2D captured images 51 and 52 are clearly separated in the 3D poses 67 and 68, allowing 3D pose tracking unit 34 to accurately assign an identification ID to each person and accurately track each person.
[0059] 12(a) and 12(b), FIGS. 13(a) and 13(b) show a comparison of 2D poses 73 to 76 and 73' to 76' estimated by 2D pose estimation units 32a and 32b from captured images 71 and 72 with 3D poses 73" to 76" generated by 3D pose generation unit 33. 2D poses 73 and 73' and 3D pose 73" are poses of the same person. Similarly, 2D poses 74 and 74' and 3D pose 74"; 2D poses 75 and 75' and 3D pose 75"; and 2D poses 76 and 76' and 3D pose 76" are poses of the same person.
[0060] 13(a) and 13(b) show that the 3D poses 73" to 76" generated by the 3D pose generator 33 can accurately map 3D poses for four closely overlapping people by using images (images of the approximately rear and approximately front sides) captured from different directions by the two entrance cameras (entrance camera 3a and entrance camera 3b). Therefore, 3D pose tracking using these 3D poses can accurately track each person in FIGS. 13(a) and 13(b). However, if multiple closely overlapping people as described above are tracked using conventional single-camera object tracking (such as SORT, DeepSORT, or Deep OC-SORT) using 2D captured images, a swapping problem occurs, in which the identification IDs assigned to the people (tracking targets) are swapped.
[0061] The reason for this difference is that even if multiple people overlap very closely in a 2D camera image (photographed image), they can still be distinguished in 3D space. 3D pose tracking can significantly reduce the occurrence of the above-mentioned swapping compared to conventional single-camera object tracking using 2D photographed images. Deep OC-SORT, mentioned above, incorporates an object appearance matching technique into OC-SORT, a "motion modeling" tracking model (a model that predicts the movement of an object based on its previous movement).
[0062] 4 extracts features of a person image or a face image ("head image" in the claims) of the approximately back side and the approximately front side of the person detected by the person detection units 31a and 31b, and sends the extracted image features to the tracking target person registration unit 36. Note that, in the claims, the term "head image" (of the tracking target person) is used for strict definition, but the term "face image" is easier to understand, so in the embodiment, the term "face image" is used.
[0063] The tracked person registration unit 36 uses the tracking results of each tracked person (person) by the 3D pose tracking unit 34 to perform a process of (additionally) registering, in the tracked person registration DB27 (tracked person registration database in the claims) of the analysis server 1, the feature amounts (image feature amounts) of the person images or facial images of the approximately back and approximately front sides of each tracked person (detected person) photographed synchronously by the entrance cameras 3a, 3b, and the identification ID of each tracked person.
[0064] In addition, the tracking target registration unit 36 may use the tracking results of each tracking target (person) by the 3D pose tracking unit 34 to register the person images or face images of the approximately back and approximately front sides of each tracking target (detected person) photographed synchronously by the entrance cameras 3a, 3b, and the identification ID of each tracking target, in the tracking target registration DB 27 of the analysis server 1.
[0065] In this embodiment, the "track," which is data of the tracking result by the 3D pose tracking unit 34, includes the movement trajectory of the tracked person (person) and feature amounts (image feature amounts) of the person's image or face image on the approximately rear and approximately front sides of the person. In this embodiment, the movement trajectory included in the track, in addition to the feature amounts (image feature amounts) of the person's image or face image on the approximately rear and approximately front sides and an identification ID, are registered (stored) in the tracking object DB 39 as information about a certain tracked person (person).
[0066] In this embodiment, only the edge device 2a (the tracked person registration unit 36 thereof) assigns (allocates) an identification ID for multi-camera object tracking, so there is no need to distinguish between an identification ID for tracking by the three-dimensional pose tracking unit 34 and an identification ID for multi-camera object tracking by the multi-camera tracking unit 39. However, if the edge devices 2c and 2d also assign (allocate) an identification ID for multi-camera object tracking, they need to distinguish between an identification ID for tracking by the three-dimensional pose tracking unit 34 (similar to a so-called local ID) and an identification ID for multi-camera object tracking by the multi-camera tracking unit 39 (a so-called global ID).
[0067] In this embodiment, for people detected in images other than those captured by the entrance cameras 3a and 3b, normal single-camera object tracking (using two-dimensional captured images) is performed. Images captured by the in-store camera 3c are sent to the edge device 2c, and the person detection unit 31c of the edge device 2c detects people in the captured images. The two-dimensional single-camera tracking unit 38a of the edge device 2c performs single-camera object tracking, which tracks people detected by the person detection unit 31c within the capture range of the in-store camera 3c. Specifically, the two-dimensional single-camera tracking unit 38a associates the bounding box of the person detected by the person detection unit 31c with the track using techniques such as SORT (Simple Online and Realtime Tracking) or DeepSORT (Simple Online and Realtime Tracking with a deep association metric). The image feature extraction unit 35b of the edge device 2c extracts features of the person image or face image of the person detected by the person detection unit 31c.
[0068] The SoC 11 of the edge device 2c (see FIG. 2) sends a query, which is a "track," which is data on the tracking result obtained by the 2D single-camera tracking unit 38a, to the multi-camera tracking unit 39 of the analysis server 1. The "track" includes image features extracted by the image feature extraction unit 35b. The multi-camera tracking unit 39 of the analysis server 1 performs multi-camera object tracking processing for the query sent from the edge device 2c. Specifically, the multi-camera tracking unit 39 matches the image features included in the query sent from the edge device 2c with the image features of each tracked target registered in the tracked target registration DB 27 to estimate which identification ID registered in the tracked target registration DB 27 the track obtained by the single-camera object tracking of the edge device 2c is associated with, and assigns the estimated identification ID to the corresponding track. Note that the above query refers to a command statement requesting the database management system (DBMS) of the tracked target registration DB 27 to perform processing such as data search.
[0069] If the CPU 21 of the analysis server 1 determines that the feature vector of the track sent as a query from the edge device 2c matches the feature vector of track 1 registered as information about a certain tracked person (person 1), the CPU 31 of the analysis server 1 considers the track sent from the edge device 2c to be that of person 1, and updates the information about person 1 registered in the tracked person registration DB 27 using the track information included in the query, using the database management system. That is, for example, the CPU 31 updates the information about person 1 registered in the tracked person registration DB 27 by updating the movement trajectory of person 1 taking into account the movement trajectory of person 1 within the shooting range of the in-store camera 3c, and updates the feature vector of person 1 using the feature vector of the track included in the query (for example, by replacing the feature vector of person 1 registered in the tracked person registration DB 27 with the feature vector included in the query, or by adding the feature vector included in the query to the feature vector of person 1 registered in the tracked person registration DB 27).
[0070] The single-camera object tracking processing performed by edge device 2d on images captured by exit camera 3d, and the multi-camera object tracking processing performed by multi-camera tracking unit 39 of analysis server 1 on queries sent from edge device 2d, are similar to the single-camera object tracking processing performed by edge device 2c on images captured by in-store camera 3c described above, and the multi-camera object tracking processing performed by multi-camera tracking unit 39 of analysis server 1 on queries sent from edge device 2c.
[0071] As mentioned above, only the edge device 2a (its tracked person registration unit 36) that processes the images captured by the entrance cameras 3a and 3b assigns (allocates) an identification ID for multi-camera object tracking to the tracked person, while the edge devices 2c and 2d that process the images captured by the in-store camera 3c and the exit camera 3d do not assign (allocate) an identification ID for multi-camera object tracking to the tracked person. This makes it possible to avoid new tracked persons emerging from images captured by cameras located inside the store S (the "building or room" in the claims), and therefore enables accurate person re-identification in the multi-camera object tracking process.
[0072] The tracking target person registration DB 27 is an image-related database called a gallery. In the first embodiment, the image-related data of the tracking target person registered in the tracking target person registration DB 27 is the feature amount (image feature amount) of an image of the tracking target person on the substantially front side and an image of the tracking target person on the substantially back side (person image or face image).
[0073] Here, the multi-camera object tracking process performed by the multi-camera object tracking system 10 according to the first embodiment will be summarized with reference to the flowchart in Fig. 14. First, the person detection units 31a and 31b of the edge device 2a shown in Fig. 4 detect people from captured images input from the entrance cameras 3a and 3b, and send the person detection images of the approximately rear side and the approximately front side to the 2D pose estimation units 32a and 32b, respectively (S1). The 2D pose estimation units 32a and 32b estimate key points of the tracked person in the person detection images of the approximately rear side and the approximately front side sent from the person detection units 31a and 31b, respectively (S2). The 3D pose generation unit 33 generates a 3D pose of the tracked person from the key points of the tracked person (in the person detection images of the approximately rear side and the approximately front side) estimated by the 2D pose estimation units 32a and 32b and the calibration parameters of the entrance cameras 3a and 3b obtained in advance (S3). Then, the three-dimensional pose tracking unit 34 shown in FIG. 4 assigns an identification ID to the person to be tracked corresponding to each three-dimensional pose generated by the three-dimensional pose generation unit 33, and tracks the person to whom this identification ID has been assigned (each of the people to be tracked) using the person's three-dimensional pose throughout the entire video frame consisting of images captured by the entrance cameras 3a and 3b (S4).
[0074] Next, the tracked person registration unit 36 uses the tracking results of each of the tracked persons by the 3D pose tracking unit 34 to register feature amounts (image feature amounts) of the person images or facial images of the approximately back and approximately front sides of each of the tracked persons captured synchronously by the entrance cameras 3a and 3b, as well as the identification IDs of each of the tracked persons, in the tracked person registration DB 27 of the analysis server 1 (S5). Then, the multi-camera tracking unit 39 of the analysis server 1 matches the image feature amounts included in the query sent from the edge device 2c with the image feature amounts of each of the tracked persons registered in the tracked person registration DB 27, thereby tracking each of the tracked persons across the shooting ranges of the multiple cameras (entrance cameras 3a and 3b, in-store camera 3c, and exit camera 3d) (performing multi-camera object tracking for each of the tracked persons) (S6).
[0075] As described above, according to the multi-camera object tracking system 10 of the first embodiment, the images captured by the two cameras (entrance camera 3a and entrance camera 3b) installed at the start point (entrance) of multi-camera object tracking are images of the approximately front and rear sides of the tracked person, and therefore the image features registered in the tracked person registration DB 27 are the image features of the approximately front and rear sides of the tracked person. This makes it possible, unlike conventional multi-camera object tracking systems, to match the image features of at least the approximately front and rear sides of the tracked person with the features of an image of the tracked person captured by a camera installed other than the tracking start point, thereby enabling stable person re-identification. Therefore, the tracked person can be accurately tracked across the capture ranges of multiple cameras.
[0076] In conventional multi-camera object tracking systems, only one camera is installed at the starting point of multi-camera object tracking, such as an entrance, and therefore the person images captured by the camera installed at the tracking starting point are only those captured at a single (camera) angle of view. As a result, the images (person images or facial images) or image features registered in the database for registering tracking targets are only images or image features of the person viewed from a specific direction. If only images or image features of such a person viewed from a specific direction (e.g., an image of the person viewed from behind) are used to match with images or image features captured by a camera installed at a point other than the tracking starting point, stable person re-identification cannot be performed.
[0077] In contrast, the multi-camera object tracking system 10 according to the first embodiment, unlike conventional multi-camera object tracking systems, can match image features of at least an image of the approximately front side and an image of the approximately rear side of the tracking target with features of an image captured by a camera installed at a location other than the tracking start point. Then, for example, if either the image feature of the approximately front side image of the tracking target or the image feature of the approximately rear side image of the tracking target registered in the tracking target registration DB 27 matches the feature of an image of the tracking target captured by a camera installed at a location other than the tracking start point, it is determined that matching of this feature has been successful. This enables successful matching regardless of the facial orientation of the person (tracking target) detected in the image captured by a camera installed at a location other than the tracking start point, thereby enabling stable person re-identification. Therefore, the tracking target can be accurately tracked across the capture ranges of multiple cameras.
[0078] Furthermore, according to the multi-camera object tracking system 10 of the first embodiment, a three-dimensional pose of a tracked person is generated, an identification ID is assigned to the tracked person corresponding to the generated three-dimensional pose, and each of the tracked people assigned this identification ID is tracked using the generated three-dimensional pose throughout the entire video frame consisting of images captured by two or more cameras installed at the entrance (the starting point of the multi-camera object tracking). As described above, even multiple people who are very close together and overlapping in two-dimensional captured images can be identified in three-dimensional space. Therefore, three-dimensional pose tracking using three-dimensional poses can reduce the occurrence of the above-mentioned swapping (the problem of the identification IDs assigned to tracked people being swapped during tracking) compared to conventional single-camera object tracking using two-dimensional captured images. This allows accurate identification IDs and image features of people (tracked people) captured in captured images to be registered in the tracked person registration DB 27 at the starting point (entrance) of the multi-camera object tracking.
[0079] Furthermore, according to the multi-camera object tracking system 10 of the first embodiment, (two-dimensional) key points of the tracked person are estimated in the images of the approximately front side and the approximately back side of the tracked person that are synchronously captured by the entrance cameras 3 a and 3 b, and a three-dimensional pose of the tracked person is generated from the estimated (two-dimensional) key points of the tracked person and the calibration parameters of the entrance cameras 3 a and 3 b. This makes it possible to generate an accurate three-dimensional pose of the tracked person.
[0080] Next, a multi-camera object tracking system 80 according to a second embodiment of the present invention will be described with reference to Fig. 15. The multi-camera object tracking system 80 according to the second embodiment differs from the multi-camera object tracking system 10 according to the first embodiment mainly in the following two points. The first difference is that what is registered in the tracked person registration DB 27 is not the image feature amounts of images (face images or person images) of the approximately front and rear sides of the tracked person, but images (face images or person images) of the approximately front and rear sides of the tracked person, and multi-camera object tracking of the tracked person is performed based on these images.
[0081] The second difference is that the edge device 2a includes an interpolated image generation and registration unit 81 that uses a generative AI model to generate a facial image or a person image of the tracked target captured from an intermediate direction between two directions based on facial images or person images of the tracked target, and registers the generated facial image or person image of the tracked target as an interpolated image in the tracked target registration DB 27 using the tracked target registration unit 36. The interpolated image generation and registration unit 81 corresponds to the "interpolated image generation and registration means" in the claims. For example, the interpolated image generation and registration unit 81 generates a facial image or person image of the tracked target captured from an intermediate direction between the two directions based on facial images or person images of the approximately front side and approximately back side of the tracked target, using a generative AI model, and registers the generated facial image or person image of the tracked target as an interpolated image in the tracked target registration DB 27 using the tracked target registration unit 36.
[0082] 15, the interpolated image generation and registration unit 81 uses images (face images or person images) of the tracked person taken from two directions (approximately the front side and approximately the back side of the tracked person) and key points in the images of the approximately front side and approximately the back side of the tracked person estimated by the two-dimensional pose estimation units 32a and 32b to generate an image (face image or person image) of the tracked person taken from a direction intermediate between the above two directions using a generative AI model. The interpolated image generation and registration unit 81 repeats the process of generating images (interpolated images) of the tracked person taken from a direction intermediate between the above two directions, thereby generating images (face images or person images) of the tracked person with various facial orientations as interpolated images, and can register them in the tracked person registration DB 27 by the tracked person registration unit 36.
[0083] In the second embodiment, the tracking target registration unit 36 registers at least a face image or a person image of the tracking target captured by the entrance cameras 3a and 3b from two directions (approximately the front side and approximately the back side of the tracking target), a face image or a person image of the tracking target captured from a direction intermediate between the two directions and generated by the interpolated image generation and registration unit 81, and an identification ID of the tracking target in the tracking target registration DB 27 of the analysis server 1. In the example shown in Fig. 15, the edge devices 2c and 2d do not have image feature amount extraction units 35b and 35c, unlike the first embodiment shown in Fig. 4, and therefore the queries sent from the edge devices 2c and 2d include not image feature amounts of the image of the tracking target but images (face images or person images) of the tracking target detected by the person detection units 31c and 31d from images captured by the in-store camera 3c and the exit camera 3d. The multi-camera tracking unit 39 of the analysis server 1 matches the image of the tracked person included in the query sent from the edge devices 2c and 2d with images of the tracked person in various facial orientations registered in the tracked person registration DB 27, thereby estimating which identification ID registered in the tracked person registration DB 27 the track obtained by the single-camera object tracking of the edge devices 2c and 2d is linked to, and assigns the estimated identification ID to the corresponding track.
[0084] 4, the edge devices 2c and 2d may be provided with image feature extraction units 35b and 35c similar to those in the first embodiment, and the queries sent from the edge devices 2c and 2d may include image feature amounts of the images of the tracked person, as in the first embodiment. In this case, the multi-camera tracking unit 39 may have a function of extracting image feature amounts from images of the tracked person with various facial poses registered in the tracked person registration DB 27, and by matching the image feature amounts of the images of the registered tracked person with various facial poses extracted by this function with the image feature amounts included in the queries sent from the edge devices 2c and 2d, it may be possible to estimate which identification ID registered in the tracked person registration DB 27 is associated with the track obtained by the single-camera object tracking of the edge devices 2c and 2d.
[0085] As described above, according to the multi-camera object tracking system 80 of the second embodiment, a facial image or a person image of the tracked person taken from an intermediate direction between two facial images or person images of the tracked person taken from two directions is generated by a generative AI model, and the generated facial image or person image of the tracked person is registered as an interpolated image in the tracked person registration DB 27. This allows images (facial images or person images) of the tracked person in various facial orientations to be generated as interpolated images and registered in the tracked person registration DB 27, so that the multi-camera tracking unit 39 can use the images of the tracked person in various facial orientations to match them with images included in a query sent from the edge devices 2c and 2d (images based on images taken from cameras other than the camera installed at the start point of multi-camera object tracking). This allows accurate linking of an identification ID to the tracked person corresponding to the query (person re-identification), thereby enabling accurate multi-camera object tracking processing.
[0086] Furthermore, according to the multi-camera object tracking system 80 of the second embodiment, an image (face image or person image) of the tracked person taken from a direction intermediate between the two directions is generated by a generative AI model using images (face images or person images) of the tracked person taken from two directions and key points of the tracked person in the two images estimated by the two-dimensional pose estimation units 32a and 32b. In this way, to generate an image of the tracked person taken from a direction intermediate between the two directions, not only the images of the tracked person taken from the two directions but also the key points of the tracked person in these two images are used, so that an image of the tracked person taken from a direction intermediate between the two directions can be accurately generated.
[0087] Next, a multi-camera object tracking system 90 according to a third embodiment of the present invention will be described with reference to Fig. 16. The multi-camera object tracking system 90 according to the third embodiment is an example of the multi-camera object tracking system according to claim 7 at the time of filing. The multi-camera object tracking system 90 according to the third embodiment differs from the multi-camera object tracking systems according to the first and second embodiments in that it estimates key points of the tracked person in images captured by all cameras in the system and performs multi-camera object tracking of the tracked person based on these key points.
[0088] 16, the tracked person registration unit 36 of the edge device 2a uses the tracking results of each tracked person by the 3D pose tracking unit 34 to perform a process of registering key points of images (face images or human images) of the approximately front and rear sides of the tracked person estimated by the 2D pose estimation units 32a and 32b, and the identification IDs of each tracked person, in the tracked person registration DB 27 of the analysis server 1. Meanwhile, in the example shown in FIG. 16, the edge devices 2c and 2d are equipped with 2D pose estimation units 91a and 91b, similar to the edge device 2a, and the queries sent from the edge devices 2c and 2d to the multi-camera tracking unit 39 of the analysis server 1 include the key points of the tracked person estimated by the 2D pose estimation units 91a and 91b. The multi-camera tracking unit 39 of the analysis server 1 then matches the key points of the tracked person contained in the query sent from the edge devices 2c and 2d with the key points of each of the tracked people registered in the tracked person registration DB 27, thereby estimating which identification ID registered in the tracked person registration DB 27 the track obtained by the single-camera object tracking of the edge devices 2c and 2d is linked to.
[0089] 16 shows an example in which the key points of the image of the tracked person estimated by the 2D pose estimation units 32a and 32b are registered in the tracked person registration DB 27 of the analysis server 1, and multi-camera object tracking of the tracked person is performed based on these key points themselves, but it is also possible to register 2D poses connecting the key points of the image of the tracked person estimated by the 2D pose estimation units 32a and 32b in the tracked person registration DB 27 of the analysis server 1, and perform multi-camera object tracking of the tracked person based on these 2D poses. In this case, the multi-camera tracking unit 39 of the analysis server 1 matches the 2D pose of the tracked person included in the query sent from the edge devices 2c and 2d with each of the 2D poses of the tracked person registered in the tracked person registration DB 27.
[0090] As described above, the multi-camera object tracking system 90 of the third embodiment estimates key points of the tracking target in images captured by all cameras (entrance cameras 3a and 3b, in-store camera 3c, and exit camera 3d), including images of the tracking target taken by entrance cameras 3a and 3b installed at the start point (entrance) of multi-camera object tracking. Based on these key points, the tracking target can be tracked across the capture ranges of all cameras. This allows key points estimated from images of at least the tracking target's approximately front and rear sides, captured by entrance cameras 3a and 3b at the start point of multi-camera object tracking, to be matched with key points in images captured by cameras installed other than the tracking start point (in-store camera 3c and exit camera 3d), enabling stable person re-identification. Therefore, the tracking target can be accurately tracked across the capture ranges of all cameras.
[0091] Variations: The present invention is not limited to the configurations of the above-described embodiments, and various modifications are possible within the scope of the invention. Next, modifications of the present invention will be described.
[0092] Variation 1: In the above embodiment, as shown in Figure 5, an example is shown in which entrance camera 3a and entrance camera 3b are positioned so that the angle between their optical axes is 110° to 140°, but the positioning of entrance camera 3a and entrance camera 3b is not limited to this, and the angle between the optical axes of these entrance cameras may be less than 110° or greater than 140°.
[0093] Variation 2: In the above embodiment, the three-dimensional pose generation unit 33 generates a three-dimensional pose of the tracked person from the key points of the tracked person estimated by the two-dimensional pose estimation units 32a and 32b and the calibration parameters of the entrance cameras 3a and 3b. However, the three-dimensional pose generation unit may directly generate a three-dimensional pose of the tracked person based on an image of the approximately front side and an image of the approximately back side of the tracked person captured synchronously by two or more cameras (more precisely, person detection images of the approximately front side and the approximately back side of the tracked person output from the person detection units 31a and 31b).
[0094] Variation 3: In the above embodiment, tracking of people within the shooting range of the two entrance cameras was performed by 3D pose tracking using 3D poses, but tracking of people within the shooting range of the two entrance cameras may also be performed by conventional (2D) single camera tracking, similar to tracking of people within the shooting range of the other cameras (in-store camera 3c and exit camera 3d).
[0095] Variation 4: In the first embodiment, the edge device 2a is provided with the tracked person registration unit 36, which registers the feature amounts (image feature amounts) of the images of the approximately back side and approximately front side of each of the tracked persons, which are synchronously photographed by the entrance cameras 3a and 3b, and the identification IDs of each of the tracked persons, in the tracked person registration DB 27 of the analysis server 1. However, the present invention is not limited to this, and a tracked person registration unit may be provided on the analysis server side, and the tracked person registration unit on the analysis server side may register the image feature amounts of the approximately back side and approximately front side of each of the tracked persons, which are sent from the edge device 2a, in the tracked person registration DB.
[0096] Variation 5: In the third embodiment, key points of the tracked person in images captured by all cameras in the system are estimated, and multi-camera object tracking of the tracked person is performed based on these key points. However, this is not limited to this. For example, multi-camera object tracking of the tracked person may be performed based on key points and image features of the tracked person in images captured by all cameras in the system. In this case, the tracked person registration unit 36 of the edge device 2a registers, in the tracked person registration DB 27 of the analysis server 1, the key points of the images of the approximately front and rear sides of the tracked person estimated by the 2D pose estimation units 32a and 32b shown in FIG. 16, the identification IDs of the tracked people, and the image features of the images of the approximately front and rear sides of the tracked person extracted by a functional block similar to the image feature extraction unit 35a shown in FIG. 4. The query sent from the edge devices 2c and 2d shown in FIG. 16 to the multi-camera tracking unit 39 of the analysis server 1 includes the key points of the tracked person estimated by the two-dimensional pose estimation units 91a and 91b shown in FIG. 16 and the image features of the tracked person extracted by functional blocks similar to the image feature extraction units 35b and 35c shown in FIG. 4. [Explanation of symbols]
[0097] 1. Analysis server (part of the "computer" in the claims) 2a Edge Device (part of the "computer" in the claims) 3 Fixed Camera (Camera) 3a Entrance camera (one of the "two or more cameras" in the claim) 3b Entrance camera (one of the "two or more cameras" in the claim) 3c In-store camera 3d exit camera 10 Multi-camera object tracking system 27 Tracking target registration database (tracking target registration database) 28 Multi-camera object tracking program (part of the "multi-camera object tracking program" in the claims) 32a, 32b 2D pose estimation unit (2D pose estimation means) 33 3D pose generation unit (3D pose generation means) 34 3D pose tracking unit (3D pose tracking means) 36 Tracking target registration unit (tracking target registration means) 80 Multi-camera object tracking system 81 Interpolated image generation and registration unit (interpolated image generation and registration means) 90 Multi-camera object tracking system Entrance
Claims
1. In a multi-camera object tracking system that performs multi-camera object tracking to track a target across the shooting ranges of multiple cameras, two or more cameras among the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked object; a three-dimensional pose generation means for generating a three-dimensional pose of the tracking target in a three-dimensional space based on at least an image of a substantially front side and an image of a substantially rear side of the tracking target captured synchronously by the two or more cameras; a three-dimensional pose tracking means for assigning an identification ID to each of the tracking targets corresponding to the generated three-dimensional poses, and tracking each of the tracking targets assigned with the identification ID using the generated three-dimensional pose in a video frame consisting of images taken by two or more cameras installed at the starting point; a tracking target registration means for registering, in a tracking target registration database, at least the image of the substantially front side and the image of the substantially back side of the tracking target photographed synchronously by the two or more cameras, or image feature quantities of these images, and an identification ID of the tracking target; A multi-camera object tracking system that tracks the person to be tracked across the shooting ranges of the multiple cameras based on at least an image of the approximately front side and an image of the approximately back side of the person to be tracked that are registered in the database for registering people to be tracked, or on the image features of these images.
2. The three-dimensional pose generation means a two-dimensional pose estimation means for estimating key points of the tracked object in at least an image of a substantially front side and an image of a substantially rear side of the tracked object captured synchronously by the two or more cameras; 2. The multi-camera object tracking system according to claim 1, wherein a three-dimensional pose of the tracked person is generated from the key points of the tracked person estimated by the two-dimensional pose estimation means and calibration parameters of the two or more cameras.
3. The multi-camera object tracking system of claim 1, wherein the starting point of the multi-camera object tracking is an entrance to a building or room, and two or more of the multiple cameras are positioned so as to capture at least an image of the approximately front side and an image of the approximately back side of a tracked person passing through the entrance.
4. A multi-camera object tracking system that performs multi-camera object tracking to track a target across the shooting ranges of multiple cameras, two or more cameras among the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked object; a tracking target registration means for registering, in a tracking target registration database, at least the image of the substantially front side and the image of the substantially back side of the tracking target photographed synchronously by the two or more cameras, and an identification ID of the tracking target; the tracking target registration means registers a head image or a person image of the tracking target in the tracking target registration database; Further provided is an interpolated image generating and registering means for generating an image of the head or person of the tracking target taken from an intermediate direction between the two directions based on an image of the head or person of the tracking target taken from two directions using a generation AI model, and registering the generated image of the head or person of the tracking target as an interpolated image in the tracking target registration database by the tracking target registration means, A multi-camera object tracking system that tracks the person to be tracked across the shooting ranges of the multiple cameras based on at least an image of the head or an image of the person to be tracked approximately from the front and approximately from the back, which are registered in the database for registering people to be tracked, and an image of the head or an image of the person to be tracked taken from a direction intermediate between the two directions generated.
5. The method further includes a two-dimensional pose estimation means for estimating key points of the tracked person in the head images or person images of the tracked person taken from the two directions, The multi-camera object tracking system of claim 4, wherein the interpolated image generation and registration means uses an image of the head or image of the person being tracked taken from the two directions and key points of the person being tracked estimated by the two-dimensional pose estimation means to generate an image of the head or image of the person being tracked taken from a direction intermediate between the two directions using a generative AI model.
6. A multi-camera object tracking program for performing multi-camera object tracking in which a tracking target person is tracked across the shooting ranges of multiple cameras, comprising: When two or more of the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked object, Computer, a three-dimensional pose generation means for generating a three-dimensional pose of the tracking target in a three-dimensional space based on at least an image of a substantially front side and an image of a substantially rear side of the tracking target captured synchronously by the two or more cameras; a three-dimensional pose tracking means for assigning an identification ID to each of the tracking targets corresponding to the generated three-dimensional poses, and tracking each of the tracking targets assigned with the identification ID in a video frame consisting of images captured by two or more cameras installed at the starting point; functioning as a tracking target registration means for registering, in a tracking target registration database, the image of at least the substantially front side and the image of the substantially back side of the tracking target photographed synchronously by the two or more cameras, or image feature quantities of these images, and an identification ID of the tracking target; A multi-camera object tracking program for tracking a person to be tracked across the shooting ranges of the multiple cameras based on at least an image of the approximately front side and an image of the approximately back side of the person to be tracked registered in the database for registering people to be tracked, or image features of these images.
7. A multi-camera object tracking program for performing multi-camera object tracking in which a tracking target person is tracked across the shooting ranges of multiple cameras, comprising: When two or more of the plurality of cameras are installed at a start point of the multi-camera object tracking and are arranged so as to capture at least an image of a substantially front side and an image of a substantially rear side of the tracked object, Computer, functioning as a tracking target registration means for registering, in a tracking target registration database, the image of at least the substantially front side and the image of the substantially back side of the tracking target photographed synchronously by the two or more cameras, and an identification ID of the tracking target; the tracking target registration means registers a head image or a person image of the tracking target in the tracking target registration database; The computer is further made to function as an interpolated image generating and registering means that generates, based on the head image or person image of the tracking target person taken from two directions, an image of the head or person image of the tracking target person taken from an intermediate direction between the two directions using a generation AI model, and registers the generated head image or person image of the tracking target person as an interpolated image in the tracking target person registration database using the tracking target person registration means; A multi-camera object tracking program for tracking a person to be tracked across the shooting ranges of the multiple cameras based on at least an image of the head or an image of the person to be tracked approximately from the front and rear sides registered in the database for registering people to be tracked, and an image of the head or an image of the person to be tracked taken from a direction intermediate between the two directions generated.
Citation Information
Patent Citations
Moving object tracing device
JP2009223434A
Image processing apparatus, image processing method, and program
JP2025076600A
Tracking apparatus, tracking system, tracking method, and recording medium
WO2022091166A1