Tracking device, tracking method, and program
The tracking device and method address irregular pedestrian movements and occlusions by clustering and re-dividing segments using LiDAR data, ensuring accurate tracking in crowded conditions with improved metrics.
Patent Information
- Application Number
- JP2023220429
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-07-09
AI Technical Summary
Existing pedestrian tracking methods using 3D point clouds struggle with accurate tracking when pedestrians move irregularly or occlude each other, leading to incomplete point clouds and merging of segments, which disrupts continuous tracking.
A tracking device and method that utilize point cloud data from multiple LiDARs to cluster segments based on density, estimate the number of people in merged segments using a CNN model, and apply Kalman filtering to maintain tracking accuracy by re-dividing segments when inconsistencies occur, combined with trajectory repair using diffusion models for non-overlapping LiDAR arrangements.
Enables robust and accurate pedestrian tracking even in crowded environments with irregular movements by maintaining tracking performance through timely re-clustering and segment division, improving metrics like MOTA to 0.977 in dense pedestrian scenarios.
Smart Images

Figure 2025103217000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for tracking a pedestrian by using point cloud data of the pedestrian acquired by a distance measuring device having an observation area.
Background Art
[0002] In recent years, in spaces such as public facilities and commercial facilities where various people come and go, the demand for detecting the flow of people has been increasing in order to respond to sudden accidents and to understand people's positions and actions to improve services. By the way, due to the remarkable development of image processing technology, in recent years, there are many systems and methods for detecting people from RGB moving images and the like. However, since they basically directly acquire information including personal information such as faces, it is difficult to eliminate the anxiety about the privacy of passers-by even when they are discarded immediately after use. In addition, since it is necessary to respond to opt-out requests from passers-by and a lot of work is required such as obtaining sufficient explanations and consents that most people can accept, there are many obstacles to implementing flow measurement by images in public spaces and semi-public spaces.
[0003] On the other hand, a method of detecting the presence and posture of a person with a lower privacy risk using a three-dimensional distance sensor such as a three-dimensional ranging sensor (LiDAR) or a depth camera that acquires only distance information to an object has attracted attention. The three-dimensional ranging sensor measures the distance to the nearest object in each direction from the sensor by using a method of irradiating an infrared pattern and camera parallax or by ToF (Time of Flight) measurement of an infrared beam to form a three-dimensional point cloud.
[0004] The inventors proposed a method for pedestrian tracking (trajectory derivation) in the downtown area of a city using 3D point cloud data captured by multiple LiDARs (Non-Patent Document 1). Most of the person detection methods using 3D point clouds aim to detect the posture of a person at a fixed position, and basically do not assume situations where pedestrians obstruct each other's view or situations where pedestrians approach each other and point clouds of multiple people are observed simultaneously. On the other hand, in the detection of the flow of people in public spaces, due to occlusion between people and sunlight, etc., the point cloud information capturing a person is often incomplete. In the method of Non-Patent Document 1, considering the detection (clustering) failure of person segments due to the incompleteness of the point cloud, such as missing point clouds and combination of point clouds due to the approach of multiple people, the estimation of two situations where point clouds of multiple pedestrians are combined and observed as a single segment and where point clouds of a single person are observed as multiple segments, and tracking by a Kalman filter are combined to achieve more robust tracking.
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] However, in the method of Non-Patent Document 1, since the movement prediction is performed on the premise that pedestrians move straight forward at a constant speed when point clouds of multiple pedestrians are combined (merged) and observed, there is a problem that it becomes difficult to continue accurate tracking when irregular movement is performed while the combined point cloud is being observed.
[0007] The present invention has been made in view of the above, and provides a tracking device, a tracking method, and a program that enable more accurate tracking to be continued by estimating the number of people in a segment in a timely manner and re - performing clustering.
Means for Solving the Problems
[0008] The tracking device according to the present invention captures point - group data for measurement periodically acquired by at least one distance - measuring device having an observation area, performs clustering to divide into segments for each pedestrian based on the point - group density of the point - group data, and a first clustering means for acquiring the observation position thereof, and uses the observation position of the pedestrians in each segment acquired this time and the predicted position this time calculated using the previous predicted position in the object for each pedestrian to be tracked, and associates the segment and the object having the minimum distance, and a tracking means for performing a tracking process using a filter on the associated object. The tracking means includes a number - estimating means for estimating the number of people included in the segment when it is determined that the association between the segment and the object is inconsistent, and a second clustering means for dividing the segment into the estimated number of parts.
[0009] In addition, the tracking method according to the present invention includes a computer capturing measurement point cloud data periodically acquired by at least one distance measuring device having an observation area, clustering to divide into segments for each pedestrian based on the point cloud density of the point cloud data, a first clustering step of acquiring the observation position thereof, and calculating the current prediction position using the observation position of the pedestrian in each segment acquired this time and the previous prediction position in the object for each pedestrian to be tracked, associating the segment and the object that minimize the current prediction position, and a tracking step of performing tracking processing using a filter on the associated object. The tracking step includes a number estimation step of estimating the number of people included in the segment when it is determined that the association between the segment and the object is inconsistent, and a second clustering step of dividing the segment into the estimated number of people.
[0010] In addition, the program according to the present invention is for causing a computer to function as the tracking device.
[0011] According to these inventions, the first clustering means captures periodic measurement point cloud data from the measuring device, and performs clustering to divide it into segments for each pedestrian based on the point cloud density of the point cloud data, and a process of obtaining the observation position thereof. Next, the tracking means uses the observation position of the pedestrian in each segment acquired this time and the current predicted position calculated using the previous predicted position in the object for each pedestrian to be tracked, and associates the segment and the object with the minimum distance, and a tracking process using a filter, for example, a Kalman filter, is performed on the associated object. Further, when the number estimation means determines that the association between the segment and the object is inconsistent, the number of people included in the segment is estimated, and the second clustering means performs a process of dividing the segment into the estimated number of people. Therefore, when a plurality of pedestrians approach and merge into one segment, and it is determined that the association between the segment and the object is inconsistent, for example, a learned model is used to estimate the number of merged people in the segment and re-divide it into a preferable number of segments, so that robust tracking becomes possible. In particular, reliability is obtained by maintaining the tracking performance even when passing through a crowded place.
[0012] Further, the number estimation means determines the above-mentioned inconsistency when it can be considered that a plurality of pedestrians have merged in the segment. According to this configuration, by taking immediate measures to eliminate the cause of tracking interruption, robust tracking is maintained.
[0013] Further, the number estimation means determines the above-mentioned inconsistency when the predicted position of the object to be associated matches the position information of the segment where the association has ended. According to this configuration, it becomes possible to obtain the occurrence of merging using the position information.
[0014] Further, the number estimation means is a learned CNN model. According to this configuration, the process for number estimation becomes easy.
[0015] Further, the CNN model decomposes the three-dimensional point cloud data into a plurality of two-dimensional planes including at least a horizontal plane, estimates the number of people on each plane, and outputs the estimated number of people from each plane with weights attached. According to this configuration, processing can be performed at high speed.
[0016] Also, the second clustering means performs clustering by the K-Means method. According to this configuration, by adopting the K-Means method, it is easy to perform clustering and calculate the centroid (center) position of the point cloud within each cluster.
[0017] Also, it is provided with the tracking device and a distance measuring device that outputs the acquired point cloud data to the first clustering means. According to this configuration, it can be provided as a tracking device including a distance measuring device.
[0018] Further, the tracking device according to the present invention includes the trajectory repair means, the distance measuring device has a plurality of them, the observation regions for obtaining respective partial trajectories are dispersed, and the trajectory repair means includes a reference information creation means for creating information corresponding to four criteria: similarity of feature amounts of the point cloud of each pedestrian, spatial consistency of the partial trajectories of the pedestrian, temporal consistency of the same trajectory, and trajectory joining using a diffusion model, a determination means for re-identifying the pedestrian based on the created information, and a trajectory estimation means for estimating a trajectory within an unobserved region based on the determination result. According to this configuration, even in a mode where a plurality of distance measuring devices are arranged with unobserved regions in between so that the observation regions for obtaining respective partial trajectories are dispersed to perform efficient and power-saving tracking, it is possible to re-identify the pedestrian by repairing the walking trajectory in the unobserved region.
Effect of the Invention
[0019] According to the present invention, accurate tracking can be continued even when point clouds of a plurality of pedestrians are observed as merged.
Brief Description of the Drawings
[0020]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Best Mode for Carrying Out the Invention
[0021] FIG. 1 is a functional configuration diagram showing an embodiment of a tracking device 1 according to the present invention. The tracking device 1 executes tracking processing on, for example, a pedestrian who is an observation target, and includes a computer including a CPU (processor). The processor functions as a control unit 10. A storage unit 100 that stores data necessary for the tracking process is connected to the control unit 10. Further, a display unit 21 having a screen for displaying an image and an operation unit 22 for giving an instruction from the outside are connected to the control unit 10. The display unit 21 confirms the instruction content and displays the processing information. The operation unit 22 may include a known touch panel laminated on the screen of the display unit 21 and / or a keyboard and a mouse.
[0022] The storage unit 100 includes at least a control program memory 101 that stores an operation program and an arithmetic program for executing the tracking process, a model parameter memory 102 that stores learned model parameters for estimating the number of people obtained by machine learning using AI (artificial intelligence), and a ranging data memory 103 that stores periodically acquired ranging data and background point cloud data obtained in advance, etc., and also has a work area and a memory area as necessary.
[0023] In this embodiment, the distance measurement unit 30 employs a three-dimensional ranging sensor (hereinafter referred to as LiDAR (Light Detection and Ranging) 31) that acquires only the distance information to an object within the observation area with a lower privacy risk. The distance measurement unit 30 includes a required number, in this embodiment, a plurality of LiDARs 31. The LiDAR 31 is arranged at a predetermined height (e.g., 3 m) from, for example, the floor surface or using a wall surface, etc., at an appropriate position and orientation so as to surround the observation area. In this embodiment, the respective orientations are set so that a part of the observation area overlaps without interruption. The observation area is, for example, a public passage or a square where pedestrians come and go, and the arrangement position and orientation of the LiDAR 31 are set so as to overlook from diagonally above. As is well known, the LiDAR 31 emits a pulsed laser beam at high speed while scanning it in the horizontal direction, and receives laser reflection pulses from pedestrians and background objects as a point cloud. Also, the LiDAR 31 scans in the vertical direction and enables transmission and reception in the depression angle direction. As a result, the distance measurement unit 30 repeatedly acquires the laser reflection pulses as three-dimensional point cloud data from the observation area at a frame period. As the performance of the LiDAR 31, for example, an infrared laser wavelength, a point cloud number of hundreds of thousands of points / second, and a frame rate (frame period) of 10 frames / second are applicable. Note that the distance measurement unit 30 may be another distance measurement device as long as it can acquire a three-dimensional point cloud. The three-dimensional point cloud data for each frame acquired by the distance measurement unit 30 is sequentially output to the control unit 10.
[0024] The control unit 10 executes tracking processing of the pedestrian to be observed from the three-dimensional point cloud acquired for each frame, and stores the result in the storage unit 100 as necessary. The control unit 10 functions as a distance measurement data acquisition unit 11, a foreground point cloud extraction unit 12, a clustering unit 13, a tracking unit 14, and an optional trajectory repair unit 15 by reading and executing the control program stored in the control program memory 101 of the storage unit 100. Note that the trajectory repair unit 15 will be described later in relation to step S17.
[0025] The ranging data acquisition unit 11 instructs the ranging unit 30 to emit the ranging laser light in synchronization with the scanning operation for each frame, and captures the laser light that has returned from the reflecting object and been received as a three-dimensional point cloud, and performs predetermined preprocessing. The preprocessing is to integrate the three-dimensional point clouds from all the LiDARs 31 to generate an overall three-dimensional point cloud. More specifically, the ranging data acquisition unit 11 constructs an absolute coordinate system for the target observation space in advance, and based on the three-dimensional positions and orientations of each LiDAR 31, converts and integrates the coordinates of each three-dimensional point cloud obtained respectively onto the absolute coordinate system.
[0026] In the integrated point cloud, in the region where the fields of view of multiple LiDARs 31 overlap, the point clouds of the same object from each LiDAR 31 may overlap. Therefore, compared with the region where the fields of view of the LiDARs 31 do not overlap, the density of the obtained point cloud becomes higher. This affects the density-based point cloud clustering described later. Therefore, the ranging data acquisition unit 11 performs a downsampling process of collecting multiple overlapping points within each voxel grid into one point (for example, the center of the grid) for the entire obtained integrated point cloud, such as by using a voxel grid filter. Note that the size of the cube-shaped voxel grid can be appropriately set based on the measurement distance accuracy of the LiDAR 31 used.
[0027] The foreground point cloud extraction unit 12 extracts the foreground point cloud including the point cloud formed by pedestrians using background subtraction. The foreground point cloud extraction unit 12 acquires, as the background point cloud, the point cloud when there are no pedestrians in the observation area for each LiDAR 31 in advance, stores it in the same absolute coordinate system, and uses it in the background subtraction process. That is, after the start of tracking, when a pedestrian enters the observation area with the LiDAR 31, a difference occurs between the acquired point cloud and the background point cloud. The foreground point cloud extraction unit 12 obtains only the point cloud of the pedestrian (foreground point cloud) by extracting the point cloud of the difference.
[0028] The clustering unit 13 applies clustering to the integrated foreground point cloud and divides it into a plurality of different pedestrian segments. The clustering unit 13 may compress the three-dimensional point cloud into a two-dimensional point cloud composed of the xy plane by removing the vertical axis in order to shorten the processing time required for clustering as necessary. This is because pedestrians existing in the public space move on a two-dimensional plane and the things they possess and use, so there is no need to consider the overlap of multiple people in the vertical direction.
[0029] In this embodiment, the clustering unit 13 applies the Euclidean clustering method, which is clustering based on point cloud density, to the point cloud converted into two dimensions composed of the xy plane to extract pedestrian segments. Here, the pedestrian segments are referred to as pedestrian segments. In addition to the point cloud, the clustering unit 13 may also display each segment together with a rectangular bounding box (a combination of the minimum and maximum values for each axis of the coordinates of each point in the segment) along the axes of the spatial coordinates.
[0030] Next, tracking will be described. Tracking includes various processes, which are executed by the tracking unit 14. First, the extracted pedestrian segments are observed, and the pedestrians who are observed and become the targets of tracking are called pedestrian objects. Hereinafter, they may simply be referred to as segments and objects. Here, tracking is performed by correcting, updating, and predicting the state of the object using a Kalman filter. Each object is represented as a four-term set consisting of the xy coordinate components of the object's position coordinates and the vector component parallel to the xy axis among the object's velocity vectors as the state of the Kalman filter.
[0031] Here, let the set of observations (pedestrian) segments detected in the k-th frame be O k ={o k,j}. j indicates the number of the observed segment. Also, let the set of objects during tracking after the processing of the k-th frame is completed be T k ={t k,iLet it be so. i indicates the number of the object being tracked. T k The object t in k,i is corrected by the Kalman filter. Hereinafter, for the set of segments O k each segment in it will simply be called segment o, and for the set of objects T k the object in it will sometimes simply be called object t. In the following description, during the processing of the k-th frame, the speed of the object newly added to the object T k shall be a predetermined value, for example, 1.0 (m / s) per second in both the x and y directions. For each object t in the object T k the predicted value (predicted position) of that object ≧ in the (k + 1)-th frame predicted from it is p k,i and the set thereof is P k+1、i Let it be so. Note that T0 and P1, P k+1 at the start of frame processing, T k are all empty sets, and it is assumed that k ≧ 1. k
[0032] Subsequently, regarding the same person determination, first, the processing when the k-th frame is issued will be described. At the start of frame processing, based on the set of objects T k-1 the predicted value P k of the object in the current (k-th) frame is obtained in advance. If the predicted value P k is empty (an empty set), T k ← O k , that is, all of the segments O k observed in the current frame are made the objects of tracking, and the processing of this frame is terminated. Then, wait for the next frame. This processing corresponds to steps S9 and S11 in FIG. 5.
[0033] When the tracking unit 14 has a non-empty set of objects T k-1 it performs the following three processes (1) to (3).
[0034] (1) The segment set O k of the current frame and the object set T k-1 Determine the association with [object] so as to minimize the "cost" based on the distance.
[0035] (2) For the object t and segment o for which the association was not given in (1), determine the cause and determine the corresponding association accordingly.
[0036] (3) Using the association information determined in (1) and (2), perform Kalman gain and state update (Kalman filter processing).
[0037] First, regarding the association in (1), for the given pair of segment o and object t, find the one that minimizes the sum of the "costs" calculated from the pair of object and segment. Note that the details of the "cost" calculation will be described later. This process corresponds to steps S21 and S23 in FIG. 6.
[0038] Here, for the number of segments (denoted as |O k |) in the current frame k and the number of objects (denoted as |T k-1 |) in the previous frame (k - 1), perform the association that satisfies the following conditions (a) to (d).
[0039] (a) Each segment is associated with at most one object.
[0040] (b) Each object is associated with at most one segment.
[0041] (c) There are only min{|O k |,|T k-1 |} associations.
[0042] (d) The total cost of the association is minimized.
[0043] In this way, ensure that all of the smaller number of objects or segments have associations.
[0044] Next, in (2), in order to determine the update method of the Kalman filter in (3), the cause of the occurrence of an object t or a segment o without association is judged. For example, the object t or the segment o changes due to occlusion (shielding), disappearance due to exiting the observation area of the pedestrian, or appearance due to entering the observation area of the pedestrian. However, by associating the cause of their occurrence with the occurrence position in the observation area and making a judgment, the purpose is to improve the accuracy of the Kalman filter in (3). For example, when disappearing or appearing at the edge of the observation area, it can be judged that entry and exit to the observation area have occurred, and the occurrence and disappearance other than at the edge of the frame can be judged to be caused by occlusion or the like. This process corresponds to step S15 in FIG. 5.
[0045] Finally, in (3), based on the obtained association, the object t in frame k is predicted using the Kalman filter. The observation includes all the points included in the segment o of the current frame k, and is represented by a six-tuple consisting of the central coordinate components of a rectangular parallelepiped (bounding box) whose sides are parallel to any of the xyz axes and the sizes of the three sides of the segment. Using these, for example, for a certain object t, even if the corresponding segment o is not found, if it is estimated that the segment has disappeared due to occlusion, tracking is continued. For the object t without association in (2), the Kalman filter is updated without an observation value. Assume that the segment o has moved at the speed of the object T k and update the Kalman filter. For the object t with an association with the segment o, the current position is updated using the Kalman filter.
[0046] Next, the details of the calculation example of "cost" will be described. For each pair of the object t in the previous frame (k - 1) and the segment o in the current frame k, their distance is set as the "cost" of that pair, and an association that minimizes the cost while satisfying the above conditions is obtained. Object T k,i and segment O k,j The cost c i,j for is defined by Equation (1).
[0047] c i,j=√( x j- x i) 2 +( y j- y i) 2 ···(1) x j , y j represent the xy coordinates of segment O k,j respectively, and x i , y i represent the xy coordinates of the predicted value p k,i of object T k,i respectively. The assignment is determined in order from the combination that minimizes this cost c i,j . If the cost of the determined combination exceeds a threshold value (set to a predetermined value, for example, 1.0 based on experiments and experience), the assignment is rejected.
[0048] Then, for the assigned pair of the pedestrian in segment O k and object T k , using the observed segment and the Kalman filter, update the current position of the pedestrian in the corresponding T k-1 and add it to T k .
[0049] Next, the processing of the unassociated object or segment shown in the above processing (2) will be described.
[0050] If there are objects or segments that are not included in the association determined based on the policy of (2) above, the following "additional association method" is applied to find their associations. The "additional association method" determines the cause of the lack of association and provides the most appropriate association according to the cause.
[0051] As the reasons why a certain segment o does not have a corresponding object t, the following three (S1) to (S3) can be considered. This processing corresponds to steps S31 to S35 in FIG. 7.
[0052] (S1) A new pedestrian enters the observation area to obtain a segment, but there is no object in the previous frame that can be associated with it.
[0053] (S2) Due to the influence of occlusion, a segment split into l (el) segments for one pedestrian is obtained, but in the previous frame, there is only one object corresponding to them (that is, no association can be obtained for the remaining (l - 1) segments).
[0054] (S3) A segment is obtained because a pedestrian not observed due to the influence of occlusion is observed again, but there is no object in the previous frame that can be associated with it.
[0055] First, in the cases of (S1) and (S3), since it can be considered that a new pedestrian has appeared, position prediction by the Kalman filter may be started from the next frame and subsequent frames. Therefore, the Kalman filter is not applied to this segment, and this segment is added to T k Next, in the case of (S2), when what was recognized as one object t i in frame (k - 1) is observed as l (el) segments due to occlusion in the current frame k (when the observation "splits"). If this is assumed to be the reason, the tracking of this object t i is terminated at the time of this split, and thereafter, the segments generated by the split are tracked respectively.
[0056] Also, there are two reasons, (O1) and (O2) below, why an object t does not have a corresponding segment o. This process corresponds to steps S41 to S45 in FIG. 5.
[0057] (O1) The pedestrian recognized as an object in frame (k - 1) exits the observation area, resulting in the loss of the corresponding segment and no association can be obtained.
[0058] (O2) When multiple people approach and pass by, only one segment can be obtained for two or more pedestrians.
[0059] First, in the case of (O1), since it can be considered that the pedestrian has exited the observation area, the Kalman filter for this segment is not applied. Also, an object that has not been associated with a segment for more than a predetermined number of frames F thr ends tracking. In the case of (O2), when n objects t l (l = 1, 2, 3, ···, n) recognized in frame (k - 1) are observed as one segment due to occlusion or the like in the current frame (when the observations "merge"). The determination that the observations have merged is when the predicted position and the observed position of the object overlap. If this is assumed to be the reason, the segments are processed in the order of (G1) to (G3) below.
[0060] Figure 2 is a diagram for explaining the number prediction (G1), position prediction (G2), and association with tracking (G3) corresponding to such (G1) to (G3) of the merged segment. Note that Figure 2(G1), Figure 2(G2), and Figure 2(G3) are diagrams for convenience of explanation, and there is no direct relationship between the point cloud images of Figure 2(G1) and Figure 2(G2) and the association diagram of Figure 2(G3).
[0061] (G1) Estimate the number of people in the merged segment using a learned, for example, convolutional neural network (CNN) model. When the predicted positions of multiple segments and objects overlap, perform number estimation for all overlapping segments.
[0062] (G2) For the corresponding segment to the estimated number of people, perform, for example, K-Means clustering and estimate the center position of each.
[0063] (G3) The center positions of the clusters (sub-segments) divided by K-Means clustering and the predicted value (position) pk,tl Associate them and apply the Kalman filter.
[0064] First, (G1) is executed by the number-of-people estimation unit 141. In (G1), for the combined segment o, the number of people included in the segment o is predicted using a convolutional neural network (CNN) model. In this number-of-people prediction, as shown in FIG. 2 (G1), 2D CNNs (xy 2D CNN, yz 2D CNN, zx 2D CNN) are constructed from the viewpoints of the top surface and the side surfaces of the segment o, respectively, and the number-of-people estimation results for each surface are combined by adding them with predetermined weights, and the final number-of-people prediction is performed (FIG. 3).
[0065] FIG. 3 is a diagram showing the overall image of the CNN model used. The input point cloud on the left side is a synthetic image acquired by LiDER31, and the images from the viewpoints of each surface (xy plane, yz plane, zx plane) are obtained by replacing the left synthetic image with the images of each viewpoint. The 2D CNN uses those that have been pre-trained for each surface. The training data for the 2D CNN uses, for example, those that contain about 5,000 to 6,000 combined segments of 1 person to a predetermined number of people, for example, 8 people each.
[0066] FIG. 4 shows the structure of the 2D CNN model used in this embodiment. The 2D CNN model uses two convolutional layers and then two fully connected layers, and outputs the probability for each number of people as to whether it is any of 1 to 8 people. After each of the two convolutional layers, a ReLU layer and a Max Pooling layer are used, and a ReLU layer is also used after the first fully connected layer. For the input of this 2D CNN model, an array of 100×100 is created from the point cloud of the viewpoint of each surface. The point cloud of the viewpoint of each surface is divided into 5 cm×5 cm square cells, and an array of 100×100 is created by setting the height of the point with the largest height among the point clouds included in each cell as the value of that cell. Then, the probability that the number of people included in the segment is from 1 to 8 people is calculated, and the output is such that the total is 100%. And as shown in FIG. 3, the number-of-people estimation results P for the top surface and the side surfaces respectivelyxy , P yz , P zx are weighted and added together to synthesize the final estimated number of people est xyz .
[0067] (G2) is executed by the clustering unit 142. In (G2), for the corresponding segment where the number estimated in (G1) is obtained, that is, the segment where merging has occurred, for example, by performing K-Means clustering, the point cloud within the segment is divided again. Note that K-Means clustering is an algorithm that classifies by finding the centroid of the data. This process corresponds to steps S47 and S49 in FIG. 8. By performing this clustering, the central position c m (m = 1, 2, 3, ···, n) of the people included in the sub-segment is estimated. In the drawing on the right side of FIG. 2 (G2), the point cloud of multiple people is decomposed, and an × mark is attached to each calculated centroid position (central position).
[0068] In (G3), the center position c m of the cluster is associated with the predicted value p k,tl of the object, and the Kalman filter is updated in the same way as for the objects for which normal association has been successful. This association is also the same as normal association, and the center position c k,tl with the closest distance to the predicted value p m is associated (see the × marks shown in FIG. 2 (G3)). For the associated object, while assuming that the size of the object is the same as the previous observation, the Kalman filter is updated assuming that it has moved to the associated center position. If the distance at the time of association is greater than or equal to the threshold (set to 0.55 based on previous experimental results or experience), the association fails, and the overlapping area of the predicted value p l of each object t k,tl and the observation o k,j is generated as a virtual segment (sub-segment), and this virtual segment is used as the object t lIt is associated as a segment. Then, the virtual segment is pseudo-observed, and the Kalman filter is updated. Note that the association between the object t and the segment o is appropriately memorized in relation, or linked using, for example, an ID code or the like. Thus, each process for executing the tracking according to the present embodiment has been described.
[0069] Subsequently, an example of the procedure of the tracking process will be described using the flowcharts shown in FIGS. 5 to 8. First, in FIG. 5, the control unit 10 determines whether or not the tracking has ended (step S1). When it is determined that the tracking has ended upon receiving an end instruction from the operation unit 22, the process proceeds to the "trajectory repair" process (described later) in step S17 described later. On the other hand, if it is not the end, the distance measurement data acquisition unit 11 continuously performs distance measurement data acquisition (point cloud synthesis) processing periodically (every time (T), for example, at a frame rate of 10 frames / second) (step S3). Note that the tracking end instruction may be in a mode of entering by interruption. Next, the foreground point cloud extraction unit 12 executes a downsampling process based on the acquired distance measurement data and a process of extracting a foreground point cloud including a point cloud based on a pedestrian using background difference (step S5).
[0070] Next, the clustering unit 13 obtains a set O of segments by applying, for example, Euclidean clustering based on the density of the point cloud acquired in the current frame corresponding to the time (T) (step S7). Subsequently, the tracking unit 14 determines whether or not an object was detected at time (T - 1), that is, in the previous frame (step S9). If it was detected, the process proceeds to step S15. On the other hand, when the tracking unit 14 determines that an object was not detected at time (T - 1), that is, in the previous frame, each segment o of the set O of segments obtained in step S7 is set as an object at time (T) (step S11), and the time is advanced to (T + 1) in step S13, and then the process returns to step S1.
[0071] In step S15, the tracking unit 14 executes the case discrimination process for tracking. That is, in the set T of objects obtained at time (T-1) and the set S of segments obtained at time (T), if it is determined that the set T of objects is empty and the set O of segments is not empty, the process proceeds to step S31. If it is determined that the set T of objects is not empty and the set O of segments is empty, the process proceeds to step S41. If it is determined that both the set T of objects and the set O of segments are empty, the process returns to step S13. And if it is determined that both the set T of objects and the set O of segments are not empty, the process proceeds to step S21 as normal processing.
[0072] First, in step S21, as shown in FIG. 6, if it is not the first time, a pair of an object t and a segment o with the minimum "cost" of distance is selected from the set T and the set O and associated (i.e., linked). Next, using the Kalman filter, the predicted value of the associated object t is estimated and set as the position of the object t at time (T), and the object t and the segment o are deleted from the groups T and O (step S23). Then the process returns to step S15.
[0073] Next, in step S31, as shown in FIG. 7, the tracking unit 14 determines whether each segment o in the group O is a sub-segment added in the CNN process of step S47 described later. If it is a sub-segment, the sub-segment is discarded (step S33), and the process returns to step S13. On the other hand, if it is not a sub-segment, after setting it as the object at time (T) (step S35), the process returns to step S13.
[0074] Next, in step S41, as shown in FIG. 8, the tracking unit 14 calculates a predicted position at a time (T) linearly predicted from the object t. Next, the tracking unit 14 determines whether or not the calculated predicted position overlaps with the position of a certain segment o that has already been associated in step S21 (step S43). If there is an overlap, since the segment o may include two or more pedestrians, the point cloud of the segment o is two-dimensionally transformed from a plurality of viewpoints on the top surface and the side surface, and the number of people m included in the segment o is estimated using a CNN model (step S45). On the other hand, if there is no overlap (No in step S43), the process returns to step S15.
[0075] Subsequently, in step S47, the tracking unit 14 uses the estimated number of people m of the segment o as a parameter, applies K-Means clustering to the point cloud of the segment o, and divides the segment o into m sub-segments. Further, the central position in the xy plane of each sub-segment is set as the sub-segment position. Next, the tracking unit 14 adds the sub-segment o to the set O (step S49), and returns to step S15.
[0076] Subsequently, Experiments 1 and 2 for accuracy evaluation will be described.
[0077] In Experiment 1, 3D point clouds were collected from a plurality of fixedly installed 3D LiDARs 31, and the performance of this tracking method when pedestrians are dense was evaluated. In Experiment 2, data reproducing a situation where a plurality of dense pedestrians move irregularly was collected for accuracy evaluation. In Experiment 2, actual traffic data was collected and accuracy evaluation was performed during a time period when the density of pedestrians is high within the campus of this university. Hereinafter, Experiment 1 will be described with reference to FIG. 9, and Experiment 2 will be described with reference to FIGS. 10 to 12.
[0078] <Experiment 1> Accuracy Evaluation for Nonlinear Movement at the Time of Segment Merging (1) Experimental Environment and Collected Data The experimental environment was set up by installing two 3D LiDARs 31 at positions that were approximately diagonal to each other so that a flat area within the university campus with dimensions of about 10m × about 10m served as the observation area (data acquisition area). In this observation area, the movements of four pedestrians (subjects) were collected as 3D point clouds. Figure 9(A) is a measurement screen diagram explaining the experimental environment, showing the arrangement of the two 3D LiDARs 31 and a perspective view of the acquired 3D point cloud. The LiDAR 31 (Livox Avia: manufactured by LIVOX) used in Experiment 1 has the performance shown in Table 1. Each LiDAR 31 was fixed at a height of approximately 3m from the floor. By integrating the foreground point clouds from the point cloud data acquired by each LiDAR 31 and applying background subtraction, an integrated foreground point cloud for the entire space was collected.
[0079] [Table 1]
[0080] The situation of Experiment 1 is shown in Figure 9(B). Note that the movements of the four pedestrians are indicated by arrows, representing one pedestrian in the upper left. That is, first, the four pedestrians walk from the four corners of the observation area towards the center and gather at the center (indicated by arrow (1) in the figure). Next, they perform a 270° rotational movement in a clockwise direction while in the gathered state (indicated by arrow (2) in the figure). Then, after the rotational movement, they walk towards the four corners to disperse (indicated by arrow (3) in the figure). This action was repeated six times.
[0081] (2) Evaluation Metrics To evaluate the tracking performance of multiple people, the following evaluation metric MOTA among the CLEAR metrics was adopted. MOTA represents the accuracy of person detection and ID assignment in multi-person tracking and is calculated by Equation (2).
[0082] MOTA = 1 - Σ k (FP k + FN k + ID SWk ) / ΣGT k ···(2) However, FP in Equation (2) k , FNk , ID SWk , GT k represents the number of false detections, the number of detection failures (misses), the number of ID swaps, and the number of true people in frame k, respectively. Also, FP represents the number of false detections over all pedestrians and all periods. FN represents the number of detection misses over all pedestrians and all periods. IDsw represents the number of ID swaps over all pedestrians and all periods. Frag represents the number of trajectories split due to misses corresponding to FN over all pedestrians and all periods.
[0083] (3) Evaluation Results Accuracy evaluation was performed in two cases for the collected data of about 90 seconds: when performing the number prediction and position prediction of the merged segment (the proposed method) and when not performing them. In the method where the number prediction and position prediction were not performed, when the segments were merged, the area where the predicted position of the pedestrian overlapped with the observation was pseudo - used as the pedestrian observation, and the Kalman filter was updated. The results are shown in Table 2.
[0084]
Table 2
[0085] By performing the number prediction and position prediction, ID SWThe number of times decreased from 14 to 3, and Frag decreased significantly from 49 to 7. It was found that long-term tracking was possible with the same ID. This is considered to be due to the fact that the predicted position at the time of combination made it possible to capture the observed position of the pedestrian more accurately, making it easier to perform tracking corresponding to complex movements. The MOTA index indicating the overall tracking performance was 0.855 when no person number prediction or position prediction was performed, but it was 0.977 with the proposed method. The factor causing this difference is that FN decreased significantly. The situation where FN frequently occurred was a scene where multiple pedestrians were densely packed, the segments merged, and multiple pedestrians were recognized as one pedestrian. In this scene, it is determined that none of the pedestrians recognized as merged except one can be accurately detected, so FN increased significantly. When the accuracy of tracking is improved by the proposed method, the predicted position can be accurately updated by the Kalman filter even when a movement other than a linear movement occurs, and correct association can be performed even when the segments are merged. As a result, accurate position prediction and association could be performed even for the merged segments, which is considered to be the factor.
[0086] <Experiment 2> Tracking in an actual crowded environment (1) Experimental environment and collected data The three-dimensional point clouds of pedestrians actually passing through the university were collected using four three-dimensional LiDARs 31 installed in an observation area on flat ground in a range of about 20m × about 10m within the university. The same three-dimensional LiDAR 31 as in Experiment 1 was used. Each LiDAR 31 was fixed at a height of about 3m from the floor. The data collected from each LiDAR 31 was applied with background difference, and the integrated foreground point cloud of the entire space was collected by integrating the foreground point clouds.
[0087] The situation of Experiment 2 is shown in Fig. 10. Fig. 10 is a perspective view that shows the environment of Experiment 2 from an overhead perspective, including the arrangement of the 3D LiDAR, the acquired 3D point cloud, and the campus facilities. In Fig. 10, the upper part is the entrance to the campus, and the elevator and stairs for moving to the upper floors are arranged on the lower right side. Pedestrians passing through this observation area where data was collected mostly move back and forth between the entrance to the campus and the elevator and stairs. Note that the four rectangular frames in the figure indicate the locations where each 3D LiDAR 31 is arranged, and each arrow indicates the observation direction. The other multiple point clouds in the figure are mainly foreground point clouds including pedestrians. During peak pedestrian hours, there may be a queue waiting for the elevator, and there was also a scene where another pedestrian crossed the queue.
[0088] The tracking performance evaluation in Experiment 2 is performed for the scenes where there are many pedestrians and particularly where there are many encounters with pedestrians, among the data collected under this situation. For the baseline method, similar to the evaluation results of Experiment 1, when the segments are combined, the area where the predicted position of the pedestrian overlaps with the observation is pseudo-used as the observation of the pedestrian, and the Kalman filter is updated. Also, the tracking performance except for the moment when pedestrians are close is not much different from the baseline. Moreover, since the number of pedestrians is large, when using the CLEAR metric shown in the evaluation metric of Experiment 1 as the evaluation metric, no significant difference was observed between the baseline and the proposed method. Therefore, two cases where a significant difference was observed in the tracking between the baseline and the proposed method at the moment when multiple pedestrians are close are described below.
[0089] (2) Tracking Results The cases where significant differences were observed are designated as Case 1 and Case 2, and are shown in FIGS. 11 and 12. FIG. 11 shows Tracking Case 1 in a scenario where actual close tracking occurred. The tracking at the baseline for comparison is represented by (a1)-(a3), and the tracking by the proposed method is represented by (b1)-(b3). FIG. 12 shows Tracking Case 2 in a scenario where actual close tracking occurred. The tracking at the baseline for comparison is represented by (c1)-(c3), and the tracking by the proposed method is represented by (d1)-(d3).
[0090] In these figures, the images represented by rectangles show the pedestrians being tracked, and the curves extending from each rectangular image represent the trajectories followed by those pedestrians. Also, the pedestrians surrounded by circular marks are the pedestrians for whom the merging of segments was observed. The pedestrians surrounded by the upper two circular marks in screen (a3) of FIG. 11 and the two circular marks in screen (c3) of FIG. 12 have failed tracking. Next, the tracking status of the pedestrians surrounded by circular marks will be described below.
[0091] In Case 1, one pedestrian (P a ) was passing from the stairs below the elevator towards the campus entrance above, and one pedestrian (p b ) was passing in the opposite direction. This passing occurred on the side of the column that was in front of the elevator. These two pedestrians approached the two people who were forming the column at the same time, and for an instant, four pedestrians were within a close distance of each other (FIGS. 11(a1)-(a2)). At the baseline, after the two pedestrians passed this location (FIG. 11(a2)), the tracking transitioned to the upper two circular marks in (a3), and it can be seen that the tracking of the pedestrian (P a ) who was passing from the stairs below towards the campus entrance above has at least failed ((P a ’) reference) (FIG. 11(a3)). On the other hand, with the proposed method, in this scenario, the pedestrians (P a ),(p b) It can be seen that the tracking was successful even after the close following (Fig. 11(b3)). This is attributed to the fact that the effect of position estimation is exerted, and tracking can be performed while more accurately capturing the position even during segment combination.
[0092] Case 2 is a scene where, immediately after Fig. 11, pedestrians (p c ) walking from the staircase towards the entrance and pedestrians (p d ) walking from the entrance towards the staircase pass by each other below the column in front of the elevator (see Fig. 12(c1) to (d1)). Pedestrians (p d ) walking from behind the column in front of the elevator towards the staircase and pedestrians (p c ) walking from the staircase towards the entrance passed by each other at a short distance below the column (Fig. 12(c2)). Due to the narrowness of the place, the moving direction changed clockwise / counterclockwise at the moment of passing by. In the baseline, after the two pedestrians passed by, it was unable to cope with the change in the tracking direction, and the IDs of these two people were swapped after passing by ((p c ’) and (p d ’) in Fig. 12(c3)). On the other hand, with the proposed method, by position estimation, the tracking direction can be appropriately changed even during the occurrence of passing by, and the IDs after passing by can be maintained ((p c ) and (p d ) in Fig. 12(d3)).
[0093] <Summary of Experiments 1 and 2> The 3D point clouds obtained from multiple LiDARs 31 were used to extract the foreground point clouds, and the 3D point cloud segments representing pedestrians were extracted by clustering. While predicting the merging and splitting of the segments representing multiple pedestrians, the positions and velocities of each pedestrian were predicted using a Kalman filter for tracking. When the merging of segments was predicted, the number of people was estimated using a 2D CNN model pre-trained for the upper and side surfaces of the segments, respectively, and the central positions of the pedestrians within the segments were calculated by K-Means clustering. Then, the tracking accuracy was improved by reassigning the central positions and the observed pedestrians. At the university, multiple 3D LiDARs were installed, and data on multiple densely packed people walking irregularly were collected. As a result of evaluating the tracking accuracy for approximately 90 seconds of this data, MOTA reached 0.977. Also, as a result of evaluating the actual pedestrians collected during a time period with a high density of passers-by within the university, it was shown that pedestrian tracking can be realized with high accuracy even in a congested public space where observations merge and irregular movements occur.
[0094] <Other Embodiments> Next, as another embodiment, a tracking technique that enables human trajectory reconstruction using point cloud feature quantities and a diffusion model will be described.
[0095] Covering human tracking over a wide area using a large number of LiDARs is not practical from the viewpoints of cost and power supply. Therefore, after dispersing and arranging multiple LiDARs, it is desired to enable human tracking in the same way as MCT (multi-camera tracking) in cameras. However, MCT using LiDAR, where it is difficult to obtain feature quantities for individual identification, has many problems in a distributed arrangement type.
[0096] The following describes a person tracking technology that uses multiple LiDARs with non-overlapping observation areas in the target area. The proposed method discovers the correspondence relationships of the partial trajectories of multiple persons obtained by each LiDAR and estimates the overall trajectories of each person within the target area. Therefore, in the proposed method, persons are re-identified based on four criteria: the similarity of the point cloud feature amounts of each person, the spatial consistency of the partial trajectories of the persons, the temporal consistency of the same trajectories, and trajectory joining (matching) using a diffusion model, and their complete trajectories are estimated.
[0097] The point cloud feature amounts are realized by extracting discriminative fixed-size feature amounts representing the shape and behavior of each point cloud using Fisher vectors applied to linear discriminant analysis. Also, in order to learn the dissimilarity of the features of different persons, training of a deep distance learning model is performed. Regarding temporal and spatial consistency, the start and end points of the partial trajectories are utilized to learn the possible transitions of the persons. This is realized by defining the probability distributions of the transition time and transition pattern and updating the distributions using a Bayesian approach. Also, as shown in FIG. 13, for trajectory joining using a diffusion model, which is a type of deep learning model for generating images, an image (the left figure) in which a part of the trajectory is missing as a mask area is redrawn using the diffusion model to draw a natural connection, and based on the drawing result, it is statistically determined which trajectories are matched.
[0098] To demonstrate the effectiveness of the proposed method, a verification experiment was conducted using 3D point cloud data collected from multiple LiDARs installed in a six-story building at the university. As a result, a high reconfiguration accuracy of the person trajectory was achieved with an F value of 0.91.
[0099] <A. System Architecture and Problem Definition> <A-1. Acquisition of Partial Trajectories> Figure 14 is a perspective view for explaining partial trajectories obtained from distributed LiDERs in another embodiment. Each LiDAR can capture a part of the target three-dimensional space, and a three-dimensional point cloud can be acquired in each frame from the data stream. For example, the three-dimensional space in which the point cloud is generated by LiDARi is called the scan space of LiDARi, and is denoted as S i 3 Next, the background subtraction method is applied to the point cloud to extract the moving object. The extracted point cloud is called the foreground point cloud, and segmentation is applied to find each person included in the foreground point cloud. Using a voxel grid filter, downsampling is applied to the foreground point cloud to convert each voxel grid cell into one virtual point. Next, a clustering algorithm is applied to the foreground point cloud to divide it into person segments. When applying clustering, the z-axis is removed to reduce processing overhead, and DBSCAN-based clustering is adopted to obtain the person segments. Note that DBSCAN is a type of clustering algorithm that groups high-density regions as the same group and treats low density as outliers.
[0100] The partial trajectory is, for example, a two-dimensional trajectory obtained as a temporal sequence of the (x, y) coordinates of the person segment corresponding to the movement trajectory of one person in the scan space S i 3 of LiDARi. For simplicity, the two-dimensional region in which the partial trajectory is obtained by LiDARi is called the trajectory region of LiDARi, and is denoted as S i 2 Generally, S i 2 is the projection of S i 3 onto the xy plane.
[0101] Here, let TR i [t,t’] and H i t represent the set of partial trajectories obtained by LiDARi in the time window [t, t'] and the set of person segments at time t, respectively. Also, tr ij [ta,tb]Let it be assumed that the partial trajectory j obtained by LiDAR i starts at time ta and ends at tb. For any partial trajectory tr ij [ta,tb] ∈TR i [t,t’] note that t ≤ ta ≤ tb ≤ t’. tr ij [ta,tb] is in S i 2 and its start and end points are on the boundary of S i 2 .
[0102] At time t, a set H of person segments is obtained from LiDAR i i t Assume that a set TR of partial trajectories for k - 1 time windows (k > 1) has been obtained. Next, find the correspondence between the partial trajectory tr ∈ TR i [t-k,t-1] and the person segment h ∈ H i [t-k,t-1] . This correspondence can be easily found by applying a Kalman filter-based tracking method, and an updated TR i t-1 can be obtained. Finally, define the temporal relationship between two partial trajectories tr i [t-k,t] and tr i,u [ta,tb] . Let tr j,v [tc,td] < tr b < t c if and only if T i,u [ta,tb] < tr j,v [tc,td] .
[0103] <A-2. Problem Definition> Here, consider a time window, for example, [t s , t e . Assume that a set L of LiDARs, their scan spaces and trajectory regions, and a map of the area where those LiDARs are installed are given. Also, assume that the partial trajectory set TR = ∪ i∈L TR i[ts,te] It is assumed that it has been obtained.
[0104] The trajectory estimation problem is formulated as a problem of finding the division of TR, and each division can form a partial trajectory sequence. For example, if tr1 = tr i,u [1,3] and tr2 = tr j,v [4,6] , tr3 = tr j,r [10,12] are included, then tr1 < tr2 < tr3 is satisfied, and the partial trajectory sequence tr1:tr2:tr3 can be constructed. <A-3. Trajectory Estimation Algorithm> For a given TR = ∪ i∈L TR i [ts,te] the present trajectory estimation algorithm operates as follows.
[0105] Prepare sets V1 and V2 of partial trajectories and a set E of pairs of partial trajectories, and initially set them all to be empty. V1 and V2 respectively correspond to sets of partial trajectories whose endpoints are connection points. Then, for a pair of partial trajectories tr u , tr v ∈TR, if tr u < tr v , then add tr u and tr v , (tr u , tr v ) to V1, V2, and E respectively. Also, calculate the affinity (∈[0, 1]) of the pairs defined and explained later, and create a weight function W: E → [0, 1].
[0106] Finally, obtain a weighted bipartite graph G = (V1 ∪ V2, E, W). Due to the way V1 and V2 are created, |V1| = |V2| holds, so the problem is to find the optimal one-to-one match between V1 and V2 on E. This corresponds to finding a subset E' of E that maximizes the sum of affinities, as shown in Equation (3).
[0107]
Equation
[0108] <B. Proposal Method> Here, for each pair of partial trajectories (tr u , tr v ), the probability that they are the movement trajectories of the same pedestrian (referred to as affinity) is defined. To calculate the affinity, four points are used: (i) the similarity of the point clouds of the two person segments of tr u and tr v , (ii) the statistical spatial feature (transition frequency) from the end point of tr u to the start point of tr v , (iii) the statistical temporal feature (travel time) from the end point of tr u to the start point of tr v , and (iv) the matching result using the diffusion model. The corresponding probabilities are denoted as P1, P2, P3, and P4 respectively, and all are within the range of [0, 1]. The affinity (denoted as A(tr1, tr2)) is the product of the above probabilities as shown in Equation (4).
[0109]
Number
[0110] Hereinafter, how the probabilities P1, P2, P3, and P4 are calculated will be explained. In addition, since the likelihood distributions of the probabilities P2 and P3 are based on statistics, that is, the prior distribution, they are updated by the Bayesian approach. This update will also be explained hereinafter.
[0111] <B-1. Similarity of Person Segments> The similarity of two person segments is defined, and by calculating it, it is determined whether a set of person segments belongs to the same person. Point cloud data is usually unordered and unstructured, and the number of points in the segments is also different. Therefore, simply adopting a learning-based similarity scheme is generally insufficient. Thus, feature extraction using Fisher vectors and similarity calculation using deep metrology learning are devised to address this problem.
[0112] <B-1-1. Feature Extraction> First, to extract the fixed-size representation of the input person segment, the Fisher vector is adopted. Specifically, the deviation from the Gaussian mixture model of the three-dimensional point cloud is calculated. The intuitive reason for using the Fisher vector for feature extraction is that it can capture the spatial formation of 3D points in space, so that the discriminative features of the person segment can be obtained. This can be realized by calculating the gradient of the log-likelihood of the sample with respect to the parameters of the GMM (Gaussian mixture distribution) model. The extracted feature representation of the Fisher vector has a fixed size that does not depend on the number of points in the point cloud of the person segment. Due to this advantage, it becomes easy to process person segments of variable sizes using learning-based similarity techniques.
[0113] Formally speaking, X i ={p t ∈R 3 ,t = 1, ···, T} represents the three-dimensional point cloud of person segment i. Here, T indicates the number of points in the segment, which may vary greatly depending on different factors such as the resolution and range of LiDAR, the scene, and the distance. Each point is assumed to come from one of C different groups representing body parts such as the head, arms, and legs, and the groups are Gaussian distributions. And the set of parameters λ of the C-component GMM is λ = {(w c , μ c , Σ c ), c = 1, ···, C}, where w c , μ c , Σ c are the weight, mean, and covariance matrix of the distribution of c th respectively. Different Gaussians are predefined and placed on the 3D grid with equal weights and standard deviations. The likelihood that a single three-dimensional point belongs to the Gaussian of c th is given by the following equation (5).
[0114]
Equation
[0115] The likelihood of a point belonging to the GMM is defined as in the following equation (6).
[0116]
Number
[0117] Given a specific GMM, under the assumption of common independence, the Fisher vector G λ X can be written as the sum of the normalized gradient statistics calculated for each point p t as shown in Equation (7).
[0118]
Number
[0119] Here, L λ is the square root of the inverse Fisher information vector. Change the variable from w c to α c to ensure that u λ (x) is a valid distribution and simplify the calculation of the gradient.
[0120]
Number
[0121] Therefore, the normalized gradient can be written as in the following Equations (9), (10), and (11).
[0122]
Number
[0123] Concatenating all these components forms the Fisher vector as shown in Equation (12).
[0124]
Number
[0125] To avoid the variation in the three-dimensional point counts in each segment, the resulting Fisher vector is normalized with a sample size \(T\) as in Equation (13).
[0126]
Number
[0127] Furthermore, the Fisher vector guarantees that the extracted features are invariant to the permutation of the input by using a symmetric function. Specifically, the Fisher vector calculates the sum of gradients, which is a symmetric function. Therefore, by referring to "max-pooling", an additional symmetric function including the minimum and maximum values is calculated to extend the basic Fisher vector (see Equation (14)). This results in a more descriptive and permutation-invariant representation. As a result, each person segment is mapped to a \(20\times54\) feature matrix. Here, 54 is the number of Gaussians and 20 is the number of features.
[0128]
Number
[0129] <B-1-2. Calculation of Similarity between Feature Quantities> To calculate the similarity between the extracted person shape feature quantities, deep distance learning is used. Here, deep distance learning using "Triplet loss" is performed. "Triplet loss" increases the distance between samples of different classes and decreases the distance between samples of the same class. Also, cosine similarity is used as the distance function.
[0130] The overall flow is as shown in FIG. 15. FIG. 15 is an explanatory diagram showing the processing procedure for calculating the similarity of two person segments. More specifically, in FIG. 15, two person segments, one above and the other below at the left end, are input, and Fisher vectors are calculated for each. At this time, by adopting the processing referring to "max-pooling" as described above, each person segment is mapped to a 20×54 feature matrix. The result is input into a trained neural network (trained NN) to obtain the coordinates in the embedding space. The output is the cosine similarity cos(x,y) of the obtained coordinates x and y. The cosine similarity is usually in the range of [-1,1], but for consistency with other feature quantities described later, it is converted to the range of [0,1]. Therefore, the probability P1 is defined as in Equation (15).
[0131] [Number]
[0132] Here, h i and h j are person segments in the two sub-trajectories of interest, respectively.
[0133] <B-2. Movement Frequency between LiDAR Areas>[ The spatial features of the two sub-trajectories represent how frequently similar transitions occurred in the past. For this reason, attention is paid to S i 2 which is the boundary of the trajectory area of each LiDARi.
[0134] At the boundary of the trajectory area S i 2 of each LiDARi, there are parts called virtual gates where pedestrians are likely to enter and exit. The virtual gates are the boundaries of the trajectory areas S i 2 such as those of doors and corridors. Each sub-trajectory needs to have its start and end points belonging to the virtual gates. Then, a transition matrix Q is constructed, where each element is the transition probability from one virtual gate to another.
[0135] tr v When there is a destination virtual gate g2 to which the starting point of tr belongs, the partial trajectory tr u and tr v The spatial feature probability P2 of tr u is defined as the ratio to the sum of the probabilities from the virtual gate g1 to which the end point of tr belongs to other gates g2. P2 is defined by the following formula (16).
[0136]
Number
[0137] <B-3. Travel Time between LiDAR Areas> Similarly, the temporal features of two partial trajectories are defined to represent the likelihood of travel time from one virtual gate to another. tr iu [ta,tb] and tr jv [tc,td] are assumed to represent two partial trajectories. The travel time is obtained as Δt = t iu [ta,tb] <tr jv [tc,td] at that time. Assuming a prior probability density function p c -t b (x), the probability can be obtained as in the following formula (17). time (x), the probability can be obtained as in the following formula (17).
[0138]
Number
[0139] <B-4. Distribution Update of Spatiotemporal Features> The update methods of the transition matrix Q (spatial feature) and the probability density function p time (x) (temporal feature) will be explained. The former is performed by a histogram, and the latter is based on the Bayesian approach.
[0140] These functions need to be updated with high reliability during operation. An example of a good pattern is a case where there is only one pedestrian moving from one end point to another starting point and no other pedestrians are observed. This phenomenon may occur in a scene with few people. The transition matrix can be easily updated with high reliability based on the distribution of recorded past transitions. In the case of Bayesian updating of the moving time probability distribution, the likelihood of the moving time P(E|H) in such a high-reliability case is similar to a normal distribution. Therefore, the prior distribution is updated using the Bayesian formula shown in Equation (18).
[0141]
Number
[0142] Here, both the prior distribution P(H) of the transition time and the created posterior distribution P(H|E) are calculated as inverse gamma distributions.
[0143] <B-5. Matching Using the Diffusion Model> By converting the trajectory data into a 2D image and training it with the diffusion model for image generation, it is possible to generate trajectories in the invisible region. The plausibility of the connection between sub-trajectories is obtained from the number of generated trajectory connection patterns, and this is used as an index for trajectory connection. Here, the open-source "Stable Diffusion" released by Stability AI is used as the base model for the diffusion model that performs image generation. Also, "Hypernetworks" are used for the training of trajectory images. When introducing the diffusion model, "Stable Diffusion WebUI" is used. In "Stable Diffusion WebUI", it is possible to easily adjust the parameters of the model, and the generated images can be visualized at once, so the data can be evaluated efficiently.
[0144] "Inpainting" of "Stable Diffusion" is used to generate the trajectory pattern of the LiDAR invisible area. This draws an image of the content along the prompt with reference to the image around the mask, and by repairing the missing part of the trajectory, it outputs a trajectory pattern of the invisible area that is naturally connected.
[0145] In the generation, "DPM2++ 2S a" was used as the sampling method. This is because the arrangement of features is uniform and dense, so smooth lines can be realized and it is suitable for drawing trajectories. In the additional learning of the trajectory pattern, 18 images with three trajectories drawn were learned, added to the model as a hypernetwork, and fine-tuning was performed. The hypernetwork originally adds a small neural network to the "Stable Diffusion" model, making it possible to change the style of the image. Therefore, by creating a hypernetwork, as shown in Fig. 16, the painting style originally suitable for paintings can be shifted to an image having only trajectory data. Fig. 16 is an exemplary four-frame image diagram generated from the same text "art by trajectory". (A) is the case without a hypernetwork, and (B) is the case with a hypernetwork. In Fig. 16(A), an artistic image is output by the original "Stable Diffusion" model, while in Fig. 16(B), an image with a style similar to the trajectory images used in the learning is output. Also, during the drawing, it is possible to prevent unrelated objects and characters from being mixed in.
[0146] <C. Evaluation> Next, the proposed system is evaluated using the created test bed. 70 LiDARs were installed on the floors from the 1st floor (1F) to the 6th floor (6F) of the campus building of the university, and an experiment using a part of them was conducted.
[0147] <C-1. System Specification, Environment, Dataset>
[0148]
Table 3
[0149] Table 3 shows the specifications of the LiDAR used in the experiments. Using the above-mentioned equipment, two types of experiments were conducted. Specifically, (i) performance evaluation using a high-density and narrow FoV LiDAR in the 1F entrance hall for comparison with the RGB camera-based method (Experiment-1), and (ii) basic performance evaluation in an intentionally synthesized scenario (2F: indoor square, Experiment-2) were carried out.
[0150] In each experiment, almost the entire floor was photographed by the installed LiDAR, but some point clouds were intentionally removed, and these areas were regarded as outside the scan area to evaluate this method. More detailed configurations will be described below.
[0151] <C-2. Data collection> The statistical values of the obtained dataset are shown in Table 4.
[0152]
Table 4
[0153] In Experiment-1, four LiDARs were used to observe pedestrians in the 1F entrance hall. Fig. 17 is a photographic screen diagram showing the experimental environment, where (A) shows the 1F entrance hall (Experiment-1) and (B) shows the 2F indoor square (Experiment-2). Area Z1 marked in (A) indicates the erased space. Also, in the figure, only three LiDARs, A, B, and C, can be seen. From September 10th to September 13th, 2022 (for 4 days), complete trajectories of a total of 2,356 observations of the number of observed people (see Table 4) were observed.
[0154] In Experiment-2, 32 general subjects with different genders and ages (from their 20s to 50s) (see Table 4) were recruited. Each subject was asked to walk along a designated route. As shown in (B), the area Z2 indicated by the rectangle is 4m × 7m. Also, there are two LiDARs, namely A and B. For each subject, about 1000 frames (= 100 seconds) of 3D point cloud data were collected. In this experiment, the LiDARs are not blocked by other subjects. Therefore, clear partial trajectories can be obtained, and the performance of matching using the partial trajectories can be evaluated.
[0155] <C-3. Experiment-1> In Experiment-1, pedestrians were tracked in a real environment over four days. As shown in Table 2, due to the uncongested environment, the accuracy of the proposed method was confirmed to be as high as an F-value of 0.91.
[0156] <C-3-1. Accuracy Comparison with RGB Method> A person re-identification system using an RGB camera was compared with the proposed method. The LiDAR on the first floor was installed at the same position and orientation as the surveillance camera considering the building structure. Also, the installed LiDAR has a FoV relatively close to that of the camera. Since it is difficult to implement an RGB-based system here, a more accurate re-identification function (see Liang Zheng, Yi Yang, and Alexander G Hauptmann. Person re-identification: Past, present and future. arXiv preprint arXiv:1610.02984, 2016.) was used instead of point cloud-based re-identification.
[0157] On the other hand, when using a camera system, tracking with a single camera is less accurate than single LiDAR-based tracking, so it is usually difficult to obtain transition patterns. As a result, the proposed method can utilize point cloud features (P1), transitions (P2), movement time (P3), and matching in the diffusion model (P4), while the camera-based method uses color-based features (P1 +) and the moving time (P3) can be utilized. Table 5 shows the results. The proposed method makes a comprehensive judgment from four types of features and achieves sufficient accuracy. On the other hand, RGB can achieve higher accuracy by leveraging the power of image-based features. The proposed method can achieve the same level of accuracy as the RGB-based method while realizing better data privacy compared to the camera-based re-identification method.
[0158]
Table 5
[0159] <C-3-2. Evaluation of Matching Accuracy Using Diffusion Model> The area Z1 shown in the trajectory image was masked using "Stable Diffusion" and redrawn, and the results are shown in Fig. 13. In this example, the matching was successful for all six generated images. However, among them, only the result in the lower left is connected to an incorrect trajectory when compared with the image before generation. As patterns where the matching fails, there were many cases where the trajectory was interrupted in the middle or the trajectory was branched. Also, the area for redrawing was set as two rectangular areas of 5m × 4m.
[0160] Regarding the test data with two trajectories to be connected, eight trajectory data were created in each of the two areas, and 20 images were generated for each trajectory data. Therefore, a total of 320 images were generated. Regarding the data with three trajectories, since it was difficult to generate the original image in which three trajectories were beautifully drawn in a narrow space, four trajectory data were created for one area. Therefore, a total of 160 images were generated.
[0161] For two areas, the generated images after image restoration were performed with two and three trajectories respectively, and are shown in Figure 18. In this example, correct matching was performed for all images. Figure 18 is an exemplary diagram showing the results of image restoration by "Stable Diffusion". The upper part shows the image before restoration, and the lower part shows the image after restoration. The figure is a trajectory image obtained by restoring the part covered by the square mask. The restored trajectory images on the lower side are the best ones selected from the 20 generated images respectively.
[0162] As an index for evaluating the accuracy, the matching rate at which matching was successful including the case where two trajectories were incorrect, and the correct rate which is the ratio of correct matching among the matched pairs were calculated (see Table 6).
[0163]
Table 6
[0164] From these results, it was found that when the number of trajectories to be connected is two, correct matching can be performed with higher accuracy than when it is three.
[0165] <C-4. Experiment-2> Thirty-two subjects walked while following the same route. Then, the point clouds of multiple subjects were synthesized to generate multiple scenarios with different numbers of subjects and timings.
[0166] In Experiment-2, the change in accuracy due to the congestion situation (Scenario 2-(a)), the change in accuracy due to the change in the synthesis delay time (Scenario 2-(b)), and the performance measurement for each feature (Scenario 2-(c)) were comprehensively examined.
[0167] <C-4-1. Created Scenarios> In Scenario 2-(a), different numbers of subjects were synthesized with an interval from subsequent subjects by delaying the walking start time of subsequent subjects by 10 seconds.
[0168] In Scenario 2-(b), the delay time was changed. This scenario mainly aims to evaluate the effect of the travel time distribution. That is, the shorter the delay time, the more difficult it is to distinguish the travel time.
[0169] In Scenario 2-(c), in order to evaluate the contribution of each feature, matching is performed using each of the four features one by one.
[0170] In all scenarios, the matching performance before and after the update of the spatial and temporal features (transition matrix Q and travel time probability distribution p time (x)) is compared. As initial values, the probabilities of Q and p time (x) are all uniform (assuming a certain range for p time (x)), and both are updated using the observed values obtained through all scenarios.
[0171] <C-4-2. Results: Scenario 2-(a)> Figure 19 is a diagram showing the change in the F-value. (A) shows the change in the F-value for each number of subjects in Scenario 2-(a), and (B) shows the change in the F-value at different time intervals in Scenario 2-(b).
[0172] Figure 19(A) shows the F-value for each number of subjects, indicating the before and after update of the transition matrix and the travel time distribution respectively. Since the initial distribution is completely uniform, the accuracy with 32 subjects is about 0.6, but it has improved significantly after the update. For normal walking speeds, 4 to 8 subjects generate an appropriate density. Looking at that value, the F-value after the update is 0.8 to 0.9, which is a sufficiently high value. From the above observations, the case of 4 subjects was selected in Scenario 2-(b).
[0173] <C-4-3. Results: Scenario 2-(b)> Figure 19(B) shows the F values for each delay time interval. As the interval is lengthened, high accuracy is obtained at 10 seconds, resulting in a very natural outcome. Also, as the interval becomes longer, the accuracy decreases, indicating that 10 seconds is optimal as the interval for distinguishing movement time and patterns.
[0174] <C-4-4. Results: Scenario 2-(c)> The F values for the case of only one feature quantity of P1, P2, P3, and P4 are evaluated, and the results are shown in Table 7.
[0175]
Table 7
[0176] The evaluation of matching P4 using the diffusion model is the same as the above <c-3-2>The results shown in item were used. In the case of four feature quantities, the F value became higher than when each feature quantity was used alone, indicating that the combination of feature quantities was effective.
[0177] <C-4-5. Results: Effect of Distribution Update>
[0178]
Table 8
[0179] Table 8 shows the accuracy before and after the update of the distribution, and the F value after the update has improved to 0.89. When moving on the diagonal of the map, the state of the distribution update is shown in Fig. 20(B). The prior distribution is a curve (1) with a large variance, and the likelihood is a curve (2) with a high reliability of the distribution calculated from the movement. The posterior distribution is curve (3), which is the result of Bayesian update.
[0180] <C-4-6. Performance of Deep Metric Learning> The basic performance of deep metric learning was investigated for 32 subjects. The model was trained with 90% of the randomly selected data and tested with the remaining 10%. The average cosine similarity of two person segments included in the test data was calculated, and the results are shown as the matrix in Fig. 21(A). In Fig. 21(A), generally, the regions not surrounded by frames (frames) have higher similarity as the density increases (+ direction), while the regions surrounded by frames have lower similarity as the density increases (- direction). More specifically, along the diagonal line from the upper left to the lower right of the figure, a region of cells with high density in the plus direction can be seen, which means that the similarity search is correct. However, it can be seen that there are also cells with high density in the plus direction, that is, high similarity, in some of the cells in the regions other than the diagonal line. To quantify this result, Table 9 shows the average value and standard deviation of the similarity between the same person and different persons on the diagonal line. From the results in Table 9, it can be seen that when the subjects are the same, the similarity values are concentrated at high values, and when they are different, the similarity values are widely dispersed. Also, as shown in Fig. 21(B), the ROC (Rate of Change) curve had an AUC (Area Under the Curve) of 0.71.
[0181]
Table 9
[0182] Here, as another embodiment, a method for tracking a person's trajectory using a plurality of LiDARs with non-overlapping observation regions was proposed and verified. Two types of experiments were conducted using the data collected in a large-scale test bed with 70 LiDARs. As a result, it was confirmed that the person's trajectory can be accurately tracked with an F value of 0.91 even in non-overlapping regions.
[0183] Finally, the subroutine of "trajectory repair" in step S17 of FIG. 5 will be described. Step S17 executes the trajectory repair processing method described in <B. Proposed Method>. The tracking device 1 shown in FIG. 1 has been described by taking an example in which the observation areas of a plurality of LiDARs 31 partially overlap. However, the observation areas of the plurality of LiDARs 31 may be in a dispersed mode via non-observation areas in addition to a partially overlapping mode. In the case of the latter dispersed mode, it is desirable to estimate the trajectory of a pedestrian in a segmented area.
[0184] Therefore, in order to execute the process of step S17, the trajectory repair unit 15 shown in FIG. 1 creates information corresponding to four criteria: the similarity of the point cloud feature amounts of each pedestrian, the spatial consistency of the partial trajectories of the pedestrian, the temporal consistency of the same trajectory, and the trajectory joining (matching) using a diffusion model. The criterion information creation unit 151, the determination unit 152 that re-identifies (determines) the pedestrian based on the created information, and the trajectory estimation unit 153 that estimates the complete trajectory based on the determination result. That is, the determination unit 152 and the trajectory estimation unit 153 determine whether the trajectory of the same pedestrian object is segmented or is a trajectory by another pedestrian object. When it is determined that the trajectory is segmented, they are joined and repaired into one trajectory. For example, in the storage unit 100, a control program for trajectory repair, a learned model, area information regarding a mask area (non-observation area), and other necessary data are stored so as to be readable.
[0185] Also, for the tracking process, a filter such as a particle filter, which performs a process similar to a Kalman filter that updates the state at the next time from the prediction and observation of the state, may be applied.
Explanation of Signs
[0186] 1 Tracking device 10 Control unit 13 Clustering unit (first clustering means) 14 Tracking unit (tracking means) 141 Number estimation unit (number estimation means) 142 Tracking unit (second tracking means) 15 Trajectory repair unit (trajectory repair means) 151 Reference information creation unit (reference information creation means) 152 Determination unit (determination means) 153 Trajectory estimation unit (trajectory estimation means) 100 Memory unit 30 Distance measurement unit 31 LiDER
Claims
1. Importing measurement point cloud data periodically acquired by at least one distance measuring device having an observation area, clustering to divide into segments for each pedestrian based on the point cloud density of the point cloud data, and first clustering means for acquiring the observation position thereof; Tracking means for associating the observation position of the pedestrian in each segment acquired this time with the segment and object having the minimum distance from the current predicted position calculated using the previous predicted position in the object for each pedestrian to be tracked, and performing tracking processing using a filter on the associated object; The tracking means is When it is determined that the association between the segment and the object is inconsistent, number estimation means for estimating the number of people included in the segment; A tracking device comprising second clustering means for dividing the segment into the estimated number of people.
2. The tracking device according to claim 1, wherein the number estimation means determines the inconsistency when it can be considered that a plurality of pedestrians have merged in the segment.
3. The tracking device according to claim 2, wherein the number estimation means determines the inconsistency when the predicted position of the object to be associated matches the position information of the segment where the association has ended.
4. The tracking device according to claim 1, wherein the number estimation means is a learned CNN model.
5. The CNN model according to claim 4 decomposes 3D point cloud data into a plurality of 2D planes including at least a horizontal plane, estimates the number of people on each plane, and outputs by weighting the estimated number of people from each plane. Tracking device.
6. The tracking device according to claim 1, wherein the second clustering means performs clustering by the K-Means method.
7. A tracking device comprising the tracking device according to claim 1 and a distance measuring device for outputting the acquired point cloud data to the first clustering means.
8. Equipped with trajectory repair means, The distance measuring device has a plurality of them, and the observation areas for obtaining respective partial trajectories are dispersed, The trajectory repair means is Criterion information creation means for creating information corresponding to four criteria: similarity of feature amounts of point clouds of each pedestrian, spatial consistency of partial trajectories of pedestrians, temporal consistency of the same trajectory, and trajectory joining using a diffusion model; Determination means for re-identifying pedestrians based on the created information; The tracking device according to claim 1, comprising trajectory estimation means for estimating a trajectory within an unobserved area based on a determination result.
9. A computer captures point cloud data for measurement periodically acquired by at least one distance measuring device having an observation area, performs clustering to divide into segments for each pedestrian based on the point cloud density of the point cloud data, and a first clustering step of acquiring the observation position thereof; associates a segment and an object such that the observation position of the pedestrian in each segment acquired this time and the predicted position this time calculated using the previous predicted position in the object for each pedestrian to be tracked are at a minimum distance, and performs a tracking process using a filter on the associated object, and a tracking step; The tracking step includes a number estimation step of estimating the number of people included in the segment and a second clustering step of dividing the segment into the estimated number of people when it is determined that the association between the segment and the object is inconsistent.
10. A program for causing a computer to function as the tracking device according to any one of claims 1 to 8.