Moving body tracking system

The moving object tracking system addresses errors in complex object tracking by identifying detection accuracy reduction positions and correcting IDs using a reference motion database, ensuring accurate trajectory tracking despite occlusion and complex movements.

JP2025150871APending Publication Date: 2025-10-09TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024052015
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Existing video analysis systems face errors in tracking moving objects when their behavior becomes complex, particularly when occlusion occurs, leading to incorrect matching between consecutive image frames.

Method used

A moving object tracking system that includes an imaging device, a moving object detection unit, a position information acquisition unit, a tracking unit, a detection accuracy reduction determination unit, and a moving object identification unit, which identifies the trajectory of a moving object by determining detection accuracy reduction positions and correcting IDs based on a reference motion database.

Benefits of technology

The system accurately tracks moving objects even when detection accuracy decreases, reducing computational load and correcting erroneous IDs, ensuring precise trajectory identification despite occlusion and complex movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025150871000001_ABST
    Figure 2025150871000001_ABST
Patent Text Reader

Abstract

To provide a moving body tracking system capable of suppressing occurrence of an error in tracking a moving body, even when an operation of the moving body is complicated.SOLUTION: A moving body tracking system for calculating a locus of a moving body on the basis of a moving image obtained by imaging the moving body moving in a predetermined area includes: an imaging apparatus for acquiring the moving image by imaging the predetermined area; a moving body detection part for detecting the moving body of the predetermined area; a positional information acquisition part for acquiring positional information of the detected moving body; a moving body tracking part for acquiring the locus by tracking the positional information; a detection accuracy reduction determination part for detecting that the degree of detection of the moving body is reduced to a predetermined degree and acquiring a detection accuracy reduction position where reduction in the degree of detection of the moving body is detected; and a moving body identification part for identifying the locus from the detection accuracy reduction position to a predetermined position as the locus of the moving body, when the moving body reaches a preset predetermined position, after acquiring the detection accuracy reduction position.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a system for tracking a moving object contained in a video or image. [Background technology]

[0002] Patent Document 1 discloses a video analysis device that detects a person region for each image frame of a video or image, and tracks the person using the matching rate of the positions and features of the previous and next person regions. The video analysis device of Patent Document 1 is configured to complement tracking information of the person in previous and next image frames. The video analysis device of Patent Document 1 recognizes a person from the person region using facial recognition or the like, and estimates a person ID by comparing the features calculated from the facial region with pre-collected features related to each person's face. Then, a score is output for changes in person recognition in the person region in other image frames. The higher the score, the more likely the same person is recognized in two person regions. Furthermore, the device detects human behavior in the image frame and assigns a behavior ID to the person region. Then, a score is output for changes in human behavior in the person region in other image frames. The higher the score, the more strongly the human behavior in the two person regions is related. The video analysis device of Patent Document 1 complements the recognition of the person ID, behavior ID, and movement ID assigned to the person region based on the sum of the costs for each image frame output in this way. This allows each person area to be associated with and tracked. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent No. 7065557 Summary of the Invention [Problem to be solved by the invention]

[0004] Patent Document 1 claims that this configuration allows robust person tracking to continue even when people overlap and become hidden in the video captured by the camera, or when person detection fails. However, the video analysis device of Patent Document 1 performs matching based on the strength of association between the moving distance between each image frame and person recognition, behavior recognition, etc. Therefore, if a person performs different actions before and after consecutive image frames, there is a risk of incorrectly matching each person before and after the occlusion. In other words, the video analysis device of Patent Document 1 is prone to errors in matching moving objects between consecutive image frames when the moving object's behavior becomes complex.

[0005] This invention has been made in light of the above technical problems, and aims to provide a moving object tracking system that can suppress errors in tracking a moving object even when the moving object's movements are complex. [Means for solving the problem]

[0006] In order to achieve the above-mentioned object, the present invention provides a moving object tracking system that images an area of ​​a predetermined extent and a moving object moving within the predetermined area, and determines the trajectory of the moving object based on the obtained video, and is characterized by comprising: an imaging device that images the predetermined area and acquires the video; a moving object detection unit that detects the moving object included in the predetermined area; a position information acquisition unit that acquires position information of the detected moving object; a moving object tracking unit that tracks the position information and acquires the trajectory; a detection accuracy reduction determination unit that detects when the detection level of the moving object in the video has decreased to a predetermined level and acquires the detection accuracy reduction position where the decrease in the detection level of the moving object is detected; and a moving object identification unit that, after acquiring the detection accuracy reduction position, identifies the trajectory from the detection accuracy reduction position to the predetermined position as the trajectory of the moving object when the moving object reaches a predetermined position that has been set in advance.

[0007] In addition, the detection accuracy degradation determination unit in this invention may be configured to determine that the degree of detection of the moving body has decreased to the predetermined degree when at least a portion of the moving body overlaps with another object in the predetermined area in the video.

[0008] In addition, the moving body detection unit in this invention may be configured to generate a moving body area that is partitioned into a rectangle so as to include all of the moving body, and the detection accuracy reduction determination unit may be configured to determine that the degree of detection of the moving body has reduced to the predetermined degree when at least a portion of the moving body area overlaps with another object in the video.

[0009] In addition, this invention may further include a database that stores a movement target position that is set in association with the moving body based on a reference movement that defines the movement of the moving body, and the moving body identification unit may be configured to, when another moving body different from the moving body that is associated with the specified position as the movement target position, determine that the other moving body is the moving body and identify the trajectory.

[0010] Furthermore, in this invention, the device may further include a training data generation unit that generates a trained model of the moving body that has learned the features of the moving body when viewed from above through machine learning, the imaging device may be configured to photograph the specified area from above, and the moving body detection unit may be configured to detect the moving body in the specified area based on the trained model. [Effects of the Invention]

[0011] A moving object tracking system according to an embodiment of the present invention is configured to detect a moving object included in a video captured by an imaging device of a predetermined area and determine the trajectory of the moving object moving within the predetermined area. When it is determined that the detection level of the moving object in the video has decreased to a predetermined level, the system acquires the detection accuracy decrease position where the decrease occurred. Then, after acquiring the detection accuracy decrease position, when the moving object reaches a predetermined position, the system identifies the trajectory from the detection accuracy decrease position to the predetermined position as the moving object's trajectory. Therefore, even if an event that decreases detection accuracy occurs, the trajectory of the moving object that has reached the predetermined position is identified as the moving object's trajectory. Therefore, even if the moving object is hidden by overlapping with other objects in consecutive image data or video data, the moving object can be accurately associated and tracked. Furthermore, since the trajectory of the moving object is identified when the detection level of the moving object has decreased to a predetermined level, the load of performing computational processing to identify the moving object's trajectory can be reduced compared to when the trajectory of the moving object is always identified.

[0012] Furthermore, when a moving object in a video that has reached a predetermined position based on a pre-stored reference motion differs from a reference moving object that is scheduled to reach the predetermined position if it moves based on the reference motion, the trajectory of the moving object in the video is identified as the reference moving object. Therefore, even if the same moving object is detected as a different moving object before and after the position where the detection accuracy of the moving object decreases, the trajectory of the moving object can be correctly identified.

[0013] When analyzing an image of a specific area captured from above, a trained model is generated that has previously learned the feature values ​​of a moving object when viewed from above. Since the moving object is detected based on the trained model, it is possible to prevent erroneous detection of a moving object captured in a video. [Brief explanation of the drawings]

[0014] [Figure 1]1 is an explanatory diagram for explaining the entire moving object tracking system according to an embodiment of the present invention; [Figure 2] 1 is a block diagram illustrating a functional configuration of a moving object tracking system according to an embodiment of the present invention. [Figure 3] 10 is an explanatory diagram for explaining a process of detecting a worker by a moving object detection unit. FIG. [Figure 4] 10 is an explanatory diagram for explaining a process of tracking a worker performed by a moving object tracking unit. FIG. [Figure 5] 10 is an explanatory diagram for explaining a process of identifying a moving object executed by a moving object identification unit. FIG. [Figure 6] 4 is a flowchart illustrating an example of control executed by the moving object tracking system according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0015] The present invention will be described below based on the embodiments shown in the drawings. Note that the embodiments described below are merely examples of specific embodiments of the present invention, and are not intended to limit the present invention.

[0016] FIG. 1 shows an overall view of a mobile object tracking system 1 according to an embodiment of the present invention. The mobile object tracking system 1 according to the embodiment of the present invention is applied to a work site 2, which is a predetermined area such as a vehicle factory or an inspection center, where a worker 3, a mobile object moving within the area, moves to each work area 4 in a predetermined order and performs predetermined tasks. The mobile object tracking system 1 is configured to acquire and analyze the trajectory of the worker 3. The mobile object tracking system 1 may be used not only in vehicle factories but also in other factories where distribution and production are performed. For example, in distribution, the mobile object tracking system 1 may be used in warehouses, distribution centers, and delivery centers where mobile objects transport items according to a predetermined work order. Alternatively, in the production field, the mobile object tracking system 1 may be used in factories where mobile objects retrieve and install required materials and parts from predetermined shelves according to a predetermined work order. The mobile object tracking system 1 includes a camera 5, a management server 6, and an information processing device 7.

[0017] The camera 5 is configured to capture an image of a predetermined area in the work site 2, and is an imaging device that captures the predetermined area from above. The camera 5 is fixed to, for example, the ceiling of the work site 2, and is positioned so as to capture an overhead image of the entire work site 2, which is the predetermined area. The camera 5 may be configured similarly to a conventionally known camera 5, for example, an RGB camera 5, and is configured to acquire and output captured images and videos as image data. The image data or video data captured by the camera 5 is output to the information processing device 7 via the management server 6.

[0018] The management server 6 manages the image data or video data acquired by the camera 5. The management server 6 also stores pre-stored recognition information related to the work site 2 and the mobile object. For example, the management server 6 acquires captured video data and outputs it to the information processing device 7. At that time, the management server 6 is configured to output learning data related to the related work site 2 and the mobile object to the information processing device 7 as necessary.

[0019] The information processing device 7 mainly comprises a processor, a communication unit, a storage unit, etc. The information processing device 7 is configured to perform calculations according to a predetermined program using data acquired from the outside and pre-stored data, and to output the results of the calculations as control command signals. For example, the information processing device 7 executes functions that meet predetermined purposes by having the processor load a program stored on a recording medium into a working area of ​​the storage unit and execute the program, and perform various controls through the execution of the program.

[0020] The processor is, for example, a CPU or a DSP. This processor is configured to control the information processing device 7 and perform various information processing operations. The main memory includes, for example, a RAM and a ROM. As described above, the main memory has a work area for the processor to execute programs. The auxiliary memory includes, for example, an EPROM or a hard disk drive. This auxiliary memory may also include a portable recording medium, i.e., a removable medium. The auxiliary memory also reads or writes various programs, data, and tables, allowing them to be freely stored in the recording medium. The auxiliary memory may also store an operating system. The communication unit is a wireless communication circuit connected to an external communication device via wireless communication for data communication. The wireless communication circuit communicates using cellular communication (mobile communication) such as 5G or 4G (LTE). The wireless communication circuit may also communicate using narrowband communication such as DSRC. The information processing device 7 also has a monitor as an output unit.

[0021] Next, a functional configuration of the moving object tracking system 1 according to an embodiment of the present invention will be described. As shown in Fig. 2, the moving object tracking system 1 includes an image acquisition unit 8, a moving object feature acquisition unit 9, a learning data generation unit 10, a moving object detection unit 11, a moving object tracking unit 12, a reference action database 13, a detection accuracy degradation determination unit, a moving object identification unit 15, and an output unit 16.

[0022] The video acquisition unit 8 acquires video data captured by the camera 5. The video data is data acquired by the camera 5 capturing an image of a predetermined area from above, and is a collection of image frames capturing the work site 2 and the worker 3 from above.

[0023] The moving object feature amount acquisition unit 9 acquires feature amounts of the worker 3 contained in the video data acquired by the video acquisition unit 8. The video data mainly includes the work area 4 and the worker 3. The moving object feature amount acquisition unit 9 acquires feature amounts of the work area 4 and the worker 3 by photographing the work site 2, which is a predetermined area, from above. Note that at this time, the moving object feature amount acquisition unit 9 may be configured to accurately acquire feature amounts of the worker 3 by performing predetermined processing and analysis on the image frames. For example, the moving object feature amount acquisition unit 9 removes image noise and distortion from multiple image frames, adjusts brightness and color, etc. to emphasize the worker 3, extracts each moving object (worker 3) from the image data, and detects feature amounts of the worker 3.

[0024] The training data generation unit 10 learns about the worker 3 and the work site 2 through machine learning. The training data generation unit 10 constructs a trained model by inputting a large amount of data about the worker 3, the work site 2, or the work area 4 into a neural network in advance and training the neural network. In this process, in addition to a pre-trained model that has been trained in advance using a dataset related to a known person, additional training is performed on the features of the worker 3 contained in video data captured from different angles. The feature quantities of the worker 3 are recorded in this manner, converted into parameters (weights), and stored. Based on the trained model obtained by such deep learning, the training data generation unit 10 is configured to recognize or detect the worker 3 from video data captured from above. In other words, the training data generation unit 10 improves the accuracy of the trained model of the worker 3 by training the dimensions of the feature quantities of the training dataset for detecting the worker 3 from the side with a predetermined accuracy and the dimensions of the feature quantities of the training data of the worker 3 viewed from above by rotating, translating, shrinking, enlarging, etc.

[0025] In the training data for the worker 3, the overall image of the person, the proportion of each part to the body, color, size, shape, etc. are learned in advance based on a training data set on people. Furthermore, by learning the feature amounts of the worker 3 from different angles, the training data generation unit 10 generates a trained model of the person.

[0026] The learning data generation unit 10 may also learn feature quantities such as color, size, and shape of installed or positioned equipment, tools, jigs, parts, etc., as well as how they appear when viewed from different angles. The learning data generation unit 10 may also learn data on a group of position information for each worker 3 in past video data, and learn the order in which the worker 3 visited each work area 4. For example, when a worker 3 performs work while moving between each work area 4, the learning data generation unit 10 may learn the positions and patterns at which different workers 3 intersect on the video data.

[0027] As shown in FIG. 3, the moving object detection unit 11 detects objects (labeled as people, vehicles, etc.) from the acquired video data based on learning data, such as previously learned object features. The moving object detection unit 11 detects objects from image frames constituting the video data using conventional methods such as YOLO and SSD. The moving object detection unit 11 is then configured to generate bounding boxes, which are rectangular moving object regions that include all of the workers 3, in the areas corresponding to the detected objects based on the coordinates and areas of the detected objects. In this way, the moving object detection unit 11 detects the workers 3 included in the video data based on the trained model of the workers 3 generated by the training data generation unit 10. Because the training data generation unit 10 has learned the features of the work site 2 and the workers 3 viewed from above, the moving object detection unit 11 analyzes the acquired video data to identify the workers 3 and the work area 4. Note that if multiple workers 3 are captured in an image frame, the moving object detection unit 11 detects all of the workers 3.

[0028] As shown in FIG. 4, the moving object tracking unit 12 tracks the image area of ​​the moving object (worker 3) detected by the moving object detection unit 11. The moving object tracking unit 12 assigns an ID to each moving object based on position information of the center coordinates in the image area of ​​each moving object. For example, if there are multiple workers 3 in an image frame, each worker 3 is distinguished and assigned a different ID. The moving object tracking unit 12 acquires position information, which is the center coordinates of the image area of ​​each moving object, for each consecutive image frame that makes up the video data. Note that the arrows shown in FIG. 4 represent the trajectory of each ID.

[0029] The moving object tracking unit 12 also stores data related to the past position information of each worker 3. For example, data such as how each worker 3 moved or where hiding occurred due to overlapping of workers 3 in the video data is stored. The moving object tracking unit 12 predicts the movement path of each worker (ID) 3 based on the past position information. Based on the predicted movement path, the moving object tracking unit 12 considers which work area 4 the worker 3 will move to and compares it with the detected series of position information of the worker 3 to track it. In this way, the moving object tracking unit 12 is configured to track moving objects while preventing the assignment of incorrect IDs. This prevents a decrease in tracking accuracy even when workers 3 overlap or intersect with each other in the image frame. Furthermore, by saving the position information of each ID, the results of tracing the position information of each ID can be output as a trajectory.

[0030] The reference motion database 13 has data on reference motions such as the movement order of the worker 3. The reference motion database 13 records data such as position information, motion information, movement path, and time required in each work area 4 when the reference worker 3 actually performs work at the work site 2. In other words, the reference motion database 13 stores correct motions that are prescribed or set as standards for the worker 3 at the work site 2 and each work area 4.

[0031] The detection accuracy degradation determination unit 14 determines that the detection accuracy of each moving object has decreased in the video data and acquires the detection accuracy degradation position, which is the position where the detection accuracy has decreased. For example, when a worker 3 at least partially overlaps another object in the video data, or when a worker 3 overlaps with or approaches another worker 3 by a predetermined distance, for example, when the two workers 3 approach each other so closely that their bounding boxes become one, the detection accuracy of each worker 3 decreases. The detection accuracy degradation determination unit 14 is configured to determine that an event has occurred in which the detection accuracy has decreased to a predetermined level, and, when the detection accuracy has decreased, acquires the position where the detection accuracy has decreased. Note that the detection accuracy of each worker 3 decreasing to a predetermined level mainly refers to the worker 3 being hidden in the video data, making it impossible to accurately track each worker 3.

[0032] As shown in FIG. 5, the moving object identification unit 15 identifies the workers 3 to whom IDs have been assigned by the moving object tracking unit 12. Specifically, the moving object identification unit 15 first stores position information for each ID assigned to each worker 3. Then, for each work area 4 or at predetermined time intervals, the moving object identification unit 15 compares the position information of the acquired ID with position information (reference position information) based on the reference motion of the worker 3 stored in the reference motion database 13. That is, the moving object identification unit 15 compares whether the reference ID located in the predetermined work area 4 when moving based on the reference motion matches the actual ID located in the predetermined work area 4 captured in the video data. If, as a result of comparing these IDs, the reference ID located in the predetermined work area 4 in the reference motion matches the actual ID located in the predetermined work area 4 in the video data, the moving object tracking unit 12 stores the position information of the ID acquired or stored without modification.

[0033] Conversely, if the reference ID and the actual ID in the video data differ in the predetermined work area 4, the moving object identification unit 15 changes the actual ID based on the reference ID. Specifically, the moving object identification unit 15 determines that the actual ID located in the predetermined work area 4 is different from the reference ID based on the reference motion, and therefore determines that the actual ID has been erroneously detected. Therefore, based on the reference ID corresponding to the reference motion, the moving object identification unit 15 corrects the actual ID in the video data to the reference ID. Furthermore, the moving object identification unit 15 traces the video data back to an image frame to which the correct ID is assigned, and corrects the erroneously assigned actual ID to the reference ID based on the reference. In other words, after acquiring the position of reduced detection accuracy, when the moving object reaches the predetermined work area 4, which is a predetermined position set in advance, the moving object identification unit 15 identifies the trajectory of the worker 3 from the position of reduced detection accuracy to the predetermined work area 4 as the trajectory of the worker 3.

[0034] As an example, a case where erroneously assigned IDs are corrected when two workers 3 move to different work areas 4 will be described with reference to FIG. 5. As shown in FIG. 5, at time t-3, a first worker A (Person A) assigned ID A and a second worker B (Person B) assigned ID B are in different positions at the work site 2. Thereafter, the first worker A and the second worker B move toward the first work area D (Area D) and the second work area E (Area E), respectively. At time t-2 during their movement, the first worker A and the second worker B approach each other or overlap in the image frame, causing the bounding boxes that separated them to merge. As a result, at time t-2, the first worker A and the second worker B are recognized as a single object and are separated by a single bounding box.

[0035] Thereafter, the first worker A and the second worker B continue to move toward their respective work areas D and E. As a result, at time t-1, the first worker A and the second worker B are once again separated and heading toward the first work area D and the second work area E, respectively. However, because they were recognized as one moving body at time t-2, at time t-1, when they are recognized as two moving bodies again, the first worker A is recognized as the second worker B, and the second worker B is recognized as a third worker C (Person C) who has been assigned a new ID C. As a result, at time t, when the first worker A and the second worker B reach the first work area D and the second work area E, respectively, the second worker B is detected to be located at the position of the first worker A, and the third worker C is detected to be located at the position of the second worker B. In other words, the first worker A is erroneously recognized as the second worker B, and the second worker B is erroneously detected as the third worker C.

[0036] The moving object identification unit 15 identifies each worker 3 when they arrive at a predetermined work area 4 or at predetermined time intervals. That is, the moving object identification unit 15 identifies the worker 3 at time t. As described above, the moving object identification unit 15 acquires data regarding the correct movement order of each worker 3 from the reference action database 13 and confirms the work area 4 in which each ID should be located at time t. That is, the moving object identification unit 15 acquires the ID that should be located in the first work area D and the ID that should be located in the second work area E. Then, the correct reference ID based on the movement order is compared with the actual IDs located in each work area 4 from the video data. As a result, it is detected that an incorrect ID has been assigned at time t. The moving object identification unit 15 corrects the actual ID based on the correct reference ID. That is, the moving object identification unit 15 corrects the ID by regarding the second worker B as the first worker A and the third worker C as the second worker B.

[0037] At time t, after correctly recognizing the workers 3 (IDs) located in each work area 4, the video data is reviewed and the IDs are corrected. That is, the actual IDs assigned to each worker 3 at time t-1 are confirmed. At this time, based on the information at time t and the data stored in the reference action database 13, it is determined whether the IDs assigned to each worker 3 at time t-1 are correct. In addition, in this case, it is determined by predicting that the movement path of the moving object will move linearly over an extremely short period of time. At time t-1, a comparison is made based on the movement order and movement path of the first worker A and the second worker B that have been stored or learned in advance, and as a result, incorrect IDs have been assigned, so the actual IDs are corrected. That is, since the first worker A is recognized as the second worker B, and the second worker B is recognized as the third worker C, these IDs are corrected. As a result, the correct IDs can be assigned to each worker 3 at time t-1 as well.

[0038] Similarly, it is determined whether the IDs are correct at time t-2 and time t-3. At time t-2, the first worker A and the second worker B have come closer to each other, and so the bounding boxes are unified, so no particular correction is made. In other words, regardless of whether they are the first worker A or the second worker B, the trajectories of their movements can be traced. Therefore, no correction is made to the IDs at time t-2. Note that, if it is necessary to accurately identify the moving objects, corrections may be made to assign bounding boxes to each worker even at time t-2, so that accurate IDs are assigned. At time t-3, the correct IDs have been assigned to the first worker A and the second worker B, so the above-mentioned corrections are not made.

[0039] In this way, the moving object identification unit 15 identifies the workers 3 at predetermined time intervals or for each predetermined work area 4 based on the correct work order (action order) and movement paths of the workers 3 stored in the reference movement database 13. Then, by correcting the incorrect IDs, the moving object tracking unit 12 can correctly track each ID and improve the accuracy of predicting the movement paths of each ID. Note that the identification of the workers 3 may be configured to be performed before and after an obscuration occurs due to multiple workers 3 overlapping each other in the video data, or a worker 3 overlapping a peripheral device.

[0040] The output unit 16 displays the movement trajectory of each worker 3 based on the tracking information of the worker 3 identified as described above. The output unit 16 outputs video data displaying a trajectory showing the movement route within a predetermined area of ​​the video data for each of the multiple workers 3 captured in the acquired video data. The video data displaying the trajectory of the movement route of such workers 3 is output from a predetermined output unit such as a monitor.

[0041] Next, a control for acquiring video data executed by the moving object tracking system 1 configured as described above, and outputting the video data with the movement trajectories of each worker 3 added thereto will be described. FIG. 6 shows a flowchart for explaining an example of this control. As shown in FIG. 6, in step S1, first, the video data to be analyzed is input. The input of the video data may be automatically acquired video data captured by the camera 5, or may be configured to input predetermined video data selected by operation of a user or the like. The video data may be, for example, video data captured from above of a work site 2 including multiple work areas 4 where workers 3 work.

[0042] Once the video data has been input, the process proceeds to step S2, where detection of objects contained in the video data is performed. In step S2, moving objects are detected label by label from the consecutive image frames that make up the video data, as described above. In step S2, workers 3 are detected from the feature amounts of objects detected from the image frames, based on learning data that has previously learned the feature amounts of workers 3, etc. For example, in step S2, learned data is acquired by weighting the feature amounts when workers 3 are viewed from above, and workers 3 are detected from video data that captures the work site 2 from above. In step S2, if multiple workers 3 are captured in the image frames, all workers 3 are detected and each is defined by a bounding box.

[0043] After detecting a worker 3 in the video data, the process proceeds to step S3. In step S3, an ID is assigned to each detected moving object, and tracking of each is initiated. At this time, multiple pieces of data relating to the position information of multiple workers 3 accumulated by analyzing past video data are referenced. That is, based on the stored group of past position information of workers 3, the location of each ID (worker 3) is predicted, as well as the location at which each worker 3 will be hidden due to overlapping in the video data. Then, by comparing the predicted movement path of each ID with the actually acquired movement path of each ID, tracking is performed to prevent erroneous assignment of each ID. After tracking of each ID is initiated in this manner, the process proceeds to step S4.

[0044] In step S4, the position information of each tracked worker 3 and each ID assigned to each worker 3 are saved. In step S4, the ID corresponding to each worker 3 acquired in each image frame and the position information of that ID are saved. Note that in step S4, the center coordinates of each worker 3 when partitioned by a bounding box are stored as the position information of each worker 3 in association with the ID.

[0045] After the process in step S4 is completed, the process proceeds to step S5, where it is determined whether the worker 3 has reached a predetermined work area 4 that has been set in advance. In step S5, data on the reference motion of the worker 3 is obtained from the reference motion database 13. In the reference motion of each worker 3, it is determined whether each worker 3 in the input video data has reached a predetermined position that has been set in advance. For example, in step S5, it is determined whether each worker 3 in the video data has reached any of the work areas 4 in the work site 2. If the determination in step S5 is NO because each worker 3 has not reached any of the work areas 4, the process returns to step S4, and the process of correlating and storing the position information of each worker 3 and the corresponding ID continues. Note that step S5 may be configured to determine whether at least one worker 3 has reached any of the work areas 4.

[0046] Conversely, if the determination in step S5 is YES because each worker 3 has reached one of the work areas 4, the process proceeds to step S6. In step S6, the saved position information of each worker 3 is compared with position information (reference position information) based on the reference movement of each worker 3 stored in advance in the reference movement database 13. That is, in step S6, the actual ID of the worker 3 who has reached one of the work areas 4 is compared with the reference ID based on the reference movement stored in the reference movement database 13.

[0047] After processing in step S6, the process proceeds to step S7, where it is determined whether the actual ID of each worker 3 who has reached any of the work areas 4 matches the reference ID based on the reference action. If the comparison in step S6 shows that the actual ID located in any of the work areas 4 matches the reference ID that would be located if the worker 3 performed the reference action, which is the correct action for the work, then the determination in step S7 is YES.

[0048] If the determination in step S7 is YES, the process proceeds to step S8, where it is determined whether the input video data has ended. If the determination in step S7 is YES, it is determined that the worker 3 who has arrived at any of the work areas 4 has been correctly detected, and that tracking of the worker 3 has been performed correctly. Therefore, the process proceeds to the next step, step S8, without performing any special processing.

[0049] When the process proceeds to step S8, it is determined whether the video data has ended. If the video data has not ended and the determination is NO in step S8, the process returns to step S3. In other words, tracking of each worker 3 in the video data continues.

[0050] Conversely, if the video data has ended and the determination in step S8 is YES, the process proceeds to step S9. In step S9, the results of tracking each worker 3 are output. That is, in step S9, video data is output that displays a group of position information for each worker 3 and a trajectory showing the route taken by each worker 3, as a result of analyzing the video data.

[0051] On the other hand, if the ID of a worker 3 located in any of the work areas 4 differs from the ID of the worker 3 located when performing the reference motion, step S7 returns NO, and the process proceeds to step S10. In step S10, the actual ID of the worker 3 located in any of the work areas 4 is corrected to a reference ID based on the reference motion. If the process proceeds to step S10, it has been determined that the actual ID has been erroneously detected because the stored or detected actual ID differs from the reference ID based on the reference motion. In other words, since each worker 3 moves through each work area 4 in accordance with the reference motion, if the reference ID and the actual ID differ, the actual ID has been erroneously detected. Therefore, in step S10, the actual ID is corrected to the reference ID, which is the correct ID, to accurately identify each worker 3.

[0052] Specifically, in step S10, the actual ID of a worker 3 located in one of the work areas 4 is corrected to a reference ID, and then the ID assigned in the previous image frame is corrected. That is, in step S10, the IDs are compared by going back through consecutive image frames to correct the ID back to the point where the ID was erroneously assigned. This allows the movement path of each worker 3 before arriving at one of the work areas 4 to be accurately tracked. For example, two workers 3 may be detected as one worker 3 due to their proximity to each other, and then later, when they are detected as two workers 3 again, an ID may be erroneously assigned. In such a case, the system goes back to the point where the two workers 3 were recognized as one worker 3, and corrects the erroneously assigned ID to the reference ID. Furthermore, if an ID predicted based on the past location information of each ID is saved in an incorrect state, the system may also be configured to correct the past ID based on that prediction.

[0053] After the process of step S10 is completed by correcting the actual ID based on the video data with the reference ID based on the reference action, the process proceeds to step S8. In step S8, it is determined whether the video data has ended, as described above. If the determination in step S8 is NO because the video data has not ended, the process returns to step S3. Conversely, if the determination in step S8 is YES because the video data has ended, the process proceeds to step S9, and the results of analyzing the video data are output. That is, as a result of analyzing the video data, video data is output that displays a group of position information for each worker 3 and a trajectory indicating the route taken by each worker 3.

[0054] As described above, the mobile object tracking system 1 according to an embodiment of the present invention detects workers 3 based on the feature amounts of each worker 3 captured in the video data and the feature amounts of the worker 3 based on the trained model. The system is configured to assign an ID to each worker 3 and track the movement path of each worker 3 by acquiring and tracking the position information of the ID in multiple image frames constituting the video data. At this time, the system determines whether the actual ID assigned to each worker 3 is correct based on a reference motion defined for the worker 3's behavior. For example, if the ID of a worker 3 actually located in a specified work area 4 differs from the ID of the worker 3 located when performing the reference motion, the system is configured to correct the actual ID of the worker 3 located in the actual area to the reference ID based on the reference motion. Therefore, even if an ID is incorrectly assigned to each worker 3, the correct ID based on the reference motion can be assigned. That is, each worker 3 in the video data can be accurately associated. Therefore, even if the workers 3 are hidden in the video data due to overlapping with each other, each worker 3 can be accurately associated and tracked, allowing for more accurate analysis or interpretation of the work site 2 based on the movement paths of each worker 3 obtained from the video data.

[0055] Although the embodiments of the present invention have been described above, the present invention is not limited to the above examples and may be modified as appropriate within the scope of achieving the object of the present invention. For example, learning data for detecting each worker 3, task sequences, and the like may be managed by the information processing device 7 instead of by the management server 6. In other words, the information processing device 7 may be configured to have all of the above-described functional configurations, and video data captured by the camera 5 may be directly output to the information processing device 7. Furthermore, a step of determining that the level of detection of each worker 3 has decreased to a predetermined level may be included between steps S4 and S5 of the above-described flowchart. [Explanation of symbols]

[0056] 1. Mobile tracking system 2. Work site (prescribed area) 3. Worker (mobile) 4 Work area (designated location) 5. Camera (imaging device) 6 Management Server 7. Information processing equipment 8. Video acquisition unit 9. Moving object feature acquisition unit 10. Training data generation unit 11 Moving object detection unit (location information acquisition unit) 12 Mobile Object Tracking Unit 13 Reference Motion Database 14. Detection accuracy degradation determination unit 15 Moving object identification unit 16 Output section

Claims

1. A moving object tracking system that captures an image of a predetermined area and a moving object moving within the predetermined area, and determines a trajectory of the moving object based on the obtained video, an imaging device that captures an image of the predetermined area to acquire the video; a moving body detection unit that detects the moving body included in the predetermined area; a location information acquisition unit that acquires location information of the detected moving object; a moving object tracking unit that tracks the position information and acquires a trajectory; a detection accuracy degradation determination unit that detects that the degree of detection of the moving object in the video has decreased to a predetermined degree and acquires a detection accuracy degradation position where the decrease in the degree of detection of the moving object has been detected; and a moving body identifying unit that, when the moving body reaches a predetermined position set in advance after acquiring the position where detection accuracy has decreased, identifies the trajectory from the position where detection accuracy has decreased to the predetermined position as the trajectory of the moving body. A mobile object tracking system.

2. 2. The moving object tracking system according to claim 1, The detection accuracy degradation determination unit determines that the degree of detection of the moving object has decreased to the predetermined degree when at least a part of the moving object overlaps with another object in the predetermined area in the video. A mobile object tracking system.

3. 2. The moving object tracking system according to claim 1, the moving object detection unit generates a moving object region partitioned into a rectangle so as to include all of the moving objects; The detection accuracy degradation determination unit determines that the degree of detection of the moving object has decreased to the predetermined degree when at least a part of the moving object area overlaps with another object in the video. A mobile object tracking system.

4. 4. A moving object tracking system according to claim 1, a database storing a movement target position set in association with the moving body based on a reference movement that defines the movement of the moving body; The moving body identification unit is configured to, when a moving body other than the moving body associated with the predetermined position as the movement target position is detected at the predetermined position, determine that the other moving body is the moving body and identify the trajectory. A mobile object tracking system.

5. 4. A moving object tracking system according to claim 1, a learning data generation unit that generates a trained model of the moving object by learning a feature amount of the moving object when viewed from above by machine learning; the imaging device is configured to capture an image of the predetermined area from above, The moving object detection unit is configured to detect the moving object in the predetermined area based on the trained model. A mobile object tracking system.

Citation Information

Patent Citations

  • Video analysis device, program, and method for tracking a person

    JP7065557B2