Monitoring system and monitoring method
The monitoring system efficiently manages work time across multiple areas by using cameras and analysis devices to detect worker locations, addressing the challenge of manual division and improving productivity through automated task time measurement.
Patent Information
- Application Number
- JP2021087234
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-05-24
- Publication Date
- 2025-12-03
- Estimated Expiration
- 2041-05-24
AI Technical Summary
Existing work analysis systems face challenges in efficiently managing work time across multiple work areas, particularly when workers move between these areas, as they require manual division of work information and are not effective for workers moving between different locations.
A monitoring system using cameras to capture images of multiple work areas, with an analysis device that detects workers and determines their location within these areas, integrating frames to measure work time accurately using evaluation indices like Intersection over Union (IoU) to define work areas.
Enables easy management of work time for each work area, reducing manual effort and improving productivity by automatically tracking worker movements and task times across various locations.
Smart Images

Figure 0007779472000002 
Figure 0007779472000003 
Figure 0007779472000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to a monitoring system and a monitoring method using a camera for monitoring work in a plurality of work areas. [Background technology]
[0002] At work sites, managers and others use video cameras to record the work of workers in order to consider ways to improve work efficiency. Then, work analysis techniques are implemented to analyze the work content of workers based on the work images and identify wasteful work. In work analysis, an analyst with the skills to perform work analysis must view the video of the work being evaluated and divide the filmed work content into parts.
[0003] Patent Document 1 discloses a technology for dividing the contents of a work state into steps and analyzing the work contents using a video playback device that plays back a video tape that records the work state captured by a video camera and a computer connected to the video playback device. Patent Document 2 also discloses a technology in which a worker or a manager links work images with the work contents.
[0004] Non-Patent Document 1 examines several methods, including a local feature extraction method in Bag-of-Features, which is one of the image feature representation methods, and a multi-class classifier construction method using Support Vector Machine, which is one of the pattern recognition models, and conducts image classification experiments with the aim of improving the accuracy of the process of classifying each frame of a video file that captures work scenes according to the work content. It also discloses that a representative work per unit time is selected based on the classification results, and the representative work is arranged in chronological order, thereby generating a work chart and calculating the work time. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 6-231137 [Patent Document 2] Patent No. 5416322 [Non-patent literature]
[0006] [Non-Patent Document 1] Hiroki Watanabe and two others, "Development of a Work Time Estimation System Using Machine Learning," Gifu Prefectural Information Technology Research Institute Research Report No. 18, pp. 15-21, 2016 Summary of the Invention [Problem to be solved by the invention]
[0007] According to the computer-aided work analysis device using video in Patent Document 1, a bird's-eye view of the work area is displayed on a display device, and the analyst analyzing the work content looks at the image on the display device, selects the relevant work detail name from the detail button area at the change of work, and clicks with the mouse to divide the work information. However, if a series of work information is long, the time required for the division work also takes a long time, which creates the problem of making the division work difficult.
[0008] In Non-Patent Document 1, the work time is calculated based on the work classification results, but the work images are classified for the inspection process, in which the worker sits in a chair and works at a desk, and the work is not classified for the worker moving between multiple work areas.
[0009] The present invention is an invention to solve the above-mentioned problems, and aims to provide a monitoring system and a monitoring method that allow workers to easily manage work time for each of multiple work areas that involve movement. [Means for solving the problem]
[0010] In order to achieve the above object, the monitoring system of the present invention comprises a camera that captures images of a plurality of work areas, and an analysis device that analyzes the video data obtained by the camera, the analysis device including a worker detection unit that detects a worker in a frame of the video data, and an analysis device that determines which work area the worker detected by the worker detection unit is in. The definition domain of and a work time measurement unit that determines whether a worker exists in each of the work areas and integrates the determined frames to measure the work time of each of the work areas, wherein the worker detection unit detects a worker in a frame of the video data as a rectangle, and the work time measurement unit calculates an evaluation index that indicates the degree of overlap between a detection rectangle that is a rectangle corresponding to the worker detected by the worker detection unit and a rectangle in a definition area of each work area, If the defined regions overlap, The present invention is characterized in that it is determined that the worker is present in an area where the evaluation index is equal to or greater than a predetermined threshold and where the evaluation index is the largest. Other aspects of the present invention will be described in the embodiments below. [Effects of the Invention]
[0011] According to the present invention, it is possible to easily manage the working time for each of a plurality of work areas in which a worker moves. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a diagram showing the configuration of a monitoring system according to a first embodiment. [Figure 2] FIG. 2 is a diagram showing a plurality of working areas according to the first embodiment. [Figure 3] FIG. 3 is a diagram showing an overview of an operation time measurement process according to the first embodiment. [Figure 4] FIG. 2 is a diagram illustrating a worker detection model according to the first embodiment. [Figure 5] FIG. 4 is a diagram showing the relationship between a person area and a work area according to the first embodiment. [Figure 6] 1A and 1B are diagrams showing the relationship between GT and PD according to the first embodiment, where (a) is the area (GT∩PD) and (b) is the area (GT∪PD). [Figure 7]FIG. 2 is a diagram illustrating an example of IoU according to the first embodiment. [Figure 8] 4 is a flowchart showing an operation time measurement process according to the first embodiment. [Figure 9] FIG. 4 is a diagram showing an example of a task time measurement result according to the first embodiment. [Figure 10] FIG. 10 is a diagram showing the results of work time for each work process according to the first embodiment. [Figure 11] FIG. 10 is a diagram showing the processing of a task classification unit according to the second embodiment. [Figure 12] 10 is a flowchart showing an operation time measurement process according to the second embodiment. [Figure 13] FIG. 11 is a diagram illustrating processing using a depth map according to the third embodiment. [Figure 14] 11 is a flowchart showing an operation time measurement process according to the third embodiment. [Figure 15] FIG. 13 is a diagram showing an object detection model using a 3D model according to the fourth embodiment. [Figure 16] FIG. 13 is a diagram showing an example of a multi-viewpoint image according to the fourth embodiment. [Figure 17] FIG. 13 is a diagram showing a combination of data sets according to the fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. First Embodiment 1 is a diagram showing the configuration of a monitoring system according to a first embodiment. The monitoring system MS includes a camera 60 that captures images of multiple work areas and an analysis device 100 that analyzes the video data obtained by the camera 60. The camera 60 and analysis device 100 are connected via a network NW such as a local area network (LAN). The monitoring system MS is a monitoring system that uses video captured from a bird's-eye view and targets work that involves movement between multiple work areas at a manufacturing site or work site.
[0014] The photographing device 60 includes a photographing device 61 that photographs the work area A1, a photographing device 62 that photographs the work area A2, etc. The photographing device 60 can photograph the work area from a bird's-eye view. The photographing device 60 is, for example, a web camera, and transmits the photographed video to the analysis device 100.
[0015] The analysis device 100 has a processing unit 10, a memory unit 20, an input unit 30, an output unit 40, and a communication unit 50. The processing unit 10 has a video data storage unit 11 that stores video data sent from a photographing device 60 in the memory unit 20, a worker detection unit 12 that detects workers in frames of the video data, a work time measurement unit 13, and the like.
[0016] The work time measurement unit 13 determines in which work area the worker detected by the worker detection unit 12 is present, and measures the work time for each work area by integrating the determined frames.
[0017] The storage unit 20 stores a video database 21, work time measurement results 22 (see FIG. 9), work time results 23 for each work process (see FIG. 10), and the like.
[0018] The video data stored in the video database 21 will now be described. When capturing video using the imaging device 60 (e.g., a web camera), if the frame rate is set to 30 fps (frames per second), 30 still images per second are stored in the video database 31, and video is managed using a time management method such as a time axis or time code. Here, we will explain using time codes. If the time code is "00:07:50:10," this means that the display position is 0 hours, 7 minutes, 50 seconds, and the 10th frame.
[0019] When the shooting device 60 starts shooting, still images (for example, images in JPEG (Joint Photographic Experts Group) format) with time codes "00:00:00:00", "00:00:00:01", ..., "00:00:00:29", "00:00:01:00", ..., "00:07:50:10" are stored in the video database 21 of the memory unit 20.
[0020] In this embodiment, a method for measuring work time for each work area is described, in which if a worker is present in the work area in each frame, the time is counted as work time in that work area, and if the worker is not present in the work area, the time is counted as non-work time. This embodiment makes it possible to easily manage work time for each of multiple work areas.
[0021] FIG. 2 is a diagram showing multiple work areas according to the first embodiment. The work area has multiple work areas (such as work area R1). In each work area, one worker is in charge of the work process in that work area. For example, there are multiple work areas R1 to R5 within work area A1 (see FIG. 1), and one worker is in charge of the work in those multiple work areas R1 to R5.
[0022] Workers engaged in production processes are inevitably required to be multi-skilled to handle a variety of tasks. A multi-skilled worker is a worker who has the skills to handle multiple tasks and processes, and each task involves different products, work locations, and working hours.
[0023] Furthermore, there is a demand for understanding the current productivity at production sites in order to improve work efficiency. Until now, manual measurement was required to understand productivity. However, manual measurement places a heavy burden on workers and is difficult to measure over long periods of time. Therefore, automatically understanding the current situation at production sites from images captured by an easily obtainable imaging device 60 is thought to reduce the burden on workers while also making previously unseen problems visible. Furthermore, visualization can lead to awareness of increased work efficiency and management improvements.
[0024] In this embodiment, the area where the task time is to be measured is defined in advance as a bounding box, as shown in Figure 2. This area is called a defined task area (e.g., task area R1). Note that a bounding box is a rectangular frame that surrounds an image or the like.
[0025] FIG. 3 is a diagram illustrating an overview of the task time measurement process according to the first embodiment. The task time measurement unit 13 inputs a video at a predetermined frame rate (e.g., 1 fps) (step S11), and the worker detection unit 12 performs worker (person) detection on the input video (step S12). The worker detection unit 12 uses Faster R-CNN (see FIG. 4) to set a person detection area surrounding the worker. The task time measurement unit 13 calculates the IoU (Intersection over Union) between the person detection area and the predefined task area shown in FIG. 2 (step S13). IoU represents the overlap rate between the two areas. Furthermore, based on these results, the task time measurement unit 13 determines the task area where the IoU between the person detection area and each predefined task area exceeds a predetermined threshold and maximizes the IoU (step S14). The task time measurement unit 13 determines whether all frames of the target video have been processed (step S15). If not, the process returns to step S12; if not, the process proceeds to step S16. The task time measurement unit 13 calculates the task time for each task area (step S17), outputs the calculation result (step S17), and ends the series of processes.
[0026] 4 is a diagram showing a worker detection model according to the first embodiment. In this embodiment, Faster R-CNN is adopted in step S13. Faster R-CNN is a general object detection model consisting of two stages: an RPN (Region Proposal Net) that estimates object candidate regions from a feature map, and a network that estimates the class labels and rectangular positions of objects present in the object candidate regions.
[0027] Backborn CNN 71 generates a feature map for the input image (input image 70). Next, RPN 72 uses the generated feature map as input to estimate object candidate regions. The estimated object candidate regions are converted into fixed-length feature vectors by ROI Pooling 73. Finally, FC 74, 75 output the class probability (Class 76) of the type of object each object candidate region is and the BB of the object position (BB 77). Note that FC is an abbreviation for fully connected layer, and BB is an abbreviation for bounding box.
[0028] In detail, the fixed-length feature vectors converted by ROI Pooling 73 are input to two-layer FCs 74 and 75 to obtain intermediate feature vectors. The intermediate feature vectors are then input to FC 75A for object class probability output and FC 75B for object position deviation, which output the object class probability and the accurate rectangular coordinates of the object. Here, a class probability of 0.9 or higher for the human class is considered to have been correctly detected as a human, and a rectangle is drawn.
[0029] FIG. 5 is a diagram showing the relationship between the person area and the work area according to the first embodiment. In FIG. 5, the bounding box drawn with a thick solid line indicates the detection result of the worker detection unit 12, and is set so as to surround the entire body of the worker. The bounding boxes of work areas R1 to R5 are the worker's work locations. The bounding boxes of the work locations are defined in advance, as shown in FIG. 2. It is necessary to determine at which work area within work area R1 to R5 the worker is working. This worker assignment is processed using the IoU described above (see step S13 in FIG. 3).
[0030] IoU is an evaluation metric for object detection. The definition of IoU (Intersection over Union) is shown in Equation (1).
number
[0031] In equation (1), GT represents the ground truth bounding box, and PD represents the predicted detection bounding box.
[0032] FIG. 6 is a diagram showing the relationship between GT and PD according to the first embodiment, where (a) shows the area (GT∩PD) and (b) shows the area (GT∪PD). From FIG. 6 and formula (1), the value determined by dividing the intersection of the two bounding box areas by the union of the areas is defined as IoU. From the above, IoU can be said to be an index that represents the overlap rate of the two areas, GT and PD, in the range of 0 to 1.
[0033] Fig. 7 is a diagram showing examples of IoU according to the first embodiment. Fig. 7(a) shows the case where IoU is 1, Fig. 7(b) shows the case where IoU is approximately 0.68, Fig. 7(c) shows the case where IoU is approximately 0.17, and Fig. 7(d) shows the case where IoU is 0.
[0034] In Figure 7, GT and PD are each a 10x10 square. In Figure 7(a), GT and PD are perfectly aligned, so area(GT∩PD) and area(GT∪PD) are both 10x10 = 100, and therefore IoU = area(GT∩PD) / area(GT∪PD) = 1. In Figure 7(b), there is a shift of 1 in each of the x- and y-axis directions. In this case, area(GT∩PD) = 9x9 = 81, and area(GT∪PD) = 10x10 + 10x10 - 9x9 = 119, so IoU = area(GT∩PD) / area(GT∪PD) = 81 / 119 ≒ 0.68. In Figure 7(c), there is a shift of 5 in each of the x- and y-axis directions. Calculating in the same way as (b), IoU = area(GT∩PD) / area(GT∪PD) ≒ 0.17. In Figure 7(d), GT and PD do not completely match, so area(GT∩PD) = 0, and therefore IoU = area(GT∩PD) / area(GT∪PD) = 0. From the above, when the two regions completely overlap, IoU is 1, when the two regions do not completely overlap, IoU is 0, and the greater the overlap rate, the larger the IoU.
[0035] The proposed method of this embodiment was evaluated using precision and recall. First, we define True Positive (TP), False Negative (FN), False Positive (FP), and True Negative (TN). A predicted value of 1 indicates that the output of the proposed method is determined to be in progress in a certain working area of a certain frame, and a correct answer value of 1 indicates that the correct answer in that frame is in progress. TP represents the number of frames with a predicted value of 1 and a correct answer value of 1, FN represents the number of frames with a predicted value of 0 and a correct answer value of 1, FP represents the number of frames with a predicted value of 1 and a correct answer value of 0, and TN represents the number of frames with a predicted value of 0 and a correct answer value of 0.
[0036] The precision is shown in formula (2), and the recall is shown in formula (3). Precision rate=TP / (TP+FP) Formula (2) Recall rate=TP / (TP+FN)...Equation (3)
[0037] Precision indicates the percentage of predicted values of 1 that are actually correct, and represents the accuracy of prediction in the working area. Recall indicates the percentage of correct values of 1 that are actually predicted, and represents the accuracy of detection in the working area.
[0038] As a result, the precision rate of the work area R1 among the work areas shown in Fig. 5 was 1, and the recall rate was 0.93, confirming the effectiveness of the method proposed in this embodiment.
[0039] Fig. 8 is a flowchart showing the task time measurement process according to the first embodiment. Fig. 8 explains steps S12 to S14 in Fig. 3 in detail. Step S12 in Fig. 3 corresponds to steps S21 and S27 in Fig. 8. Step S13 in Fig. 3 corresponds to step S22 in Fig. 8. Step S14 in Fig. 3 corresponds to steps S23 to S25 and S28 in Fig. 8.
[0040] The worker detection unit 12 performs worker detection (person detection) on the input video (step S21). If no person is detected in the frame (step S21, No), the work time measurement unit 13 sets a flag indicating that the person is out of frame (step S29).
[0041] On the other hand, if a person is detected in the frame (step S21, Yes), the task time measurement unit 13 calculates the IoU (Intersection over Union) between the person detection area and each predefined task area shown in Fig. 2 (step S22).The task time measurement unit 13 then determines whether there is an IoU that is equal to or greater than a threshold (step S23), and if all IoUs are less than the threshold (step S23, No), it flags the frame as being outside the task time (step S28).
[0042] On the other hand, if there is any frame with an IoU greater than or equal to the threshold (step S23, Yes), the frame is determined to be within the work time (step S24). Assuming that the monitoring system is operated under conditions where one worker is present in the video, if the IoU of multiple defined work areas is greater than or equal to the threshold, the area with the largest IoU is determined to be the work location (step S25). This makes it possible to determine for each frame of the video whether it is within the work time, outside the work time, or out of the frame.
[0043] FIG. 9 is a diagram showing an example of a work time measurement result according to the first embodiment. In the work time measurement result 22 shown in FIG. 9, the work area of the frame is determined every 1 fps. Specifically, at "20XX / 05 / 10 9:00:00:01" (first frame, 9:00:00, May 10, 20XX), it is determined to be out of frame. At "20XX / 05 / 10 9:00:03:01" (first frame, 9:00:03, May 10, 20XX), it is determined to be outside of work hours. In this frame, a person is determined, but all IoUs are below the threshold. At "20XX / 05 / 10 9:01:10:01" (first frame, 9:01:10, May 10, 20XX), it is determined that a worker is in the work area R1.
[0044] Compiling the above results, we can see that the work time in work area R1 was 185 seconds, in work area R2 it was 110 seconds, in work area R3 it was 250 seconds, in work area R4 it was 170 seconds, and in work area R5 it was 70 seconds. We can also see that it took 155 seconds to move between work areas.
[0045] FIG. 10 is a diagram showing an example of a work time result for each work process according to the first embodiment. The work time result 23 for each work process shown in FIG. 10 includes the work process ID, work area, work start time, work end time, work duration, camera ID, storage location of video captured by the image capture device 60, and worker ID. As shown in FIG. 10, for work process A0001, there are five work areas. For example, work in work area R1 started at 9:01:10 and ended at 9:04:15. That is, the work time for work area R1 is 3 minutes and 5 seconds (185 seconds). Furthermore, the folder destination for storing images captured by camera ID C0001 is e:¥C0001¥20xx0510¥090110, and the worker ID is MS001. The worker ID can be automatically registered by holding a worker ID or similar document over the image capture device 60 before starting work.
[0046] In the case of operation A0002, there are three operation areas. For example, operation area R1 started at 9:20:15 and finished at 9:25:25. In other words, the operation time for operation area R1 was 5 minutes 10 seconds (310 seconds). Furthermore, the folder destination for storing images captured by camera ID C0002 is e:\C0002\20xx0510\092015, and the worker ID is MS002.
[0047] Referring to Figure 10, if we focus on the fact that there is a gap of more than one minute between the start and end times of the work areas, in the case of work process A0001, it took one minute and 20 seconds between work area R3 and work area R4. Also, in the case of work process A0002, it took one minute and 50 seconds between work area R1 and work area R2. From this, it can be inferred that there is room for improvement in the distance between work areas and the layout.
[0048] According to the first embodiment, the working time for each of a plurality of working areas can be easily managed.
[0049] Second Embodiment In the second embodiment, compared to the first embodiment, a task classification process for classifying whether the task is in progress or not in progress is added to the processing unit 10. This allows the task time to be determined with even greater accuracy.
[0050] Fig. 11 is a diagram showing the task classification process according to the second embodiment. The task classification process performs two-class classification: task and non-task. This task classification process can handle situations where a worker overlaps with the task area but is not currently working, such as when a moving worker crosses the task area.
[0051] The task classification unit process performs task classification using a convolutional neural network for image classification that receives a single frame as input. AlexNet, VGGNet, and ResNet are implemented as typical architectures of neural networks used for image classification, and the task classification unit 14 to be used was selected by training using a task classification dataset and comparing the accuracy. As a result, VGGNet was adopted in this embodiment.
[0052] AlexNet is based on Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton, “Imagenet classification with deep convolutional neural networks,” NIPS, 2012. VGGNet is based on Karen Simonyan and Andrew Zisserman, “Very deep convolutional networks for large-scale image recognition,” ICLR, 2015. ResNet is based on Kaiming He, Georgia Gkioxari, Piotr Dollar, and Ross Girshick, “Mask R-CNN,” ICCV, 2017.
[0053] In Figure 11, two layers (FC1, FC2) of FC (fully connected layer) are used. In addition, a softmax function is used to convert and output multiple output values so that the sum of the values is 1.0 (=100%). As a result, the range of each output value is 0.0 to 1.0.
[0054] Fig. 12 is a flowchart showing the task time measurement process according to the second embodiment. In Fig. 12, the same processes as in Fig. 8 are denoted by the same reference numerals.
[0055] The worker detection unit 12 performs worker detection (person detection) on the input video (step S21). If no person is detected in the frame (step S21, No), the work time measurement unit 13 sets a flag indicating that the person is out of frame (step S29).
[0056] On the other hand, if a person is detected in the frame (step S21, Yes), the task time measurement unit 13 calculates the IoU (Intersection over Union) between the person detection area and each predefined task area shown in Fig. 2 (step S22).The task time measurement unit 13 then determines whether there is an IoU that is equal to or greater than a threshold (step S23), and if all IoUs are less than the threshold (step S23, No), it flags the frame as being outside the task time (step S28).
[0057] On the other hand, if there is any frame with an IoU greater than or equal to the threshold (step S23, Yes), a determination is made as to whether or not work is in progress using task classification processing (step S26). If work is not in progress (step S26, No), the process proceeds to step S28. If work is in progress (step S26, Yes), the frame is deemed to be within work time (step S24). Assuming that the monitoring system is operated under conditions in which one worker is present in the video, if the IoU of multiple defined work areas is greater than or equal to the threshold, the area with the largest IoU is determined to be the work location (step S25). This makes it possible to determine for each frame of the video whether the worker is within work time, outside work time, or out of frame.
[0058] According to the second embodiment, by applying the task classification process, task times for each of a plurality of task areas can be managed with even greater accuracy.
[0059] <Third embodiment> Compared to the first embodiment, the third embodiment adds a process for determining the working area using a depth map, taking into account the depth from the image capturing device 60. This allows the working time to be determined with even greater accuracy.
[0060] FIG. 13 is a diagram showing processing using a depth map according to the third embodiment. A depth map and a binary image of the human region are estimated from the input frame, and a depth image of only the human pixels is output by multiplying each element of both images. After cropping the image with the detection rectangle, a depth histogram is generated. The generated depth histogram is compared with a pre-registered depth histogram template of the work image for each region, and a flag indicating whether or not the object is likely to remain is output. This allows for the overlap of the detection rectangle and the work region to be handled by taking depth information into account, as well as the overlap of the work region itself.
[0061] More specifically, as shown in Figure 13, worker pixels are estimated using a semantic segmentation model and a depth map is estimated using a depth estimation model. The depth map containing only worker pixels is then extracted by multiplying the two images. The depth map containing only worker pixels is then cropped using the detection rectangle to calculate a depth histogram. The Euclidean distance between the calculated depth histogram and the template depth histogram for all work areas where the IoU exceeds a threshold is calculated, and a flag is output indicating that work is in progress in the area with the smallest distance. This makes it possible to measure work time taking into account the depth direction.
[0062] We use MiDaS (v2.1), proposed by Ranftl et al., as a model capable of depth estimation without annotating and fine-tuning depth maps to suit the environment. In their research, Ranftl et al. proposed a loss function and optimization method that allows multiple datasets used in depth estimation training to be treated as a single large dataset, making it possible to train the system while taking into account environmental factors such as indoors and outdoors, differences in subjects such as stationary and moving objects, and differences in output numerical data such as relative depth and absolute depth.
[0063] To estimate worker pixels, we use Mask R-CNN, the de facto standard semantic segmentation model. Mask R-CNN is an improved version of the object detection model Faster R-CNN for instance segmentation. Faster R-CNN outputs object position and class probability in the final Fully Connected Layer (FC), but Mask R-CNN estimates target object pixels within the candidate region by adding a new branch for instance segmentation.
[0064] MiDaS references Rene Ranftl, Katrin Lasinger, David Hafner, Konrad Schindler and Vladlen Koltun, “Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-shot Cross-dataset Transfer,” TPAMI, 2020. Mask R-CNN is based on Kaiming He, Georgia Gkioxari, Piotr Dollar and Ross Girshick, “Mask R-CNN,” ICCV, 2017.
[0065] Fig. 14 is a flowchart showing the work time measurement process according to the third embodiment. In Fig. 12, the same processes as in Fig. 8 are denoted by the same reference numerals. The worker detection unit 12 performs worker detection (person detection) on the input video (step S21). If no person is detected in the frame (step S21, No), the work time measurement unit 13 sets a flag indicating that the person is out of frame (step S29).
[0066] On the other hand, if a person is detected in the frame (step S21, Yes), the task time measurement unit 13 calculates the IoU (Intersection over Union) between the person detection area and each predefined task area shown in Fig. 2 (step S22).The task time measurement unit 13 then determines whether there is an IoU that is equal to or greater than a threshold (step S23), and if all IoUs are less than the threshold (step S23, No), it flags the frame as being outside the task time (step S28).
[0067] On the other hand, if the IoU is greater than or equal to the threshold (step S23, Yes), the frame is determined to be within the work time (step S24). Among multiple predefined work areas determined by the depth map, the generated depth histogram is compared with pre-registered depth histogram templates of the work images in each area to determine the closest one as the work location (step S27). This makes it possible to determine whether each frame of the video is within the work time, outside the work time, or out of frame.
[0068] According to the third embodiment, by taking depth information into consideration, it is possible to manage the work time for each of a plurality of work areas with even greater accuracy.
[0069] <Fourth embodiment> As explained above, the worker detection unit 12 in the first embodiment employs the object detection method Faster R-CNN. Before object detection, the system is trained using a dataset annotated with actual objects. In this embodiment, in order to study the improvement of object detection accuracy, a combination of a dataset created using a 3D model in addition to actual objects is studied. In this example, a cart at a work site is used as the object.
[0070] In machine learning, annotation refers to adding metadata to data to give it meaning. By annotating large amounts of data and adding correct data (teaching data), it is possible to determine what is correct about the machine learning model.
[0071] FIG. 15 is a diagram showing an object detection model using a 3D model according to the fourth embodiment. The process for the object detection model is as follows: First, a training dataset is created by annotating the carts present in the captured video. Annotation is performed by adding a bounding box (BB). Next, the annotated dataset is used for training by the object detection unit, Faster R-CNN. Finally, the video to be detected is input into the trained object detection unit to perform detection.
[0072] FIG. 16 is a diagram showing an example of a multi-viewpoint image according to the fourth embodiment. Various 3D models of the cart were created using 3D CAD software, as shown in FIGS. 16(a) to 16(d). After the 3D models were created, the created 3D models were photographed from various viewpoints to create multi-viewpoint images in order to increase the variety of the data set. In this embodiment, images from 200 viewpoints were created.
[0073] FIG. 17 is a diagram showing combinations of datasets according to the fourth embodiment. Learning is performed by combining the generated dataset with a dataset annotated with a real cart. In FIG. 17, learning was performed using six different combinations of datasets. These combinations include one using only 250 images of real carts, one combining 200 images of real carts with 50 3D models, and one combining 150 images of real carts with 100 3D models.
[0074] As a result of an experiment on cart detection using 3D models, six different datasets were trained and detection was performed on multiple videos. When the training results were compared between datasets with 250 and 150 images annotated with real carts, more carts were detected when training with 3D models included than when training with only images annotated with real carts. Also, when there were no images of real carts, no carts were detected. Therefore, it was found that while including 3D models resulted in higher detection accuracy, data on real carts was also required.
[0075] According to the fourth embodiment, it has been found that it is preferable for the object detection unit (for example, the worker detection unit 12) to use annotation teaching data that combines a real model and a three-dimensional model in the learning process of object detection.
[0076] <Consideration of analysis speed> The speed of the task time measurement and calculation process for the first embodiment was verified using a compact, low-cost embedded computer, JetsonTX2 (TG731-PC), equipped with a GPU (Graphics Processing Unit). The processing time was examined based on the processing time required to input 1,000 frames of captured video. The frame rate was calculated by dividing the processing time by the number of input images (1,000). As a result, the processing speed for person detection alone was a maximum of 3.1 fps and a minimum of 2.8 fps, while the processing speed for person detection, IoU calculation, and work area determination was a maximum of 3.1 fps and a minimum of 2.7 fps. From the above, we confirmed that the task time measurement and calculation process shown in Figure 9 is capable of real-time processing, even though it is based on an input video of 1 fps, and confirmed the effectiveness of the monitoring system MS of this embodiment. [Explanation of symbols]
[0077] 10 Processing section 11 Video data storage unit 12 Worker detection unit 13 Work time measurement section 20 Memory section 21 Video Database 22 Work time measurement results 23 Work time results for each work process 30 Input section 40 Output section 50 Communications Department 60,61,62 Imaging device 70 input images 71 Backborn CNN 72 RPN 73 ROI Pooling 74,75 FC(Fully connected layer) 76 classes 77 BB (Bounding Box) 100 Analyzer A1, A2 work area IoU Intersection over Union NW Network R1, R2, R3, R4, R5 working area MS Monitoring System
Claims
1. an imaging device for imaging a plurality of work areas; an analysis device that analyzes the video data obtained by the imaging device, The analysis device a worker detection unit that detects a worker in a frame of the video data; a work time measurement unit that determines in which definition area of a work area the worker detected by the worker detection unit is present, and measures the work time of each work area by integrating the determined frames, the worker detection unit detects a worker in a frame of the video data as a rectangle; The work time measurement unit calculates an evaluation index indicating the degree of overlap between a detection rectangle, which is a rectangle corresponding to the worker detected by the worker detection unit, and a rectangle of a definition region of each work region, and when the definition regions partially overlap, determines that the evaluation index is equal to or greater than a predetermined threshold and that a worker is present in the region with the largest evaluation index. A monitoring system characterized by:
2. The work time measurement unit determines whether the worker is working by using a convolutional neural network.
2. The monitoring system according to claim 1.
3. An imaging device that images multiple work areas; an analysis device that analyzes the video data obtained by the imaging device, The analysis device a worker detection unit that detects a worker in a frame of the video data; a work time measurement unit that determines in which definition area of a work area the worker detected by the worker detection unit is present, and measures the work time of each work area by integrating the determined frames, the worker detection unit detects a worker in a frame of the video data as a rectangle; The work time measurement unit calculates an evaluation index indicating a degree of overlap between a detection rectangle, which is a rectangle corresponding to the worker detected by the worker detection unit, and a rectangle of a definition region of each work region, and when determining that a worker is present in an region where the evaluation index is equal to or greater than a predetermined threshold and where the evaluation index is the largest, When the defined regions overlap, the work time measurement unit measures a depth map of the worker by depth estimation, calculates a depth histogram of the depth map, and classifies the depth direction based on the depth histogram. A monitoring system characterized by:
4. The work time measurement unit sets a flag in each frame indicating whether the work area is being worked on, is not being worked on, or is outside the photographing range.
2. The monitoring system according to claim 1.
5. The worker detection unit, in the object detection learning process, Using annotation teaching data that combines real models and 3D models 2. The monitoring system according to claim 1.
6. The task time measurement unit measures task time using task start times and task end times for each task area based on the determined frames.
2. The monitoring system according to claim 1.
7. A monitoring method for a monitoring system having an image capturing device that captures images of a plurality of work areas and an analysis device that analyzes video data obtained by the image capturing device, comprising: The analysis device detects workers in frames of the video data as rectangles, calculates an evaluation index indicating the degree of overlap between a detection rectangle corresponding to the detected worker and a rectangle in a definition area of each work area, and if the definition areas partially overlap, determines that the worker is present in an area where the evaluation index is equal to or greater than a predetermined threshold and where the evaluation index is the largest, and measures the work time of each work area by integrating the determined frames. A monitoring method comprising:
8. The analysis device determines whether the worker is working using a convolutional neural network. The monitoring method according to claim 7 .
9. A monitoring method for a monitoring system having a camera that captures images of a plurality of work areas and an analysis device that analyzes video data obtained by the camera, comprising: The analysis device detects a worker in a frame of the video data as a rectangle, calculates an evaluation index indicating the degree of overlap between a detection rectangle corresponding to the detected worker and a rectangle in a definition area of each work area, determines that a worker is present in an area where the evaluation index is equal to or greater than a predetermined threshold and where the evaluation index is the largest, and adds up the determined frames to measure the work time of each work area, When the defined regions overlap, the analysis device measures a depth map of the worker by depth estimation, calculates a depth histogram of the depth map, and classifies the depth direction based on the depth histogram. A monitoring method comprising:
10. The analysis device sets a flag in each frame indicating whether each of the work areas is being worked on, is not being worked on, or is outside the imaging range. The monitoring method according to claim 7 .
11. The analysis device, in the object detection learning process, Using annotation teaching data that combines real models and 3D models The monitoring method according to claim 7 .
12. The analysis device measures the work time using the work start time and the work end time of each work area based on the determined frames. The monitoring method according to claim 7 .
Citation Information
Patent Citations
Stainless steel rolled wire and rod having fine grain structure and method of making same
JP1979016322A
Computer supported work analyzer using video equipment
JP1994231137A
Data set for learning functions with image as input
JP2019032820A
Machine tool, behavior type determination method and behavior type determination program
JP2020187708A
Device, method, and program for processing information, and recording medium
JP2020204819A