Operation Analysis Device

The work analysis device automates the adjustment of criteria for accurate work classification by integrating object detection and joint position estimation, addressing inefficiencies and accuracy issues in existing technologies.

JP7794952B2Active Publication Date: 2026-01-06FANUC LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2024511148
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2026-01-06
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

Existing work classification technologies require complex classification models with low interpretability, significant computational resources, and manual adjustment of criteria, leading to inefficiencies and accuracy issues in identifying worker tasks from video data.

Method used

A work analysis device that includes a work label assignment unit, object detection annotation unit, object detection learning unit, work judgment parameter calculation unit, and work judgment unit to automatically adjust and determine criteria for accurate work judgment, utilizing object detection models and joint position estimation to minimize errors in task identification.

Benefits of technology

Enables automatic adjustment of criteria for high-accuracy work classification, reducing manual effort and computational complexity while improving task identification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794952000001
    Figure 0007794952000001
  • Figure 0007794952000002
    Figure 0007794952000002
  • Figure 0007794952000003
    Figure 0007794952000003
Patent Text Reader

Abstract

This invention automatically adjusts and derives an assessment criterion (parameter) for accurately assessing work. This work analysis device comprises: a work labelling unit that assigns, to video data including work of a worker, a work label indicating the work of the worker; an object detection annotation unit that, with respect to the video data to which the work label was assigned, annotates an object related to the work of the worker; an object detection learning unit that generates an object detection model that performs object detection from video data of an object annotated by the object detection annotation unit; an object detection unit that uses the object detection model to detect an object from video data; a work assessment parameter calculation unit that assesses work in video data to which a work label was assigned, and calculates an assessment criterion which minimizes the error as compared to the assigned work label; and a work assessment unit that uses the object detection model and the assessment criterion to assess work of a worker in newly input video data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a work analysis device. [Background technology]

[0002] In factories, operational data on machine tools and other equipment can be obtained, but data on worker work cannot be obtained. Therefore, in order to improve work, consider introducing robots, and realize digital twins of factories, it is necessary to visualize the work of workers, and technology that can automatically recognize what workers were doing from video footage of their work is important. In this regard, there is known a technology that performs machine learning using learning data consisting of input data of images of workers performing tasks and label data of the tasks performed by the workers in the images, generates a trained model for identifying tasks from images, and uses the trained model to identify which task is being performed in the image being analyzed (see, for example, Patent Document 1). Also, there is known a technology for identifying the position of a worker's hand from depth-accelerated image data captured by a depth sensor, and for identifying the position of an object from image data captured by a digital camera, and for identifying the details of the action performed by the worker during work. For example, see Patent Document 2. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-67981 [Patent Document 2] International Publication No. 2017 / 222070 Summary of the Invention [Problem to be solved by the invention]

[0004] However, classification models such as the trained model in Patent Document 1 have the problem of being complex and having low interpretability. Furthermore, in order to detect the tools (objects) used in an image for work classification as in Patent Document 2, a large amount of calculation is required because the entire image must be scanned. Furthermore, to accurately identify the tasks that workers are performing, it is necessary to adjust the criteria (parameters) for task identification and manually search and annotate images of various task scenes, which is time-consuming. Additionally, there is also the issue of whether manual searching will improve the accuracy of task identification.

[0005] Therefore, there is a demand for a function that automatically adjusts and determines the criteria (parameters) for accurately judging work. [Means for solving the problem]

[0006] One aspect of the work analysis device disclosed herein is a work analysis device that analyzes the work of a worker, and includes: a work label assignment unit that assigns a work label indicating the work of the worker to video data including the work of the worker; an object detection annotation unit that annotates the video data to which the work label has been assigned, with an object related to the work of the worker; an object detection learning unit that generates an object detection model that detects objects from the video data of the objects annotated by the object detection annotation unit; an object detection unit that detects the objects from the video data using the object detection model; a work judgment parameter calculation unit that performs work judgment on the video data to which the work label has been assigned, and calculates judgment criteria that minimize an error with respect to the assigned work label; and a work judgment unit that judges the work of the worker in newly input video data using the object detection model and the judgment criteria. [Effects of the Invention]

[0007] According to one aspect, the criteria (parameters) for determining the work with high accuracy can be automatically adjusted and determined. [Brief explanation of the drawings]

[0008] [Figure 1]1 is a functional block diagram showing an example of the functional configuration of an operation analysis system according to a first embodiment. [Figure 2] FIG. 10 is a diagram illustrating an example of a working table. [Figure 3] FIG. 10 is a diagram illustrating an example of a user interface for assigning a work label. [Figure 4] 10A and 10B are diagrams showing examples of video data in different router states; [Figure 5] FIG. 10 is a diagram illustrating an example of a determination result of an operation determination. [Figure 6] FIG. 10 is a diagram illustrating an example of erroneous detection. [Figure 7] FIG. 2 is a diagram showing an example of an image area in video data. [Figure 8] FIG. 10 is a diagram illustrating an example of the operation of a moving object detection unit. [Figure 9] 10 is a flowchart illustrating a parameter calculation process of the work analysis device. [Figure 10] 10 is a flowchart illustrating an analysis process of the work analysis device. [Figure 11] FIG. 10 is a functional block diagram showing an example of the functional configuration of an operation analysis system according to a second embodiment. [Figure 12] FIG. 10 is a diagram showing an example of joint position information in a frame image. [Figure 13] FIG. 10 is a diagram illustrating an example of the operation of a joint position task estimation model. [Figure 14] 10 is a flowchart illustrating a parameter calculation process of the work analysis device. [Figure 15] 10 is a flowchart illustrating an analysis process of the work analysis device. DETAILED DESCRIPTION OF THE INVENTION

[0009] A first embodiment and a second embodiment of the work analysis device will be described in detail with reference to the drawings. Here, each embodiment has in common the configuration that a work label indicating the work of a worker is assigned to video data (video) in which the worker's work has been captured in advance, the worker annotates the video data to which the work label has been assigned an object (tool) related to the work, and an object detection model is generated that detects the object from the video data of the annotated object. However, in determining the work of a worker in the first embodiment, the generated object detection model is used to determine the work of a worker in video data to which a work label has been assigned, and a determination criterion that minimizes the error from the assigned work label is calculated, and the work of the worker in newly input video data is determined using the object detection model and the calculated determination criterion. In contrast, the second embodiment differs from the first embodiment in that joint position information about the worker's joints is estimated, a joint position work estimation model that estimates the work of the worker based on the estimated joint position information and the assigned work label is generated, and the determination criterion is calculated to minimize the error from the work label based on a value related to the accuracy of object detection in work determination using the object detection model and the classification probability of the work estimated from the joint position in work determination using the joint position work estimation model, and the work of the worker in newly input video data is determined using the object detection model, the joint position work estimation model, and the determination criterion. In the following, the first embodiment will be described in detail first, and then the second embodiment will be described, focusing on the differences from the first embodiment.

[0010] First Embodiment FIG. 1 is a functional block diagram showing an example of the functional configuration of the work analysis system according to the first embodiment. As shown in FIG. 1, the work analysis system 100 includes a work analysis device 1 and a camera 2.

[0011] The work analysis device 1 and camera 2 may be connected to each other via a network (not shown) such as a LAN (Local Area Network) or the Internet. In this case, the work analysis device 1 and camera 2 are equipped with a communication unit (not shown) for communicating with each other via such a connection. The work analysis device 1 and camera 2 may also be directly connected to each other via a connection interface (not shown) via a wired or wireless connection. Furthermore, in FIG. 1, the work analysis device 1 is connected to one camera 2, but it may be connected to two or more cameras 2.

[0012] The camera 2 is a digital camera or the like, and captures two-dimensional frame images of workers, tools, and other objects (not shown) projected onto a plane perpendicular to the optical axis of the camera 2 at a predetermined frame rate (e.g., 30 fps). The camera 2 outputs the captured frame images as video data to the work analysis device 1. Note that the video data captured by the camera 2 may be visible light images such as RGB color images, grayscale images, and depth images.

[0013] <Work analysis device 1> The work analysis device 1 is a computer known to those skilled in the art, and as shown in Fig. 1, includes a control unit 10 and a memory unit 20. The control unit 10 also includes an work registration unit 101, an work label assignment unit 102, an object detection and annotation unit 103, an object detection learning unit 104, an work judgment parameter calculation unit 105, an object detection and annotation proposal unit 106, and an work judgment unit 107. The work judgment unit 107 also includes an object detection unit 1071 and a moving object detection unit 1072.

[0014] The storage unit 20 is a storage device such as a ROM (Read Only Memory) or an HDD (Hard Disk Drive). The storage unit 20 stores an operating system and application programs executed by the control unit 10 (described later). The storage unit 20 also includes a video data storage unit 201, a work registration storage unit 202, and an input data storage unit 203.

[0015] The video data storage unit 201 stores video data of workers and objects such as tools captured by the camera 2.

[0016] The work registration memory unit 202 stores a work table that associates tools (objects) detected by an object detection unit 1071 (described later) with the corresponding work of the worker, which is registered in advance by the work registration unit 101 (described later) based on input operations by a user such as a worker via an input device (not shown), such as a keyboard or touch panel, included in the work analysis device 1. FIG. 2 is a diagram illustrating an example of the working table. As shown in FIG. 2, the work table has storage areas for "objects" and "work." In the storage area for "object" in the work table, tool names such as "Ryuta (registered trademark)" and "sandpaper" are stored. In the storage area for "tasks" in the work table, tasks such as "routing" and "sanding" are stored.

[0017] The input data storage unit 203 stores, for example, a set of frame image data in which a tool (object) annotated by the object detection annotation unit 103 described later is associated with the image range in which the tool appears among the frame images of the video data, as input data when the object detection learning unit 104 described later generates an object detection model.

[0018] The control unit 10 includes a CPU, a ROM, a RAM (Random Access Memory), a CMOS memory, and the like, which are configured to be able to communicate with each other via a bus, and are well known to those skilled in the art. The CPU is a processor that controls the entire work analysis device 1. The CPU reads system programs and application programs stored in ROM via the bus and controls the entire work analysis device 1 in accordance with the system programs and application programs. As shown in FIG. 1 , the control unit 10 is configured to implement the functions of a work registration unit 101, a work label assignment unit 102, an object detection and annotation unit 103, an object detection and learning unit 104, a work judgment parameter calculation unit 105, an object detection and annotation proposal unit 106, and a work judgment unit 107. The work judgment unit 107 is also configured to implement the functions of an object detection unit 1071 and a moving object detection unit 1072. The RAM stores various data, such as temporary calculation data and display data. The CMOS memory is backed up by a battery (not shown) and is configured as a non-volatile memory that retains its memory state even when the work analysis device 1 is powered off.

[0019] The work registration unit 101 registers, in the work table shown in FIG. 2, a relationship between a tool to be used (an object to be detected) and a work to be performed using the tool (an object) (a work to be recognized), for example, based on an input operation by a user such as a worker via an input device (not shown) of the work analysis device 1.

[0020] For example, when a user views video data (video data) including the work of a worker stored in the video data storage unit 201, the work label assignment unit 102 assigns a work label to the video data (video data) indicating what work the worker is doing and the name of the work. FIG. 3 is a diagram showing an example of a user interface 30 for assigning a work label. As shown in Figure 3, the user interface 30 has an area 301 for playing video data (video) stored in the video data storage unit 201, a play / stop button 302, a slide 303, an area 310 for chronologically displaying work labels assigned to the video data by the work label assignment unit 102, a router button 321 for displaying tools to be annotated by the object detection annotation unit 103 described later, a micro router button 322, a sandpaper button 323, a rag button 314, and a completion button 330 for completing the assignment of work labels and / or annotation of objects.

[0021] Specifically, the work label assignment unit 102 displays the user interface 30 on a display device (not shown) such as an LCD included in the work analysis device 1, and plays back the video data (moving image data) stored in the video data storage unit 201 in an area 301 of the user interface 30. The user operates a playback / stop button 302 or a slide 303 via an input device (not shown) of the work analysis device 1 to check the video data, and when the user confirms that the worker is performing the task of "routing" in the video data from 13:10 to 13:13, the user inputs the task name of "routing," and the work label assignment unit 102 assigns the task label of "routing" to the video data from 13:10 to 13:13. Furthermore, if a user observes a worker performing microrouting in the video data from 13:13 to 13:18, the user enters the work name "microrouting," and the work label assignment unit 102 assigns the work label "microrouting" to the video data from 13:13 to 13:18. If a user observes a worker performing "sanding" in the video data from 13:18 to 13:20, the user enters the work name "sanding," and the work label assignment unit 102 assigns the work label "sanding" to the video data from 13:18 to 13:20. If a user observes a worker performing "cleaning" in the video data from 13:20 to 13:22, the user enters the work name "cleaning," and the work label assignment unit 102 assigns the work label "cleaning" to the video data from 13:20 to 13:22. The work label assignment unit 102 may display the work label assignment results in area 310 in chronological order on a display device (not shown) of the work analysis device 1. Then, the work label assignment unit 102 outputs the video data to which the work labels have been assigned to the object detection and annotation unit 103.

[0022] The object detection and annotation unit 103 annotates, for example, tools (objects) related to the work of the worker in the video data to which the work label has been added. Specifically, the object detection annotation unit 103 displays, for example, in area 301 of the user interface 30, frame images (still images) separated at a predetermined interval that show a router tool (object) from video data that has been assigned the task label "router operation" from time 13:10 to time 13:13, or frame images (still images) separated at any interval specified by the user. It is preferable that the frame images (still images) to be displayed are set at predetermined intervals or at arbitrary intervals so that there are, for example, about 20 frames per work label. This eliminates the need for the user to check hours of video data, allowing the user to perform the work efficiently and reducing the burden on the user.

[0023] 3, based on the user's input operation, the object detection annotation unit 103 acquires the image range (bold rectangle) of the tool (object) appearing in each frame image (still image), and annotates the tool (object) as a router when the router button 321 or the like is pressed. Note that the object detection annotation unit 103 also acquires the image range of the tool (object) appearing in each frame image (still image) of video data to which the task labels "micro-routing," "sanding," and "cleaning" are assigned, and annotates the tool (object), just as in the case of "routing." When the object detection annotation unit 103 completes the annotation of the image range containing the tool (object) and the tool (object) for all frame images (still images) of the video data to which a work label has been assigned, and the user presses the complete button 330, the object detection annotation unit 103 stores in the input data storage unit 203 a set of frame image data (hereinafter also referred to as "annotated frame image data") that associates the image range of the frame image (still image) containing the tool (timestamped) with the annotated tool (object) from the video data (video data) for the time when each work was performed (the time from the start to the end of the work).

[0024] The object detection learning unit 104 generates an object detection model for detecting objects from video data of annotated objects. Specifically, the object detection learning unit 104 performs known machine learning using, for example, annotated frame image data stored in the input data storage unit 203 as input data and teacher data in which the annotated tools (objects) are used as label data, to generate an object detection model that is a trained model such as a neural network. The object detection learning unit 104 stores the generated object detection model in the storage unit 20.

[0025] The task judgment parameter calculation unit 105 uses the object detection model generated by the object detection learning unit 104 to perform task judgment on the video data to which task labels have been assigned, and calculates a judgment criterion that minimizes the error with the assigned task label. Specifically, the task determination parameter calculation unit 105 sets initial values ​​of parameters as determination criteria for each task registered in the task table of Fig. 2. The parameters include, for example, the number of seconds X that indicates that a task is being performed for X seconds after an object is detected, a threshold value related to the accuracy of object detection for determining that the task of "routing" is being performed, and a threshold value related to the accuracy of object detection for determining that the task of "sanding" is being performed. By including the number of seconds X that indicates that a task is being performed for X seconds after an object is detected in the parameters, the task analysis device 1 can, for example, determine that a task using the tool is being performed if the tool (object) has been detected within the last X seconds, even if the task cannot be detected using the video data.

[0026] The task determination parameter calculation unit 105 inputs annotated frame image data from other video data to which task labels have been assigned and stored in the input data storage unit 203 into an object detection model to detect tools (objects). The task determination parameter calculation unit 105 determines the task based on the object detection results and the task table of FIG. 2, and calculates the error between the determined task and the correct task label. The task determination parameter calculation unit 105 then calculates an evaluation index, such as an F1 score for the parameter value, for each task based on the error calculated for all annotated frame image data, and calculates parameter values ​​for each task using Bayesian optimization or the like so as to maximize the calculated evaluation index for each task.

[0027] The object detection annotation suggestion unit 106 uses the parameters (judgment criteria) calculated by the task judgment parameter calculation unit 105 to perform task judgment on the video data to which task labels have been assigned, and proposes a frame image (still image) to be annotated based on the judgment result of the task judgment. For example, in the case of object detection of a router, if the system is trained only on video data of the router being held by a worker, as shown in the upper part of Figure 4, the accuracy of object detection of a router placed on a workbench, as shown in the lower part of Figure 4, will decrease. Therefore, video data annotated in a wide variety of scenes is needed, but it is time-consuming for users to search for various scenes. Therefore, the object detection annotation suggestion unit 106 suggests frame images (still images) that are suitable for automatic annotation, as described below. Specifically, the object detection annotation proposal unit 106 performs work judgment using image data that associates an annotated tool (object) with the image range in which the tool is captured in another video data that has a work label attached and is stored in the input data storage unit 203. Fig. 5 is a diagram showing an example of the determination result of task determination. The upper part of Fig. 5 shows the time series of correct task labels assigned to the different video data. The middle part of Fig. 5 shows the determination result of the worker's task for the image data by the object detection annotation proposing unit 106 using the object detection model and parameters. The lower part of Fig. 5 shows the object detection result for the image data using the object detection model.

[0028] As shown in FIG. 5, there is a period between 13:40 and 13:43 when the correct task label is "switching," in which "switching" is not determined (detected) in the task assessment results. This is because, for example, a frame image (still image) showing a switch was not extracted even though it existed between the time X seconds after the parameter object was detected and the time X seconds after the task was performed. Therefore, in order to increase the value related to the accuracy of object detection, the object detection annotation proposal unit 106 extracts a frame image (still image) showing a switch from other video data during the period when "switching" was not determined (detected) in the task assessment results. Similar to the object detection annotation unit 103, the object detection annotation proposal unit 106 displays the extracted frame image (still image) on the user interface 30, acquires the image area of ​​the switch in the extracted frame image (still image) based on the user's input operation, and annotates it as "switching" when the switch button 321 is pressed. The object detection annotation proposing unit 106 stores in the input data storage unit 203 image data that associates the image range of a frame image (still image) in which the router is shown (to which a timestamp has been added) with the annotated router.

[0029] Furthermore, as shown in Figure 5, during the time period from 13:43 to 13:50, when the correct task label was "sanding," there were times when the task was erroneously determined (misdetected) as "rubbing paper" in the task determination results, and times when "sanding" was not determined (detected). This erroneous determination (misdetection) of "rubbing paper" was caused by a false detection of a router in object detection for the frame image (still image) at 13:43. Furthermore, the reason "sanding" was not determined (detected) was because a frame image (still image) showing sandpaper was present but not extracted between the time the parameter object was detected and the number of seconds X, which is the time the task is assumed to be performed for X seconds. Therefore, in order to increase the value related to the accuracy of object detection, the object detection annotation proposal unit 106 extracts, from the other video data, a frame image (still image) in which sandpaper appears around 13:43 and a frame image (still image) in which sandpaper appears during a time when "sanding" was not determined (detected). The object detection annotation proposal unit 106 displays each of the extracted frame images (still images) on the user interface 30, acquires the image range of sandpaper in each frame image (still image) based on a user's input operation, and annotates the tool (object) as sandpaper when the sandpaper button 323 is pressed. The object detection annotation proposal unit 106 stores image data in the input data storage unit 203 that associates the image range of the frame image (still image) in which sandpaper appears (to which a timestamp has been added) with the annotated sandpaper. This allows for improved accuracy in object detection without requiring the user to go to the trouble of searching through various scenes.

[0030] Note that the object detection annotation proposing unit 106 may extract a frame image (still image) in which the tool (object) appears, even when the reliability of object detection of the tool (object), which is a value related to the accuracy of object detection, is low, equal to or lower than a predetermined value (e.g., 20%). The object detection annotation proposing unit 106 may display the extracted frame image (still image) on the user interface 30, acquire the image range of the tool (object) in the extracted frame image (still image) based on an input operation by the user, and annotate the tool (object).

[0031] Thereafter, the object detection learning unit 104 performs machine learning using image data including frame images (still images) extracted (proposed) by the object detection annotation proposal unit 106 and annotated with tools (objects), and updates the object detection model. The task determination parameter calculation unit 105 inputs the annotated frame image data including the frame images (still images) extracted (proposed) by the object detection annotation proposal unit 106 into the updated object detection model to determine the task and calculates the error between the assigned correct task label and the task determination result. The task determination parameter calculation unit 105 calculates an evaluation index, such as an F1 score of the parameter value, for each task based on the calculated error, and recalculates the parameter value for each task using Bayesian optimization or the like so as to maximize the calculated evaluation index for each task. For example, the object detection learning unit 104 and the task determination parameter calculation unit 105 repeat the process until there are no more frame images (still images) extracted (proposed) by the object detection annotation proposal unit 106, or until the number of frame images (still images) extracted (proposed) by the object detection annotation proposal unit 106 falls below a predetermined number. Then, object detection learning unit 104 outputs the generated object detection model to object detection unit 1071, which will be described later, and activity judgment parameter calculation unit 105 outputs the calculated parameters to activity judgment unit 107, which will be described later.

[0032] The work determination unit 107 determines the work of the worker in the video data newly input from the camera 2 using the object detection model and the set parameters (determination criteria). Specifically, the work determination unit 107 inputs, for example, a frame image (still image) of video data newly input from the camera 2 to an object detection model of the object detection unit 1071 (described later) and to a moving object detection unit 1072 (described later). The work determination unit 107 determines the work of the worker based on the tool (object) detection result output from the object detection model, the detection result of the moving object detection unit 1072, the work table of Fig. 2, and parameters. Note that if the work determination unit 107 cannot detect a tool (object) from a frame image (still image) of video data, but the tool (object) was detected immediately prior to that frame image within X seconds, the work determination unit 107 may determine the work of the worker in that frame image based on a parameter X, which is the number of seconds X that indicates that work has been performed for X seconds since the object was detected. Furthermore, the work determination unit 107 may determine that the work of the worker is "no work" when, for example, a value related to the object detection accuracy, such as the object detection reliability or class classification probability, output from the object detection model of the object detection unit 1071 is equal to or lower than a preset threshold (for example, 70%). For example, as shown in Fig. 6, when the work determination unit 107 receives an object detection result of "sandpaper" and "confidence 40%" in a case where the worker is simply touching a workpiece in the video data, the work determination unit 107 may determine that the work of the worker is "no work" because the reliability is equal to or lower than a threshold (for example, 70%). This will reduce false positives of tasks.

[0033] The object detection unit 1071 has an object detection model generated by the object detection learning unit 104, inputs frame images (still images) of newly input video data from the camera 2 into the object detection model, and outputs values ​​related to the accuracy of object detection, such as reliability, along with the detection results of the tool (object).

[0034] The moving object detection unit 1072 detects moving objects such as workers and tools based on changes in pixel brightness and the like in a specified image area of ​​each frame image (still image) of the video data newly input from the camera 2. Specifically, the moving object detection unit 1072 may determine that a worker in the video data is working if there is movement such as a change in pixel brightness in the image area indicated by the thick rectangle in the frame image (still image) as shown in Figure 7. Furthermore, the moving object detection unit 1072 may determine that the worker is working continuously if it detects movement periodically at intervals of less than X seconds (for example, 5 seconds) as shown in the dashed rectangle, as shown in the upper part of Fig. 8. Then, as shown in the lower part of Fig. 8, if the object detection unit 1071 detects a tool (object) such as a router from a frame image (still image) at a time indicated by a shaded rectangle during a period in which movement of a moving object is detected, the moving object detection unit 1072 may determine that work is being performed using the detected tool (object) during that period. On the other hand, if the moving object detection unit 1072 does not detect any movement for more than X seconds, it may determine that the worker is not working.

[0035] <Parameter calculation process of the work analysis device 1> Next, the operation of the parameter calculation process of the work analysis device 1 according to the first embodiment will be described. 9 is a flowchart illustrating the parameter calculation process of the work analysis device 1. The flow shown here is executed when a user such as a worker registers a new tool (object) and work in the work table.

[0036] In step S1, the work label assignment unit 102 plays back the video data including the work of the worker stored in the video data storage unit 201 on the user interface 30, and assigns a work label indicating the work being performed by the worker to the video data based on input operations by the user.

[0037] In step S2, the object detection and annotation unit 103 acquires the image range of the tool (object) appearing in frame images (still images) separated at predetermined intervals for each work label from the video data to which work labels have been assigned in step S1, and annotates the tool (object). The object detection and annotation unit 103 stores annotated frame image data in the input data storage unit 203, which associates the image range of the frame image (still image) in which the tool appears (with a timestamp assigned) from the video data (video data) for the time when each work was performed (the time from the start to the end of the work) with the annotated tool (object).

[0038] In step S3, the object detection learning unit 104 generates an object detection model for detecting objects from the annotated frame image data annotated in step S2.

[0039] In step S4, the task determination parameter calculation unit 105 inputs annotated frame image data of other video data to which task labels have been assigned and which are stored in the input data storage unit 203 into the object detection model, and detects the tool (object).

[0040] In step S5, the task determination parameter calculation unit 105 determines the task performed by the worker based on the object detection result in step S4 and the task table.

[0041] In step S6, the task judgment parameter calculation unit 105 calculates the error between the correct task label and the judgment result in step S5 for each task.

[0042] In step S7, an evaluation index such as an F1 score of the parameter value is calculated for each task based on the errors calculated for all video data.

[0043] In step S8, the task judgment parameter calculation unit 105 calculates the parameters for each task by Bayesian optimization or the like so as to maximize the evaluation index for each task.

[0044] In step S9, the object detection annotation proposing unit 106 performs a task determination on the other video data to which the task label has been assigned, using the parameters (determination criteria) calculated in step S8.

[0045] In step S10, the object detection annotation proposing unit 106 determines, based on the determination result of step S9, whether there is a frame image (still image) to propose in order to increase the value related to the object detection accuracy in locations where the value related to the object detection accuracy is low, such as erroneous detection or non-detection. If there is a frame image (still image) to propose, the process returns to step S2, and the processes of steps S2 to S9 are performed again, including the proposed frame image (still image). On the other hand, if there is no frame image (still image) to propose, the work analysis device 1 sets the object detection model generated in step S3 in the object detection unit 1071, sets the parameters calculated in step S8 in the work determination unit 107, and ends the parameter calculation process.

[0046] <Analysis process of work analysis device 1> Next, the operation of the analysis process of the work analysis device 1 according to the first embodiment will be described. 10 is a flowchart illustrating the analysis process of the work analysis device 1. The flow shown here is repeatedly executed while video data is being input from the camera 2.

[0047] In step S21, the object detection unit 1071 inputs a frame image (still image) of video data newly input from the camera 2 into the object detection model to detect a tool (object).

[0048] In step S22, the moving object detection unit 1072 detects moving objects such as workers and tools from changes in pixel brightness and the like in a specified image area of ​​each frame image (still image) of the video data newly input from the camera 2.

[0049] In step S23, the work determination unit 107 determines the work of the worker based on the tool (object) detection result in step S21, the moving object detection result in step S22, the set parameters, and the work table.

[0050] As described above, the work analysis device 1 according to the first embodiment can automatically adjust the criteria to accurately judge the work. In other words, the optimal parameters are calculated automatically as long as the user simply labels the work and annotates the objects. Furthermore, when the accuracy of task determination is insufficient, the task analysis device 1 can automatically suggest frames in a video that, if annotated, will improve the accuracy of task determination. The first embodiment has been described above.

[0051] Second Embodiment Next, a second embodiment will be described. In the first embodiment, a generated object detection model is used to perform a task determination of a worker's task in video data to which a task label has been assigned, and a determination criterion that minimizes the error from the assigned task label is calculated. The task of the worker in newly input video data is then determined using the object detection model and the calculated determination criterion. In contrast, the second embodiment differs from the first embodiment in that joint position information about the worker's joints is estimated, a joint position task estimation model is generated that estimates the worker's task based on the estimated joint position information and the assigned task label, and a determination criterion is calculated to minimize the error from the task label based on a value related to the accuracy of object detection in task determination using the object detection model and the task classification probability estimated from the joint position in task determination using the joint position task estimation model. The object detection model, the joint position task estimation model, and the determination criterion are used to determine the task of the worker in newly input video data. As a result, the work analysis device 1A according to the second embodiment can automatically adjust the judgment criteria to judge the work with high accuracy. The second embodiment will be described below.

[0052] Fig. 11 is a functional block diagram showing an example of the functional configuration of the work analysis system according to the second embodiment. Elements having the same functions as those of the work analysis system 100 in Fig. 1 are given the same reference numerals, and detailed descriptions thereof will be omitted. As shown in FIG. 11, the work analysis system 100 includes a work analysis device 1A and a camera 2. The camera 2 has the same functions as the camera 2 in the first embodiment.

[0053] <Work analysis device 1A> 11, the work analysis device 1A includes a control unit 10a and a memory unit 20. The control unit 10a also includes a work registration unit 101, a work label assignment unit 102, an object detection and annotation unit 103, an object detection and learning unit 104, a work judgment parameter calculation unit 105a, a joint position estimation unit 108, a joint position work learning unit 109, and a work judgment unit 107a. The work judgment unit 107a also includes an object detection unit 1071, a moving object detection unit 1072, and a joint position work estimation unit 1073. The memory unit 20 also includes a video data memory unit 201, a work registration memory unit 202, and an input data memory unit 203. The memory unit 20, the video data memory unit 201, the work registration memory unit 202, and the input data memory unit 203 have functions equivalent to those of the memory unit 20, the video data memory unit 201, the work registration memory unit 202, and the input data memory unit 203 in the first embodiment. In addition, the work registration unit 101, the work label assignment unit 102, the object detection annotation unit 103, and the object detection learning unit 104 have functions equivalent to those of the work registration unit 101, the work label assignment unit 102, the object detection annotation unit 103, and the object detection learning unit 104 in the first embodiment. Moreover, the object detection unit 1071 and the moving object detection unit 1072 have the same functions as the object detection unit 1071 and the moving object detection unit 1072 in the first embodiment.

[0054] The joint position estimation unit 108 estimates joint position information related to the joint positions of the worker for each frame image (still image) of the video data to which the task label has been assigned and which is stored in the input data storage unit 203. Note that the frame images may be extracted from the video data at appropriate intervals. For example, if the frame rate of the video data is 60 fps, the frame images may be extracted at approximately 24 fps. Specifically, the joint position estimation unit 108 uses a known method (e.g., Kosuke Kanno, Kenta Oku, and Kyoji Kawagoe, "Motion Detection and Classification Method from Multidimensional Time Series Data," DEIM Forum 2016 G4-5, or Shohei Uezono and Satoshi Ono, "Feature Extraction of Multimodal Sequence Data Using LSTM Autoencoder," Materials from the Study Group of the Japanese Society for Artificial Intelligence, SIG-KBS-B802-01, 2018) to estimate, as joint position information, time series data such as the coordinates and angles of the joints of the worker's hands, arms, etc., for each frame image (still image) of video data to which task labels have been assigned and which is stored in the input data storage unit 203. 12 is a diagram showing an example of joint position information in a frame image, in which the joint position information is shown when the worker is sanding.

[0055] The joint position task learning unit 109 performs machine learning using, for example, the joint position information estimated by the joint position estimation unit 108 as input data and the task labels assigned by the task label assignment unit 102 as label data, to generate a joint position task estimation model that estimates the tasks of the worker. For example, the joint position work learning unit 109 generates a joint position work estimation model so that when the joint position information of the worker's right hand in FIG. 12 shows one back and forth motion in 0.3 seconds or so as shown in FIG. 13, it is determined that the worker is sanding. The joint position task learning unit 109 may generate a rule base based on the joint position information estimated by the joint position estimation unit 108 and the task labels assigned by the task label assignment unit 102.

[0056] The task determination parameter calculation unit 105a calculates a determination criterion (parameter) based on a value related to the accuracy of object detection in task determination using an object detection model and the task classification probability estimated from the joint position in task determination using a joint position task estimation model, so as to minimize the error with the task label. Specifically, the task determination parameter calculation unit 105a, for example, similarly to the task determination parameter calculation unit 105 of the first embodiment, sets initial values ​​of parameters as determination criteria for each task registered in the task table of FIG. 2. The task determination parameter calculation unit 105a inputs annotated frame image data of other video data to which task labels have been assigned and stored in the input data storage unit 203 into an object detection model, detects tools (objects), and acquires values ​​related to object detection. The task determination parameter calculation unit 105a determines the task of the worker based on the object detection results and the task table of FIG. 2. The task determination parameter calculation unit 105a also estimates joint position information of the worker for each frame image (still image) of the same other video data and inputs the estimated joint position information into the joint position task estimation model, thereby estimating the task of the worker and acquiring a classification probability estimated from the joint positions. Then, the task judgment parameter calculation unit 105a calculates the value of the parameter (judgment criterion) by Bayesian optimization or the like so as to minimize the error between the task classification probability calculated using the following equation (1), where a is the weighting coefficient of the task classification probability estimated from the joint position and b is the weighting coefficient of the value related to the accuracy of object detection, and the correct task label. Task classification probability = a(task classification probability estimated from joint positions) + b (value related to the accuracy of object detection) (1) Here, the parameters include, for example, the number of seconds X for which the task is assumed to be performed for X seconds after object detection, a weight a for the task classification probability estimated from the joint position, and a weight b for the value related to the accuracy of object detection. The work judgment parameter calculation unit 105a outputs the calculated parameters to the work judgment unit 107a (described later) and sets them therein.

[0057] The task determination unit 107a determines the task of the worker in the video data newly input from the camera 2 using the object detection model, the joint position task estimation model, and the set parameters (determination criteria). Specifically, the work determination unit 107a inputs, for example, frame images (still images) of video data newly input from the camera 2 to an object detection model in the object detection unit 1071 and to the moving object detection unit 1072. The work determination unit 107a determines the work of the worker based on the detected tool (object), the work table of FIG. 2, and parameters, and acquires a value related to the accuracy of object detection. The work determination unit 107a also estimates joint position information of the worker for each frame image (still image) of the same newly input video data, and inputs the estimated joint position information to a joint position work estimation model in the joint position work estimation unit 1073, which will be described later. The work determination unit 107a acquires the estimation results of the worker's work and the classification probability of the work estimated from the joint positions from the joint position work estimation unit 1073, which will be described later. Then, the task determination unit 107a calculates the task classification probability from the task classification probability estimated from the acquired joint positions and values ​​related to the accuracy of object detection, the set parameters, and equation (1), and determines the task performed by the worker based on the calculated classification probability and the detection result of the moving object detection unit 1072.

[0058] The joint position work estimation unit 1073 has a joint position work estimation model generated by the joint position work learning unit 109, inputs the joint position information estimated by the work determination unit 107a to the joint position work estimation model, and outputs the estimation result of the worker's work and the classification probability of the work estimated from the joint position to the work determination unit 107a.

[0059] <Parameter Calculation Process of Work Analysis Device 1A> Next, the operation of the parameter calculation process of the work analysis device 1A according to the second embodiment will be described. 14 is a flowchart illustrating the parameter calculation process of the work analysis device 1A. Note that the processes from step S31 to step S33 are the same as the processes from step S1 to step S3 in FIG. 9, and detailed description thereof will be omitted.

[0060] In step S34, the joint position estimation unit 108 estimates the joint position information of the worker for each frame image (still image) of the video data to which the task label has been assigned and which has been stored in the input data storage unit 203.

[0061] In step S35, the joint position task learning unit 109 performs machine learning using the joint position information estimated in step S34 as input data and the task labels assigned in step S31 as label data, to generate a joint position task estimation model that estimates the task of the worker.

[0062] In step S36, the work judgment parameter calculation unit 105a inputs annotated frame image data of another video data to which a work label has been assigned and stored in the input data storage unit 203 into the object detection model, and obtains the detected tool (object) and values ​​related to the accuracy of object detection.

[0063] In step S37, the task determination parameter calculation unit 105a determines the task performed by the worker based on the object detection result in step S36 and the task table.

[0064] In step S38, the task judgment parameter calculation unit 105a estimates the joint position information of the worker from a frame image (still image) of the same different video data.

[0065] In step S39, the task determination parameter calculation unit 105a inputs the joint position information estimated in step S38 into the joint position task estimation model, and obtains the estimation result of the worker's task and the classification probability estimated from the joint position.

[0066] In step S40, the task judgment parameter calculation unit 105a calculates the value of the parameter (judgment criterion) by Bayesian optimization or the like so as to minimize the error between the task classification probability calculated by equation (1) and the correct task label.

[0067] <Analysis process of work analysis device 1A> Next, the operation of the analysis process of the work analysis device 1A according to the second embodiment will be described. 15 is a flowchart illustrating the analysis process of the work analysis device 1 A. The flow shown here is repeatedly executed while video data is being input from the camera 2.

[0068] In step S51, the object detection unit 1071 inputs a frame image (still image) of video data newly input from the camera 2 into the object detection model, detects a tool (object), and acquires a value related to the accuracy of object detection.

[0069] In step S52, the moving object detection unit 1072 detects moving objects such as workers and tools from changes in pixel brightness and the like in a specified image area of ​​each frame image (still image) of the video data newly input from the camera 2.

[0070] In step S53, the joint position operation estimation unit 1073 estimates the joint position information of the worker for each frame image (still image) of the newly input video data.

[0071] In step S54, the joint position task estimation unit 1073 inputs the joint position information estimated in step S53 into the joint position task estimation model to estimate the task of the worker and obtain the task classification probability estimated from the joint position.

[0072] In step S55, the task determination unit 107a calculates the task classification probability from the task classification probability estimated from the joint positions acquired in steps S51 and S54 and values ​​related to the object detection accuracy, the moving object detection result in step S52, the set parameters, and formula (1), and determines the task of the worker based on the calculated classification probability.

[0073] As described above, the work analysis device 1A according to the second embodiment can automatically adjust the criteria to accurately judge the work. In other words, the optimal parameters are calculated automatically as long as the user simply labels the work and annotates the objects. The second embodiment has been described above.

[0074] The first and second embodiments have been described above, but the work analysis devices 1, 1A are not limited to the above-described embodiments and include modifications and improvements within the scope of achieving the objectives.

[0075] <Variation 1> In the first and second embodiments, the work analysis apparatus 1, 1A is connected to one camera 2, but this is not limiting. For example, the work analysis apparatus 1, 1A may be connected to two or more cameras 2.

[0076] <Variation 2> In the above-described embodiments, the task analysis devices 1 and 1A have all the functions, but this is not limiting. For example, a server may include some or all of the task registration unit 101, task label assignment unit 102, object detection and annotation unit 103, object detection and learning unit 104, task judgment parameter calculation unit 105, object detection and annotation proposal unit 106, task judgment unit 107, object detection unit 1071, and moving object detection unit 1072 of the task analysis device 1, or some or all of the task registration unit 101, task label assignment unit 102, object detection and annotation unit 103, object detection and learning unit 104, task judgment parameter calculation unit 105a, joint position estimation unit 108, joint position task learning unit 109, task judgment unit 107a, object detection unit 1071, moving object detection unit 1072, and joint position task estimation unit 1073 of the task analysis device 1A. Furthermore, the functions of the work analysis devices 1, 1A may be realized using a virtual server function or the like on the cloud. Furthermore, the work analysis devices 1, 1A may be configured as a distributed processing system in which the functions of the work analysis devices 1, 1A are distributed among a plurality of servers as appropriate.

[0077] <Variation 3> Furthermore, for example, in the above-described embodiment, the work analysis apparatus 1A does not include the object detection annotation proposing unit 106, but may include the object detection annotation proposing unit 106. By doing so, when the accuracy of task determination is insufficient, the task analysis device 1A can automatically suggest frames in a video that, if annotated, will improve the accuracy of task determination.

[0078] Note that the functions included in the work analysis devices 1 and 1A in the first and second embodiments can be realized by hardware, software, or a combination of these. Here, "realized by software" means that the functions are realized by a computer reading and executing a program.

[0079] The program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can be supplied to a computer via wired communication paths such as electric wires and optical fibers, or via wireless communication paths.

[0080] In addition, the steps of writing a program to be recorded on a recording medium include not only processes that are performed chronologically in accordance with the order, but also processes that are not necessarily performed chronologically but are performed in parallel or individually.

[0081] In other words, the work analysis device of the present disclosure can take on a variety of different embodiments having the following configurations.

[0082] (1) The work analysis device 1 disclosed herein is a work analysis device that analyzes the work of a worker, and includes: a work label assignment unit 102 that assigns a work label indicating the work of the worker to video data including the work of the worker; an object detection annotation unit 103 that annotates objects related to the work of the worker to the video data to which the work label has been assigned; an object detection learning unit 104 that generates an object detection model that detects objects from the video data of objects annotated by the object detection annotation unit 103; an object detection unit 1071 that detects objects from the video data using the object detection model; a work judgment parameter calculation unit 105 that performs work judgment on the video data to which the work label has been assigned and calculates judgment criteria that minimize the error with the assigned work label; and a work judgment unit 107 that judges the work of the worker in newly input video data using the object detection model and judgment criteria. According to this work analysis device 1, it is possible to automatically adjust the evaluation criteria in order to accurately evaluate the work.

[0083] (2) The work analysis device 1 described in (1) may be provided with an object detection annotation suggestion unit 106 that uses the judgment criteria calculated by the work judgment parameter calculation unit 105 to perform work judgment on video data to which work labels have been assigned, and suggests frame images to be annotated based on the judgment results of the work judgment. By doing so, when the accuracy of task determination is insufficient, the task analysis device 1 can automatically suggest frame images in a video that, if annotated, will improve the accuracy of task determination.

[0084] (3) The task analysis device 1A described in (1) or (2) may include a joint position estimation unit 108 that estimates joint position information related to the joint positions of the worker, a joint position task learning unit 109 that creates a joint position task estimation model that estimates the task of the worker based on the joint position information estimated by the joint position estimation unit 108 and the task label information assigned by the task label assignment unit 102, and a joint position task estimation unit 1073 that estimates the task from the joint position information based on the joint position task estimation model created by the joint position task learning unit 109, wherein the task judgment parameter calculation unit 105a calculates a judgment criterion that minimizes the error with the task label based on a value related to the accuracy of object detection in task judgment using the object detection model and the task classification probability estimated from the joint positions in task judgment using the joint position task estimation model, and the task judgment unit 107a may judge the task of the worker in newly input video data using the object detection model, the joint position task estimation model, and the judgment criterion. By doing so, the work analysis device 1A can achieve the same effect as (1).

[0085] (4) The work analysis device 1, 1A described in any one of (1) to (3) may further include a moving object detection unit 1072 that detects a moving object in newly input video data, and the work judgment unit 107, 107a may judge whether the worker's work is continuing based on the time interval at which the moving object detection unit 1072 detects the moving object. By doing so, the work analysis device 1, 1A can more accurately determine the work of the worker.

[0086] (5) In the work analysis device 1, 1A described in any one of (1) to (4), the judgment criteria may include at least the time for which it can be estimated that work using the tool (object) is continuing after the tool (object) is detected, and a threshold value related to the accuracy of object detection. By doing so, the work analysis device 1, 1A can accurately determine the work of the worker even when a tool (object) is not detected. [Explanation of symbols]

[0087] 1, 1A work analyzer 2 Cameras 10, 10a Control section 101 Work Registration Department 102 Working label assignment unit 103 Object detection and annotation part 104 Object detection learning unit 105, 105a Work judgment parameter calculation unit 106 Object detection annotation suggestion unit 107, 107a Work judgment section 1071 Object detection unit 1072 Motion detection unit 1073 Joint position and task estimation unit 108 Joint position estimation unit 109 Joint Position Task Learning Section 20 Memory section 201 Video data storage unit 202 Work registration memory unit 203 Input data storage unit

Claims

1. A work analysis device that analyzes work performed by a worker using an object including a tool, comprising: a task label assigning unit that assigns task labels to video data captured at a predetermined frame rate, the task labels identifying tasks performed by the worker using an object including the tool, and the tasks being performed by the worker using an object including the tool, to video data during a time period in which the tasks can be confirmed; an object detection and annotation unit that annotates, from the video data to which the work label has been assigned, an object including the tool related to the work of the worker for each frame image data that is a still image separated at predetermined intervals and in which an object including the tool to which the work label has been assigned is shown; an object detection learning unit that uses frame image data, which is the still image showing an object including the tool annotated by the object detection annotation unit, as input data, and generates an object detection model that performs object detection from training data in which the names of objects including the annotated tool are used as label data; an object detection unit that detects objects including the tool from the frame image data using the object detection model; a task judgment parameter calculation unit that performs task judgment on the video data to which the task label is assigned based on the detection results of objects, including tools, detected by the object detection unit and parameters set as judgment criteria, and calculates judgment criteria including the parameters that minimize an error with the assigned task label; and an operation determination unit that determines an operation performed by the worker in newly input video data using the object detection model and the determination criterion; A work analysis device comprising:

2. 2. The work analysis device according to claim 1, further comprising: an object detection annotation suggestion unit that uses the judgment criteria calculated by the work judgment parameter calculation unit to perform work judgment on the video data to which the work labels have been assigned, and proposes frame images to be annotated based on the judgment results of the work judgment.

3. a joint position estimation unit that estimates joint position information related to joint positions of the worker; a joint position task learning unit that creates a joint position task estimation model that estimates the task of the worker based on the joint position information estimated by the joint position estimation unit and the task label information assigned by the task label assignment unit; and a joint position and task estimation unit that estimates tasks from the joint position information based on the joint position and task estimation model created by the joint position and task learning unit, the task determination parameter calculation unit calculates the determination criterion based on a value related to accuracy of object detection in the task determination using the object detection model and a task classification probability estimated from joint positions in the task determination using the joint position task estimation model, so as to minimize an error with respect to the task label; 3 . The work analysis device according to claim 1 , wherein the work determination unit determines the work performed by the worker in newly input video data by using the object detection model, the joint position work estimation model, and the determination criterion.

4. a moving object detection unit that detects a moving object in the newly input video data; 4. The work analysis device according to claim 1, wherein the work determination unit determines whether the work of the worker is continuing based on a time interval at which the moving object detection unit detects the moving object.

5. 5. The work analysis device according to claim 1, wherein the criteria include at least a time period during which it can be estimated that work using the object is continuing after the object is detected, and a threshold value related to the accuracy of object detection.

Citation Information

Patent Citations

  • Information device, program and method for estimating usage of food or seasoning

    JP2020135417A

  • Work analysis device and work analysis method

    JP2021067981A

  • Germination determination device and program

    JP2021157548A

  • Work analysis device, work analysis method, and computer-readable recording medium

    WO2017222070A1

  • Teaching signal generation device, model generation device, object detection device, teaching signal generation method, model generation method, and program

    WO2020101036A1