Task determination apparatus, task determination method, and task determination program
The work determination device accurately determines work completion by analyzing displacement data of detected objects across multiple images, addressing the challenge of hidden objects and enhancing precision.
Patent Information
- Application Number
- JP2024003967
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-28
AI Technical Summary
Existing work determination devices struggle to accurately determine the completion of work when the object being worked on is hidden by the operator's hand.
A work determination device that acquires first and second images of a work process, detects objects in each image, and uses a trained determination model to analyze displacement data of the objects' positions to determine completion, utilizing object detection algorithms like YOLO and neural networks for accurate determination.
Enables precise determination of work completion by analyzing displacement features of objects across images, regardless of the object's visibility, reducing errors and improving accuracy.
Smart Images

Figure 2025110181000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a work determination device, a work determination method, and a work determination program.
Background Art
[0002] In recent years, there has been proposed a work determination device that photographs the work of an operator or a work robot in a factory or the like and determines whether the work has been appropriately completed based on the image.
[0003] For example, Patent Document 1 discloses a work determination device that accurately determines whether work is being performed correctly and prevents work errors and omissions. This work determination device determines whether a part held by an operator is a part to be attached to a workpiece in the work.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, since the work determination device of Patent Document 1 determines the completion of work based on an image in which an operator holds a part, it has been difficult to accurately determine the completion of work when, for example, the part is hidden by the operator's hand.
[0006] An object of the present disclosure is to provide a work determination device, a work determination method, and a work determination program that accurately determine the completion of work.
Means for Solving the Problems
[0007] The work determination device according to the present disclosure includes an acquisition unit that acquires a first image and a second image that capture a work, a detection unit that detects a first object included in the first image and a second object included in the second image, and a determination unit that inputs determination displacement data including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model and outputs a determination result indicating whether the work is completed. The determination model is trained with training data including, as positive examples, a dataset of training displacement data including the position of the object at the start of the work and the position of the object at the end of the work, and a determination result indicating the completion of the work.
[0008] The work determination method according to the present disclosure includes acquiring a first image and a second image that capture a work, detecting a first object included in the first image and a second object included in the second image, and inputting determination displacement data including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model and outputting a determination result indicating whether the work is completed. The determination model is trained with training data including, as positive examples, a dataset of training displacement data including the position of the object at the start of the work and the position of the object at the end of the work, and a determination result indicating the completion of the work.
[0009] The work determination program according to the present disclosure includes acquiring a first image and a second image that capture a work, detecting a first object included in the first image and a second object included in the second image, and inputting determination displacement data including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model and outputting a determination result indicating whether the work is completed. The determination model is trained with training data including, as positive examples, a dataset of training displacement data including the position of the object at the start of the work and the position of the object at the end of the work, and a determination result indicating the completion of the work.
Advantages of the Invention
[0010] According to the present disclosure, it is possible to accurately determine the completion of work.
Brief Description of Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Modes for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0013] <First Embodiment> [Overview] As shown in FIG. 1, the work determination device 1 acquires a first image F1 and a second image F2 that capture the work. Here, it is assumed that the first image F1 and the second image F2 are acquired by capturing the work in the first process among a plurality of processes. When the work determination device 1 acquires the first image F1 and the second image F2, it detects a first object P1 included in the first image F1 and a second object P2 included in the second image F2. Note that the work determination device 1 has a determination model M1 trained with training data that includes, as positive examples, a dataset of training displacement data including the position of the object at the start of the work and the position of the object at the end of the work, and a determination result indicating the completion of the work. The work determination device 1 inputs determination displacement data including the position of the first object P1 detected in the first image F1 and the position of the second object P2 detected in the second image into the determination model, and outputs a determination result indicating whether the work has been completed normally.
[0014] [Work support device] Next, the configuration of a work support device including the work determination device 1 according to the present disclosure will be described in detail. As shown in FIG. 2, the work support device 11 includes a photographing unit 12, a work determination device 1, and a notification unit 13.
[0015] The photographing unit 12 is arranged in the work environment and captures the work of the worker, and is composed of, for example, a camera. The photographing unit 12 may be attached to the worker himself / herself (head, clothes, etc.) so as to capture, for example, the vicinity of the worker's hands.
[0016] The work determination device 1 determines whether the work has been completed normally based on the image captured by the photographing unit 12. The work determination device 1 may be built into the photographing unit 12, for example. Also, the work determination device 1 may be connected to the photographing unit 12 via a network. At this time, the work determination device 1 may be arranged in the work environment or may be arranged in an external server or the like.
[0017] The notification unit 13 is arranged in the working environment and notifies the operator of the determination result of the work determination device 1. For example, when it is determined by the work determination device 1 that the work is not completed, the notification unit 13 may notify the operator that the work is not completed. The notification unit 13 may be composed of, for example, a speaker or a display. Also, the notification unit 13 may be built into the photographing unit 12.
[0018] [Work determination device] Next, the configuration of the work determination device 1 will be described in detail. As shown in FIG. 3, the work determination device 1 includes an acquisition unit 2, a detection unit 3, and a determination unit 4.
[0019] The acquisition unit 2 acquires the first image F1 and the second image F2 obtained by photographing the work from the photographing unit 12. At this time, the acquisition unit 2 may acquire the first image F1 taken at the start of the work and the second image F2 taken at the end of the work. Also, the acquisition unit 2 may acquire a video of the entire work, that is, a plurality of frame images.
[0020] The detection unit 3 detects the position of the first object P1 included in the first image F1 and the position of the second object P2 included in the second image F2. For example, as shown in FIG. 4, the detection unit 3 may input the first image F1 acquired by the acquisition unit 2 into an object detection model and output detection information 10 of the first object P1 detected from the first image F1. Similarly, the detection unit 3 may input the second image F2 acquired by the acquisition unit 2 into an object detection model and output detection information 10 of the second object P2 detected from the second image F2. Here, it is assumed that the first image F1 includes two first objects P1a and P1b, and the second image F2 includes one second object P2. Note that the number of objects included in the image changes according to the detection and is not limited to a specific number.
[0021] The detection information 10 may include, for example, the positions L1a, L1b, L2 of an object, the type T of the object, the detection reliability C of the object, or the shooting time. Note that the object may include, for example, an article (e.g., a working object, etc.), a person, or the like. Also, the positions L1a, L1b, L2 of the object are calculated based on the detection position of the object, and may indicate, for example, the position of the object (e.g., xy coordinates) with respect to a predetermined reference position. At this time, the reference position may be set at a mark provided on a workbench or at a corner of the workbench. Also, when the position of the image does not change, such as when photographed by a fixed camera, the positions L1a, L1b, L2 of the object may indicate the position of the object in the image (e.g., the coordinate position of the image). The type T of the object may be output based on, for example, the feature amounts of a pre-learned object or person (operator). For example, the type T of the object may be output based on the feature amounts of a working object used in the work (e.g., the working object at the start of work, the working object at the end of work) or a person. At this time, the type T of the object may include identification information (e.g., name) indicating the type of the working object, identification information indicating a person, or the like. The detection reliability C may be calculated based on, for example, the probability that the detected object is of a specific type. For example, the detection reliability C may include the reliability of being a working object or a person used in the work (e.g., the reliability of being the working object at the start of work, the reliability of being the working object at the end of work, the reliability of being a person). The shooting time may be calculated based on, for example, the frame position of the captured video. For example, the shooting time may indicate the shooting time with respect to the start time of work, the shooting time with respect to the end time of work, or the like.
[0022] The object detection model may be pre-trained to output the detection information 10 of the target object when an image including the target object (e.g., the working object at the start of work, the working object at the end of work, a person, etc.) is input. For example, the object detection model may be trained with training data including a dataset of an image including the target object and the detection information 10. The object detection model may be composed of, for example, an object detection algorithm such as YOLO.
[0023] Specifically, when the first image F1 and the second image F2 are input, the object detection model sets a bounding box B for the first image F1 and the second image F2. Then, the object detection model may estimate the reliability that the set bounding box B surrounds the first objects P1a and P1b and the second object P2 respectively. Further, the object detection model divides the first image F1 and the second image F2 into a plurality of regions respectively, and calculates the probability that the first objects P1a and P1b and the second object P2 existing in the regions are of a specific type T (for example, a working object at the start of work, a working object at the end of work, a person, etc.). Subsequently, the object detection model calculates the detection reliability C of the first objects P1a and P1b and the second object P2 based on the reliability of the bounding box B and the probabilities that the first objects P1a and P1b and the second object P2 are of the specific type T. Then, the object detection model may detect the first objects P1a and P1b from the first image F1 and the second object P2 from the second image F2 based on the detection reliability C of the objects included in the first image F1 and the second image F2. Also, the object detection model may detect a preset reference position from the first image F1 and the second image F2, and calculate the positions L1a and L1b of the first objects P1a and P1b and the position L2 of the second object P2 based on the reference position.
[0024] At this time, the detection unit 3 may perform object detection processing on the first image F1 and the second image F2, and show the positions L1a and L1b based on the endpoints of the first objects P1a and P1b and show the position L2 based on the endpoints of the second object P2, for example. Also, the detection unit 3 may further detect the joint points (such as the skeleton, hands, fingers, etc.), that is, the key points of the worker W from the first image F1 or the second image F2. When the joint points of the worker W are detected, the detection unit 3 may detect the first object P1 or the second object P2 from the part other than the detection positions of the joint points of the worker W.
[0025] The determination unit 4 inputs displacement data for determination including the positions L1a and L1b of the first objects P1a and P1b and the position L2 of the second object P2 into the determination model M1, and outputs a determination result indicating whether the work is completed. Here, the determination model M1 is trained with training data including, as positive examples, a dataset of training displacement data including the positions of the objects at the start of the work and the positions of the objects at the end of the work, and determination results indicating the completion of the work. The determination model M1 may be realized by, for example, a neural network.
[0026] For example, the training displacement data may include typical displacement patterns between the position coordinates of the object at the start of the work and the position coordinates of the object at the end of the work. Note that data augmentation such as noise addition, rotation, translation, and affine transformation may be performed on the positions of the objects for the training displacement data.
[0027] Here, the training displacement data may be composed of a first image capturing the object at the start of the work and a second image capturing the object at the end of the work. Also, the training displacement data may be composed of detection information 10 detecting the object at the start of the work included in the first image and detection information 10 detecting the object at the end of the work included in the second image. For example, the training displacement data may be composed of data including the coordinates of the ends of the object. At this time, the training displacement data will have a dimension obtained by multiplying the total number N of objects, the number K of end points of each object, and the number C of types of detection information 10 (N×K×C).
[0028] The determination model M1 may learn displacement feature amounts (e.g., feature vectors) of an object at the start of work and an object at the end of work by inputting training displacement data into a point cloud processing model such as PointNet, for example. At this time, the determination model M1 may be learned so that feature vectors of substantially the same distribution can be obtained regardless of differences in the training displacement data. Further, the determination model M1 may learn displacement feature amounts by reconstructing the input training displacement data. For example, the determination model M1 may learn displacement feature amounts by an auto encoder structure that restores the input training displacement data after compressing it.
[0029] Further, the determination model M1 may be trained by training data that further includes, as negative examples, a data set of training displacement data including mis-positions other than the position of the object at the start of work and mis-positions other than the position of the object at the end of work, and a determination result indicating non-completion of the work.
[0030] Further, the determination model M1 may be trained by training data that further includes, as negative examples, a data set of training displacement data including the position of the object at the start of work and the position of the object during work, or the position of the object during work and the position of the object at the end of work, and a determination result indicating non-completion of the work.
[0031] Further, the determination displacement data may be composed of, for example, one piece of data pre-processed to include both the position of the object at the start of work and the position of the object at the end of work. Further, the determination displacement data may be composed of a plurality of pieces of data each including the position of the object at the start of work and the position of the object at the start of work.
[0032] Next, the hardware configuration of the work determination device 1 will be described in detail. For example, as shown in FIG. 5, the work determination device 1 includes a storage device 5, a processor 6, a user interface (UI) device 7, and a communication device 8 that are interconnected via a bus B.
[0033] Note that the program or instruction for realizing various functions and processes of the work determination device 1 may be downloaded from any external device via a network or the like. Further, the program or instruction for realizing various functions and processes of the work determination device 1 may be provided from a removable storage medium such as a CD-ROM (Compact Disk-Read Only Memory) or a flash memory.
[0034] The storage device 5 may be realized by one or more non-transitory storage media such as a random access memory, a flash memory, or a hard disk drive. And the storage device 5 may store files, data, etc. used for the execution of the program or instruction together with the installed program or instruction.
[0035] The processor 6 may be realized by one or more CPUs (Central Processing Unit), GPUs (Graphics Processing Unit), or processing circuitry, etc. that may be composed of one or more processor cores. The processor 6 executes various functions and processes of the work determination device 1 according to data such as programs, instructions, programs, or parameters (for example, parameters necessary for the execution of the instruction) stored in the storage device 5.
[0036] The user interface device 7 may be composed of an input device such as a keyboard, a mouse, a camera, or a microphone, an output device such as a display, a speaker, a headset, or a printer, or an input / output device such as a smartphone, a tablet, or a touch panel.
[0037] The communication device 8 is realized by various communication circuits that execute wired and / or wireless communication processing with a communication network such as an external device, the Internet, a LAN (Local Area Network), or a cellular network.
[0038] Note that the above-described hardware configuration is merely an example, and the work determination device 1 according to the present disclosure may be implemented by any other appropriate hardware configuration. Further, since the work support device 11 has the same hardware configuration as the work determination device 1, the description thereof is omitted.
[0039] [Work Support Process] Next, with reference to the flowchart shown in FIG. 6, the work support process by the work support device 11 will be described.
[0040] First, in step S1, the imaging unit 12 images the work of the worker. For example, as shown in FIG. 1, the imaging unit 12 images the work of the first step among the work consisting of a plurality of steps. At this time, the worker may output a signal indicating the start or end of the work (for example, a button press signal, voice, etc.) to the work support device 11. When the imaging unit 12 images the work, it transmits the captured image (for example, a moving image) to the work determination device 1. For example, the imaging unit 12 may sequentially transmit the captured image to the work determination device 1 while imaging the work, or may transmit the captured image to the work determination device 1 at a predetermined timing after imaging.
[0041] Subsequently, based on the captured image transmitted from the imaging unit 12, the acquisition unit 2 of the work determination device 1 acquires the first image F1 and the second image F2 in step S2. For example, the acquisition unit 2 may select the first image F1 and the second image F2 from among a plurality of captured images transmitted from the imaging unit 12.
[0042] Specifically, the acquisition unit 2 may select the first image F1 based on the start information indicating the start of the work in the first step, and select the second image F2 based on the end information indicating the end of the work in the first step. For example, the acquisition unit 2 may select the first image F1 and the second image F2 based on the signal from the worker corresponding to the start or end of the work. Further, the acquisition unit 2 may receive a detection signal from a detection unit (not shown) that detects the start or end of the work, and select the first image F1 and the second image F2 based on the detection signal. For example, the detection unit may detect the start or end of the work based on the position of an object, the working time, or the like.
[0043] At this time, the acquisition unit 2 may select a plurality of first images F1 and a plurality of second images F2. For example, the acquisition unit 2 may select a plurality of first images F1 that constitute a video of a predetermined time including the work start time. Further, the acquisition unit 2 may select a plurality of second images F2 that constitute a video of a predetermined time including the work end time.
[0044] When the acquisition unit 2 acquires the first image F1 and the second image F2, it outputs the first image F1 and the second image F2 to the detection unit 3.
[0045] When the detection unit 3 inputs the first image F1, in step S3, it detects the first object P1 included in the first image F1. Further, when the detection unit 3 inputs the second image F2, it detects the second object P2 included in the second image F2. At this time, since the acquisition unit 2 selects the first image F1 at the start of the work and the second image F2 at the end of the work, the detection unit 3 can easily detect the object used at the start of the work and the object used at the end of the work.
[0046] Specifically, the detection unit 3 may input the first image F1 into an object detection model and output detection information 10 (for example, the position of the object, the type of the object, the detection reliability, or the shooting time, etc.) of the first object P1 included in the first image F1. Further, the detection unit 3 may input the second image F2 into the object detection model and output detection information 10 of the second object P2 included in the second image F2.
[0047] Here, consider the case where the detection unit 3 inputs a plurality of first images F1 from the acquisition unit 2. When the detection unit 3 inputs the plurality of first images F1, it may detect the first object P1 from among the plurality of first images F1 based on the detection reliability C of the object detected from the plurality of first images F1. For example, as shown in FIG. 7A, when the detection unit 3 receives the plurality of first images F1 captured from time T0 to time T1 with respect to the work start time from the acquisition unit 2, it inputs each of the plurality of first images F1 into an object detection model. When the object detection model receives the plurality of first images F1, as shown in FIG. 7B, it sets a bounding box B in the first image F1. Then, the object detection model calculates the detection reliability C of the first objects P1a and P1b based on the reliability of the bounding box B and the probability that the first objects P1a and P1b are of a specific type T.
[0048] When the detection reliability C of the first objects P1a and P1b is calculated for each of the plurality of first images F1, the detection unit 3 selects a specific first image F1 from among the plurality of first images F1 based on the detection reliability C, as shown in FIG. 7C. For example, the detection unit 3 may select the first image F1 with the highest detection reliability C. Also, when a plurality of objects P1a and P1b are detected in the plurality of first images F1, the detection unit 3 may select a specific first image F1 with the highest average value of the detection reliability C of the first objects P1a and P1b. Further, in the plurality of first images F1, an image including the first object P1a in the image with the highest detection reliability of the first object P1a and the second object P1b in the image with the highest detection reliability of the second object P1b may be synthesized to obtain a specific first image F1.
[0049] In this way, the detection unit 3 selects a specific first image F1 from among the plurality of first images F1 based on the detection reliability C. Therefore, the detection unit 3 can select the first image F1 including the first objects P1a and P1b, and can suppress the non-detection of the first objects P1a and P1b. Also, since the detection unit 3 narrows down to a specific first image F1, the inference cost of the determination unit 4 can be suppressed.
[0050] When the detection unit 3 selects a specific first image F1, it detects first objects P1a and P1b from the specific first image F1 based on the detection reliability C. Thereby, the detection unit 3 can surely detect the first objects P1a and P1b included in the first image F1.
[0051] At this time, the detection unit 3 may further detect the key points (such as the skeleton, hands, fingers, etc.) of the worker W, that is, the key points, from the first image F1. When the key points of the worker W are detected, the detection unit 3 may detect the first objects P1a and P1b from portions other than the detection positions of the key points of the worker W.
[0052] In addition, the detection unit 3 may detect a plurality of first objects P1a and P1b from different first images F1. For example, when a plurality of first objects P1a and P1b are included in a plurality of first images F1, the detection unit 3 may select a specific first image F1 for detecting each of the first objects P1a and P1b from among the plurality of first images F1 based on the detection reliability C. For example, as shown in FIG. 8, there may be a case where the first objects P1a and P1b are hidden by the body of the worker W and are not detected. When the detection reliability C is lower than a predetermined value, the detection unit 3 may determine that the first objects P1a and P1b are not detected. Also, when the positions of the key points of the worker W overlap the positions of the first objects P1a and P1b, the detection unit 3 may determine that the first objects P1a and P1b are not detected. When there is a first image F1 in which the first object P1a or P1b is not detected, the detection unit 3 may detect the first object P1a from the first image F1a with the highest detection reliability C of the object P1a, and detect the first object P1b from the first image F1b with the highest detection reliability C of the object P1b. Thereby, the detection unit 3 can more surely detect the first objects P1a and P1b.
[0053] Similar to the first image F1, the detection unit 3 may detect the second object P2 from among the plurality of second images F2 based on the detection reliability C of the objects detected from the plurality of second images F2. At this time, the detection unit 3 may select a plurality of second objects P2 from different second images F2. For example, when the plurality of second objects P2 are included in the plurality of second images F2, the detection unit 3 may select a specific second image F2 for detecting each second object P2 from among the plurality of second images F2 based on the detection reliability C. Thereby, the detection unit 3 can surely detect the second object P2.
[0054] When the detection unit 3 detects the first objects P1a and P1b included in the first image F1 and the second object P2 included in the second image F2, the detection unit 3 outputs a detection result including the positions L1a and L1b of the first objects P1a and P1b and the position L2 of the second object P2 to the determination unit 4. The detection result may be composed of, for example, point cloud data or image data.
[0055] When the determination unit 4 inputs the detection result from the detection unit 3, in step S4, the determination unit 4 inputs displacement data for determination including the positions L1a and L1b of the first objects P1a and P1b detected in the first image F1 and the position L2 of the second object P2 detected in the second image F2 to the determination model M1.
[0056] For example, the displacement data for determination may include the detection information 10 for detecting the first objects P1a and P1b from the first image F1 and the detection information 10 for detecting the second object P2 from the second image F2. Thereby, the determination unit 4 can easily indicate the positions L1a and L1b of the first objects P1a and P1b and the position L2 of the second object P2. Note that the displacement data for determination will have a dimension obtained by multiplying the total number N of the first objects P1a and P1b, the second object P2, the number K of end points of each object, and the number C of types of the detection information 10 (N×K×C).
[0057] When the determination displacement data is input, the determination model M1 may extract the displacement feature amounts of the first objects P1a and P1b and the second object P2. Examples of the displacement feature amounts include feature vectors indicating the displacements of the first objects P1a and P1b and the second object P2. For example, the determination model M1 may regard the determination displacement data as point cloud data, and input the determination displacement data into a point cloud processing neural network (such as PointNet, etc.), thereby extracting the displacement feature amounts of the first objects P1a and P1b and the second object P2. Thereby, the determination model M1 can surely select the displacement feature amounts effective for the completion determination of the work.
[0058] Further, the determination model M1 may extract the displacement feature amounts of the first objects P1a and P1b and the second object P2 by performing a max pooling process inside a neural network (such as a point cloud processing neural network). Thereby, the determination model M1 can more surely select the displacement feature amounts effective for the completion determination of the work.
[0059] Subsequently, in step S5, the determination unit 4 outputs a determination result as to whether or not the work has been completed normally. In this way, the determination unit 4 inputs the determination displacement data including the positions L1a and L1b of the first objects P1a and P1b detected in the first image F1 and the position L2 of the second object P2 detected in the second image F2 into the determination model M1. Thereby, the determination unit 4 outputs a determination result as to whether or not the work has been completed normally.
[0060] Since the determination unit 4 makes a determination based on the displacement data of the object at the start and end of the operation, it can accurately determine the completion of the operation regardless of the arrangement position of the imaging unit 12, as compared with the case of making a determination using one image (one position of the object). Also, the position of the object at the start of the operation and the position of the object at the end of the operation are relatively stable within a certain range (with little variation) compared to during the operation. Therefore, the determination unit 4 can accurately determine the completion of the operation. Further, the determination unit 4 can easily train the determination model M1. Thus, the determination unit 4 can robustly determine the completion of the operation with high accuracy regardless of the arrangement position of the imaging unit 12 or the variation in the operation. Also, the determination unit 4 can automatically determine the completion of the operation using the determination model M1.
[0061] At this time, the determination model M1 extracts the displacement feature amounts of the first objects P1a and P1b and the second object P2 by performing a max pooling process on the displacement data for determination. For this reason, the determination unit 4 can extract displacement feature amounts that accurately indicate the displacements of the first objects P1a and P1b and the second object P2.
[0062] Also, the determination model M1 determines whether the operation has been completed normally based on the difference between the displacement feature amounts of the object at the start and end of the operation and the displacement feature amounts of the first objects P1a and P1b and the second object P2. For this reason, the determination unit 4 can more accurately determine the completion of the operation.
[0063] In addition, the determination model M1 may further output non-completion of the work. For example, the determination model M1 may be further trained with training data that includes training displacement data including misaligned positions other than the position of the object at the start of the work and misaligned positions other than the position of the object at the end of the work, and a data set of determination results indicating non-completion of the work as negative examples. Also, the determination model M1 may be further trained with training data that includes training displacement data including the position of the object at the start of the work and the position of the object during the work, or the position of the object during the work and the position of the object at the end of the work, and a data set of determination results indicating non-completion of the work as negative examples. Thereby, the determination unit 4 can accurately determine non-completion of the work.
[0064] In addition, the determination unit 4 may determine whether the work has been completed normally by inputting determination displacement data further including the position of the joint points of the worker W detected in the first image F1 (for example, the positions of the hands, fingers, etc.) and the position of the joint points of the worker W detected in the second image F2 into the determination model M1. At this time, the determination model M1 may be trained with training data that includes a data set of training displacement data further including the position of the worker W at the start of the work and the position of the worker at the end of the work, and a determination result indicating completion of the work as a positive example. Thereby, the determination unit 4 can more accurately determine completion of the work.
[0065] When the determination unit 4 outputs a determination result, it may transmit the determination result to the notification unit 13. For example, when the determination unit 4 does not determine that the work has been completed normally (when it determines that the work is not completed), it may transmit the determination result to the notification unit 13. On the other hand, when the determination unit 4 determines that the work has been completed normally, it may not transmit the determination result to the notification unit 13. Also, the determination unit 4 may transmit all determination results to the notification unit 13.
[0066] When the notification unit 13 receives a determination result, it notifies the worker W of the determination result. For example, when the notification unit 13 receives a determination result indicating non-completion of the work, it may notify the worker W of the determination result. At this time, the notification unit 13 may notify the determination result by sound, text, video, or the like.
[0067] In this way, each time the first image F1 and the second image F2 are acquired by the acquisition unit 2, the work determination device 1 repeats the above detection and determination. As a result, as shown in FIG. 1, the work determination device 1 sequentially determines whether the work in the first process has been completed normally.
[0068] Note that the work determination device 1 may acquire the first image F1 and the second image F2 that capture the work in the second process and the work in the third process, and determine whether the work in the second process and the work in the third process have been completed normally. Further, a plurality of work determination devices 1 corresponding to the work in the first process to the third process may be arranged, and the plurality of work determination devices 1 may each determine whether the work in the first process to the third process has been completed normally. In this way, the work determination device 1 can suppress work errors by accurately determining the completion of each process.
[0069] According to the present embodiment, the determination unit 4 inputs determination displacement data including the positions L1a and L1b of the first objects P1a and P1b detected in the first image F1 and the position L2 of the second object P2 detected in the second image F2 into the determination model M1. Then, the determination unit 4 outputs a determination result indicating whether the work has been completed. Therefore, the determination unit 4 can accurately determine the completion of the work.
[0070] <Second Embodiment> [Overview] The imaging unit 12 images a range including the object on which the worker W works and the worker W. At this time, the detection unit 3 may further detect the joint points of the worker W included in the first image F1 and the joint points of the worker W included in the second image F2.
[0071] [Work Support Device] The imaging unit 12 is arranged in the work environment so as to image a range including the object on which the worker W works and the worker W. For example, the imaging unit 12 may be fixedly installed in the work environment.
[0072] The acquisition unit 2 acquires a first image F1 including the first object P1 and the operator W, and a second image F2 including the second object P2 and the operator W.
[0073] The detection unit 3 detects the positions L1a and L1b of the first objects P1a and P1b included in the first image F1, and the position L2 of the second object P2 included in the second image F2. At this time, the detection unit 3 may further detect the joint points of the operator W included in the first image F1 and the joint points of the operator W included in the second image F2.
[0074] For example, as shown in FIG. 9, the detection unit 3 may input the first image F1 acquired by the acquisition unit 2 into an object detection model, and output the detection information 10 of the first object P1 and the operator W detected from the first image F1. Similarly, the detection unit 3 may input the second image F2 acquired by the acquisition unit 2 into an object detection model, and output the detection information 10 of the second object P2 and the operator W detected from the second image F2.
[0075] Specifically, when the first image F1 and the second image F2 are input into the object detection model, the object detection model sets bounding boxes B for the first objects P1a and P1b and the operator W included in the first image F1, and the second object P2 and the operator W included in the second image F2, respectively. Then, based on the detection reliability C, the object detection model detects the first objects P1a and P1b and the operator W from the first image F1, and detects the second object P2 and the operator W from the second image F2.
[0076] At this time, the detection unit 3 may further detect the position of the joint point K1 of the operator W included in the first image F1 and the position of the joint point K2 of the operator W included in the second image F2, that is, the position of the key point of the operator W. The joint point may indicate, for example, the center position of the joint, the center or the end point of a part such as the eye or the nose.
[0077] When the joint point K1 or K2 of the operator W is detected, the detection unit 3 may detect the first object P1 or the second object P2 from a portion other than the joint point K1 or K2 of the operator W.
[0078] Further, when the position of the joint point K1 or K2 of the operator W overlaps with the first objects P1a and P1b or the second object P2, the detection unit 3 may regard the first objects P1a and P1b or the second object P2 as undetected. Then, the detection unit 3 may select the first image F1 or the second image F2 in which the position of the joint point K1 or K2 of the operator W does not overlap with the first objects P1a and P1b or the second object P2, and detect the first objects P1a and P1b or the second object P2.
[0079] Also, the determination unit 4 inputs displacement data for determination including the positions L1a and L1b of the first objects P1a and P1b detected by the detection unit 3 and the position L2 of the second object P2 into the determination model M1. At this time, the displacement data for determination may include the position of the joint point K1 of the operator W included in the first image F1 and the position of the joint point K2 of the operator W included in the second image F2. Thereby, the determination unit 4 can more accurately determine the completion of the work.
[0080] According to the present embodiment, the detection unit 3 further detects the position of the joint point K1 of the operator W included in the first image F1 and the position of the joint point K2 of the operator W included in the second image F2. Therefore, the determination unit 4 can more accurately determine the completion of the work.
[0081] <Third Embodiment> [Overview] The photographing unit 12 photographs the work of the work robot. For example, the photographing unit 12 may photograph a range including the object on which the work robot works and the work robot. Then, the detection unit 3 may further detect the joint points of the work robot included in the first image F1 and the joint points of the work robot included in the second image F2.
[0082] [Work Support Device] The photographing unit 12 photographs the work of the work robot. For example, the photographing unit 12 is arranged in the work environment so as to photograph a range including the object on which the work robot works and the work robot. For example, the photographing unit 12 may be fixedly installed in the work environment.
[0083] The acquisition unit 2 acquires a first image F1 including the first object P1 and the working robot, and a second image F2 including the second object P2 and the working robot.
[0084] The detection unit 3 detects the positions L1a and L1b of the first objects P1a and P1b included in the first image F1 and the position L2 of the second object P2 included in the second image F2. At this time, the detection unit 3 may further detect the joint points of the working robot included in the first image F1 and the joint points of the working robot included in the second image F2.
[0085] When the joint points of the working robot are detected, the detection unit 3 may detect the first object P1 or the second object P2 from a portion other than the joint points of the working robot.
[0086] Further, when the positions of the joint points of the working robot overlap with the first objects P1a and P1b or the second object P2, the detection unit 3 may regard the first objects P1a and P1b or the second object P2 as undetected. Then, the detection unit 3 may select the first image F1 or the second image F2 in which the positions of the joint points of the working robot do not overlap with the first objects P1a and P1b or the second object P2, and detect the first objects P1a and P1b or the second object P2.
[0087] Further, the determination unit 4 inputs displacement data for determination including the positions L1a and L1b of the first objects P1a and P1b detected by the detection unit 3 and the position L2 of the second object P2 into the determination model M1. At this time, the displacement data for determination may include the positions of the joint points of the working robot included in the first image F1 and the positions of the joint points of the working robot included in the second image F2. Thereby, the determination unit 4 can more accurately determine the completion of the work.
[0088] According to the present embodiment, the detection unit 3 further detects the positions of the joint points of the working robot included in the first image F1 and the positions of the joint points of the working robot included in the second image F2. Therefore, the determination unit 4 can more accurately determine the completion of the work.
[0089] In the above-described Embodiments 1 to 3, the determination unit 4 outputs only the determination result, but it may output other information. For example, the determination unit 4 may further have an explainable model that outputs the reason for the determination by the determination model M1 by inputting the determination result output from the determination model M1. The explainable model may output the reason for the determination result, for example, based on the start state and the end state of the object used in the work. Further, in the case of a determination result indicating non-completion of the work, the explainable model may output a method for guiding to correct work. The explainable model may output, for example, a clue as to which part should be focused on to correct the work by visualizing the object that contributed to the work completion determination among the displacement data for determination input to the determination model M1. At this time, the explainable model may utilize, for example, Class Activation Mapping (CAM). Also, a multi-modal base model capable of inputting images and languages may be used to guide the worker to correct work. Specifically, when the determination model M1 determines non-completion, an image including the second image F2, an image of the work end state at the time of training, and a question sentence such as "which object should be moved and how to move it to make the work completion state" are input to the multi-modal base model. Thereby, the determination model M1 may output from the multi-modal base model a method for guiding to correct work.
[0090] As described above, specific examples of the present disclosure have been described in detail, but these are merely examples and do not limit the scope of the claims. The technology described in the claims includes various modifications and changes of the specific examples illustrated above.
Industrial Applicability
[0091] The work determination device according to the present disclosure can be used for a device that determines the completion of work.
Explanation of Signs
[0092] 1 Work determination device 2 Acquisition unit 3 Detection unit 4 Determination unit 5 Storage device 6 Processor 7 User Interface Device 8 Communication Device 10 Detection Information 11 Work Support Device 12 Imaging Unit 13 Notification Unit B Bounding Box C Detection Reliability F1,F1a~F1c First Image F2 Second Image K1,K2 Joint Point L1a,L1b First Position L2 Second Position M1 Judgment Model P1,P1a,P1b First Object P2 Second Object T Type W Operator
Claims
1. An acquisition unit that acquires a first image and a second image that capture an operation; A detection unit that detects a first object included in the first image and a second object included in the second image; A determination unit that inputs displacement data for determination including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model, and outputs a determination result indicating whether the operation is completed; Comprising: The determination model is a work determination device trained by training data including, as positive examples, a dataset of training displacement data including the position of an object at the start of an operation and the position of the object at the end of the operation, and a determination result indicating the completion of the operation.
2. The acquisition unit according to claim 1, wherein the acquisition unit acquires the first image based on start information indicating the start of an operation, and acquires the second image based on end information indicating the end of the operation.
3. The acquisition unit acquires a plurality of the first images captured at the start of an operation and a plurality of the second images captured at the end of the operation, The detection unit detects the first object from among the plurality of the first images based on the detection reliability of the object detected from the plurality of the first images, and detects the second object from among the plurality of the second images based on the detection reliability of the object detected from the plurality of the second images. The work determination device according to claim 2.
4. When a plurality of the first objects are included in the plurality of the first images, the detection unit selects a specific first image for detecting each of the first objects from among the plurality of the first images based on the detection reliability. When a plurality of the second objects are included in the plurality of the second images, the detection unit selects a specific second image for detecting each of the second objects from among the plurality of the second images based on the detection reliability. The work determination device according to claim 3.
5. The detection unit according to claim 1, wherein the detection unit detects the first object and the second object based on the detection reliability of the objects included in the first image and the second image.
6. The displacement data for determination according to claim 1 includes detection information for detecting the first object from the first image and detection information for detecting the second object from the second image.
7. The determination model regards the displacement data for determination as point cloud data, extracts displacement feature amounts of the first object and the second object by inputting the displacement data for determination into a point cloud processing neural network, and determines whether the work has been completed based on the displacement feature amounts. The work determination device according to claim 1.
8. The determination model in the work determination device according to claim 1 extracts displacement feature amounts of the first object and the second object by performing max pooling processing inside the neural network.
9. The determination model in the work determination device according to claim 1 determines whether the work has been completed based on a difference between displacement feature amounts of an object at the start of work and an object at the end of work and displacement feature amounts of the first object and the second object.
10. The determination model in the work determination device according to claim 1 is further trained by training data that further includes, as negative examples, a dataset of training displacement data including positions other than the position of the object at the start of work and positions other than the position of the object at the end of work and a determination result indicating non-completion of the work.
11. The determination model in the work determination device according to claim 1 is further trained by training data that further includes, as negative examples, a dataset of training displacement data including the position of the object at the start of work and the position of the object during work, or the position of the object during work and the position of the object at the end of work, and a determination result indicating non-completion of the work.
12. In the work determination device according to claim 1, the determination unit outputs the determination result to a notification unit that is arranged in the work environment and notifies the worker of the determination result.
13. The work determination device according to claim 1 further has an explainable model that outputs the reason for the determination by the determination model by inputting the determination result.
14. obtaining a first image and a second image that photograph the work; detecting a first object included in the first image and a second object included in the second image; inputting displacement data for determination including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model, and outputting a determination result indicating whether the work has been completed; including The determination model is trained by training data that includes a dataset of training displacement data including the position of an object at the start of work and the position of the object at the end of work, and a determination result indicating completion of the work as a positive example, and is a work determination method executed by a computer.
15. Obtaining a first image and a second image that capture the work; Detecting a first object included in the first image and a second object included in the second image; Inputting determination displacement data including the position of the first object detected in the first image and the position of the second object detected in the second image into a determination model, and outputting a determination result indicating whether the work has been completed; including; The determination model is a work determination program executed by a computer, which is trained by training data that includes a dataset of training displacement data including the position of an object at the start of work and the position of the object at the end of work, and a determination result indicating completion of the work as a positive example.
Citation Information
Patent Citations
Work determination apparatus and work determination method
JP2022010646A