Annotation device, annotation system, training system, and annotation method

The annotation device and method address the challenges of accurately annotating complex objects by generating learning datasets from actual end effector operations, reducing work costs and improving annotation accuracy.

WO2025115163A1PCT designated stage expired Publication Date: 2025-06-05KAWASAKI JUKOGYO KK
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/042876
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-30
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing annotation technologies face challenges in accurately discriminating objects from images, especially when objects have complex shapes or are made of materials that are difficult to detect, leading to increased work costs and potential failures in robotic operations.

Method used

An annotation device and method that acquire images of objects and corresponding operation information from end effectors, determining if a predetermined operation has been executed, and generating a learning dataset that includes the image and information about the end effector's position and posture during the operation.

Benefits of technology

This approach reduces annotation work costs and improves accuracy by using actual operation data, enabling the generation of effective learning models for robotic tasks regardless of object material or shape complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023042876_05062025_PF_FP_ABST
    Figure JP2023042876_05062025_PF_FP_ABST
Patent Text Reader

Abstract

This annotation device comprises: a first control unit that generates a dataset for training; and a storage unit that stores the dataset. The first control unit acquires a first image including an object captured by an imaging device, acquires action information related to the action of an end effector on the object so as to determine whether or not the end effector has executed a predetermined first action on the object, and if the end effector has executed the first action, generates, as a dataset, the first image and first information of the end effector including a position corresponding to the first action, and causes the storage unit to store the generated dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Annotation device, annotation system, learning system, and annotation method

[0001] The present disclosure relates to an annotation device, an annotation system, a learning system, and an annotation method.

[0002] Devices that generate training data for use in machine learning are known. For example, Japanese Patent Application Laid-Open No. 2021-71878 describes an annotation device that includes an imaging device, a robot, a control unit, a designation unit, a coordinate processing unit, and a storage unit. The control unit controls the robot to acquire learning images of multiple objects with different positional relationships between the imaging devices. The position of the object is designated by a user using an input unit as the designation unit. Alternatively, the position of the object is designated using an image processing algorithm using an image processing unit as the designation unit. The storage unit converts the position of the object in the robot coordinate system to the position of the object in the image coordinate system at the time of image capture, and stores the converted position together with the learning image.

[0003] The technology described in JP 2021-71878 A can automatically obtain the position of an object in the image coordinate system of a captured image by using the known position of the object in the robot coordinate system. However, depending on the shape and material of the object, it may be difficult to identify the object from the image. In addition, if the user specifies the position, there is a risk of increased operational costs. Furthermore, even if the position of the object can be obtained from the image, if the object has a complex shape, it may be difficult to output the position and posture of the robot to perform a specified process on the object.

[0004] The present disclosure has been made to solve at least part of the above-mentioned problems, and can be realized, for example, in the following forms.

[0005] According to a first aspect of the present disclosure, there is provided an annotation device. The annotation device includes a first control unit that generates a learning dataset and a storage unit that stores the dataset. The first control unit acquires a first image including an object captured by an imaging device. The first control unit acquires operation information regarding an operation of an end effector with respect to the object and determines whether the end effector has performed a predetermined first operation with respect to the object. When the end effector has performed the first operation, the first control unit generates the first image and first information about the end effector as the dataset. The first information includes a position corresponding to the first operation of the end effector. The control unit stores the generated dataset in the storage unit.

[0006] According to a second aspect of the present disclosure, an annotation system is provided. The annotation system includes an imaging device, a detection unit, and an annotation device. The imaging device captures a first image including an object. The detection unit detects motion information of an end effector that performs a predetermined motion on the object. The annotation device includes a first control unit that generates a learning dataset, and a storage unit. The first control unit acquires the motion information and determines whether the end effector has performed a predetermined first motion on the object. When the end effector has performed the first motion, the first control unit generates the first image and first information as the dataset. The first information includes a position corresponding to the first motion. The first control unit stores the generated dataset in the storage unit.

[0007] According to a third aspect of the present disclosure, there is provided an annotation method including: acquiring a first image including an object captured by an imaging device; acquiring operation information related to an operation of an end effector on the object; determining whether the end effector has performed a predetermined first operation on the object; and, if the end effector has performed the first operation, generating the first image and first information about the end effector including a position corresponding to the first operation as a dataset for supervised learning.

[0008] 1 is a schematic diagram of an annotation system according to a first embodiment; FIG. 2 is a block diagram of the annotation system according to the first embodiment; FIG. 3 is a flowchart showing annotation processing; FIG. 4 is a flowchart showing learning processing; FIG. 5 is a diagram showing an example of an input operation; FIG. 6 is a diagram showing an interlocking tool attached to the fingers of a user; FIG. 7 is a diagram showing a state in which a workpiece is grasped by two fingers attached with an interlocking tool; FIG. 8 is a diagram showing a pen-shaped interlocking tool;

[0009] <First embodiment> <Configuration of annotation system> Fig. 1 is a diagram showing an example of an annotation system 10 in the first embodiment. Fig. 2 is a schematic block diagram of the annotation system 10. The annotation system 10 mainly includes an imaging device 30, a detection unit 70, and an annotation device 100. The annotation system 10 of this embodiment further includes a robot control device 25, an operation unit 29, and an input device 40. The annotation system 10 is configured to generate a data set of an image including an object and first information including the position of the end effector when the end effector performs a predetermined first action on the object.

[0010] In this embodiment, the end effectors E1 and E2 applied to the annotation system 10 are gripping tools configured to grip a workpiece W as an object. The first operation is an operation (gripping operation) in which the end effectors E1 and E2 grip the workpiece W. Note that the "operation" can also be referred to as a process or an action. In this embodiment, the first information includes the position and orientation of the end effectors E1 and E2 when they grip the workpiece W. The first information in this embodiment is three-dimensional information. The position is position coordinate information in a coordinate system with a predetermined reference point as the origin. The orientation is information on the orientation (angle) of the end effectors E1 and E2.

[0011] The end effector E1 is a robot hand detachably attached to the robot arm 20. The robot arm 20 is a vertical articulated robot with six degrees of freedom. Two adjacent joints J, J are connected by a link L. The joint J includes known components such as a servo motor, an encoder, and a reducer. The joint J operates under the control of the robot control device 25.

[0012] The robot control device 25 is configured as a computer including a CPU and a memory. The CPU of the robot control device 25 functions as a robot control unit 26 by expanding and executing a program stored in the memory. The robot control unit 26 controls the operations of the robot arm 20 and the end effector E1.

[0013] The operation unit 29 is connected to the robot control device 25. The operation unit 29 accepts user operations related to the operation of the end effector E1. The operation unit 29 may be a so-called teaching pendant. By operating the operation unit 29 by the user, the end effector E1 can grip and move the workpiece W, and can also release the gripped workpiece W. In other words, the end effector E1 can perform so-called pick-and-place operations.

[0014] The end effector E2 is a linked tool that is supported by the fingers of the user u and moves in conjunction with the movement of the fingers of the user u. The end effector E2 is a generally V-shaped tool that has a base E21 on the side of the user u's fingers and two tip ends E22 on the opposite side of the base E21. The two tip ends E22 of the end effector E2 move toward or away from each other in response to the movement of the user's fingers. This allows the end effector E2 to grip and move the workpiece W, and also to release the gripped workpiece W.

[0015] A first sensor 21 and a second sensor 22 are provided at the tip of the end effector E1 and the tip E22 of the end effector E2, respectively. The first sensor 21 and the second sensor 22 function as a detection unit 70 that detects operation information related to the operation of the end effectors E1 and E2. In this embodiment, the first sensor 21 and the second sensor 22 are pressure sensors. The detection value of the first sensor 21 changes depending on the degree of contact of the end effector E1 with the workpiece W. Therefore, the gripping operation of the end effector E1 can be estimated from the detection result of the first sensor 21. Similarly, the detection value of the second sensor 22 changes depending on the degree of contact of the end effector E2 with the workpiece W. Therefore, the gripping operation of the end effector E2 can be estimated from the detection result of the second sensor 22.

[0016] The imaging device 30 includes a camera having an optical system and a two-dimensional imaging element. The imaging device 30 is attached to a wall, ceiling, or the like around the robot arm 20. The imaging device 30 captures images of the inside of the container C at predetermined intervals to obtain images before the end effectors E1, E2 grasp the workpiece W, images when they grasp the workpiece W, images after they grasp the workpiece W, and the like. The image before the end effectors E1, E2 grasp the workpiece W is also referred to as a "first image." Note that in this embodiment, the first image includes the workpiece W but does not include the end effectors E1, E2.

[0017] The input device 40 is configured to be able to input various instructions and information to the annotation device 100. The instructions include, for example, an instruction to start annotation processing or learning processing, which will be described later. As the input device 40, for example, various input terminals such as a touch panel or a push button device can be applied.

[0018] 2, the annotation device 100 includes a CPU (Central Processing Unit) 110, which is a processor, a memory 120, and an interface circuit 130. The annotation device 100 is communicably connected to peripheral devices including a robot control device 25, an imaging device 30, and an input device 40 via the interface circuit 130. These communications can be performed by wireless communication or wired communication using a known communication method.

[0019] The memory 120 includes a volatile memory and a non-volatile memory. An annotation program P1 is stored in the memory 120. The CPU 110 functions as a first control unit 111 by expanding and executing the annotation program P1 stored in the memory 120. The first control unit 111 executes the annotation process.

[0020] 1 and 2, the annotation device 100 is connected to a learning device 200. The learning device 200 includes a CPU 210, which is a processor, a memory 220, and an interface circuit 230. The learning device 200 is communicatively connected to the annotation device 100 and the robot control device 25 via the interface circuit 230.

[0021] The memory 220 includes a volatile memory and a non-volatile memory. A learning program P2 is stored in the memory 220. The CPU 210 functions as a second control unit 211 by expanding and executing the learning program P2 stored in the memory 220. The second control unit 211 executes the learning process using a dataset D collected by the annotation process. The annotation device 100 and the learning device 200 constitute a learning system 300.

[0022] <Annotation Processing> Fig. 3 is a diagram showing a flowchart of the annotation processing executed by the first control unit 111. In this embodiment, the annotation processing is started by inputting a start instruction to the annotation device 100 via the input device 40. As described above, the annotation system 10 can be applied with the end effector E1, which is a robot hand, and the end effector E2, which is an interlocking tool. First, the annotation processing when the robot hand (end effector E1) is applied as the end effector that grips the workpiece W will be described.

[0023] <Annotation Process Using Robot Hand> When the annotation process is started, in step S10, the first control unit 111 acquires a first image from the imaging device 30. In this embodiment, when the annotation process is started, the imaging device 30 captures an image of the workpiece W in the container C from above in response to a request from the first control unit 111, and transmits the image as a first image to the annotation device 100. In this embodiment, the first image is a color image. The first image may be a distance image, or may be a color image and a distance image.

[0024] In step S20, the first control unit 111 determines whether or not operation information of the end effector E1 has been acquired via the detection unit 70. The operation information is information regarding the gripping operation of the end effector E1 with respect to the workpiece W. With respect to the end effector E1, the operation information is the detection value of the first sensor 21. In this embodiment, when the annotation process is started, the first control unit 111 requests the robot control device 25 to transmit the operation information. When the end effector E1 is activated (on), the robot control device 25 repeatedly transmits the detection value of the first sensor 21 to the annotation device 100 as operation information. If the operation information has been acquired, the first control unit 111 makes a positive determination in step S20 and proceeds with the annotation process to step S30. If the operation information has not been acquired, the first control unit 111 makes a negative determination in step S20 and ends the annotation process.

[0025] In step S30, the first control unit 111 uses the operation information to determine whether the end effector E1 has performed the first operation. The first control unit 111 determines that the first operation has been performed if the detection value of the first sensor 21 satisfies a predetermined determination condition. In this embodiment, the first control unit 111 determines that the first operation has been performed if the detection value of the first sensor 21 is equal to or greater than a first threshold value stored in the memory 120. The first threshold value is determined based on the lower limit of the detection value of the first sensor 21 when the end effector E1 grasps the workpiece W. The first threshold value is determined in advance through experiments or simulations. Step S30 is also a step in which the first control unit 111 determines whether an "event" called the first operation has occurred. If the detection value of the first sensor 21 is equal to or greater than the first threshold value, the first control unit 111 makes a positive determination in step S30 and proceeds to step S40.

[0026] On the other hand, if the detection value of the first sensor 21 is less than the first threshold, the first control unit 111 makes a negative determination in step S30 and ends this processing routine. Note that if the first control unit 111 makes a negative determination in step S30, it may end this processing routine and then repeatedly execute the annotation processing. Alternatively, the first control unit 111 may return the processing to step S20 and repeat the acquisition of operation information in step S20 and the determination in step S30 until a positive determination is made in step S30.

[0027] In step S40, the first control unit 111 acquires first information corresponding to the first operation. The first information is information including the position of the end effector E1 when it performed the first operation. In this embodiment, the first control unit 111 acquires the position and posture when the first operation was performed via the robot control device 25. "When the first operation was performed" can also be rephrased as the timing when the first operation was performed or the time when the first operation was performed. In this embodiment, "when the first operation was performed" means "when the end effector E1 grasps the workpiece W" and "before the end effector E1 moves the workpiece W."

[0028] In this embodiment, the first information includes the gripping position and gripping orientation when the end effector E1 grips the workpiece W, that is, when the detection value of the first sensor 21 becomes equal to or greater than the first threshold. The first information may also include the opening width of the end effector E1. Note that the first control unit 111 may acquire the detection value of the first sensor 21 as well as the position and orientation of the end effector E1 when the detection value is detected, and sequentially store them in the memory 120. The first control unit 111 may acquire, as the first information, the position and orientation of the end effector E1 corresponding to the detection value of the first sensor 21 when a positive determination is made in step S30, by referring to the memory 120.

[0029] Upon acquiring the first information, the first control unit 111 proceeds to step S45 of the annotation process and determines whether the first operation was successful. Furthermore, the first control unit 111 generates second information indicating whether the first operation was successful. A state in which the first operation was successful means that the end effector E1 performed the first operation on the workpiece W, resulting in the workpiece W reaching a predetermined state. A state in which the first operation was unsuccessful means that the end effector E1 performed the first operation on the workpiece W, but the workpiece W did not reach the predetermined state. In this embodiment, the predetermined state is a state in which the workpiece W is gripped and moved by the end effector E1. The first control unit 111 determines that the first operation was successful if the end effector E1 was able to grip the workpiece W and then move it. The first control unit 111 determines that the first operation has failed if the position or posture of the workpiece W relative to the end effector E1 shifts or the workpiece W falls during the process of moving the workpiece W after the end effector E1 has gripped the workpiece W. The first control unit 111 can determine whether the end effector E1 continues to grip the workpiece W using the detection result of the first sensor 21. The first control unit 111 may also determine that the first operation has been successful if the end effector E1 has gripped the workpiece W and moved it a predetermined distance or more. In this case, the first control unit 111 may acquire the amount of movement of the end effector E1 via the robot control unit 26 or the imaging device 30.

[0030] In step S50, the first control unit 111 generates a data set D for learning the first image and the first information. Generating the data set D also means annotating the first image with the gripping position and gripping posture of the end effector E1. In this embodiment, the first control unit 111 includes the second information generated in step S40 in the data set D. The first control unit 111 stores the generated data set D in the memory 120. In this manner, the annotation process by the first control unit 111 ends.

[0031] <Annotation Processing Using an Interlocking Tool> Next, an annotation processing when an interlocking tool (end effector E2) is used as the end effector that grips the workpiece W will be described. The annotation processing using the end effector E2 differs from the annotation processing using the end effector E1 mainly in the manner in which operation information is acquired ( FIG. 3 , step S20) and the manner in which first information is acquired ( FIG. 3 , step S40). Specifically, after acquiring the first image in step S10, the first control unit 111 determines in step S20 whether or not a detection value of the second sensor 22 provided on the end effector E2 has been acquired. The second sensor 22 may be configured to repeatedly transmit detection values ​​to the annotation device 100, for example, when the user turns on the second sensor 22.

[0032] When the detection value of the second sensor 22 is acquired (step S20, YES), the first control unit 111 determines in step S30 whether the detection value of the second sensor 22 is equal to or greater than a second threshold value stored in the memory 120. The second threshold value is determined based on the lower limit of the detection value of the second sensor 22 when the end effector E2 grips the workpiece W. The second threshold value is determined in advance through experiments or simulations. When the detection value of the second sensor 22 is equal to or greater than the second threshold value, the first control unit 111 makes a positive determination in step S30 and proceeds to step S40. On the other hand, when the detection value of the second sensor 22 is less than the second threshold value, the first control unit 111 makes a negative determination in step S30 and ends this processing routine.

[0033] In step S40, the first control unit 111 acquires the position and orientation of the end effector E2 when the first operation is performed via the imaging device 30. In response to a request from the first control unit 111, the imaging device 30 repeatedly captures images from above the container C and transmits them to the first control unit 111. The first control unit 111 stores the captured images and the detection values ​​of the second sensor 22 in the memory 120. The first control unit 111 can acquire the gripping position and gripping orientation of the end effector E2 as first information by analyzing the images corresponding to the detection results of the second sensor 22 when the first operation is performed. In another embodiment, the imaging device 30 may be equipped with an image processing device, and may analyze the images to acquire the position and orientation of the end effector E2 and transmit the analyzed images to the annotation device 100. Other steps of the annotation process using the end effector E2 as an interlocking tool are similar to those of the annotation process using the end effector E1 as a robot hand, and therefore will not be described here.

[0034] By performing the annotation process described above on a plurality of works W, a plurality of data sets D are collected in the memory 120. Next, the learning process using the data sets D will be described.

[0035] 4 is a flowchart showing the learning process executed by the second control unit 211 of the learning device 200. The second control unit 211 executes the learning process when the learning device 200 is instructed to start the learning process. The start instruction may be input to the learning device 200 directly or indirectly from the input device 40. Alternatively, the start instruction may be input to the learning device 200 from the annotation device 100 when the annotation process is completed or when the number of collected datasets D reaches a predetermined number.

[0036] In step S100, the second control unit 211 acquires a dataset D from the annotation device 100 and selects a dataset D suitable for the learning process. In the present embodiment, the second control unit 211 selects, from the plurality of datasets D, a dataset D in which the second information indicates the success of the first operation. The second control unit 211 may further be configured to determine whether the first image is blurred using a predetermined determination criterion. The second control unit 211 may exclude, from the plurality of datasets D, a dataset D that includes a blurred first image.

[0037] In step S200, the second control unit 211 performs supervised learning using the selected dataset D to generate a learning model and stores it in the memory 220. If a learning model is already stored in the memory 220, the second control unit 211 updates the learning model. By using the learning model generated in this manner, it becomes possible to output, in response to an input of a first image, a gripping position and a gripping posture for the end effectors E1 and E2 to perform a gripping operation on the workpiece W included in the first image. Note that in step S100, the second control unit 211 may select all datasets D, including not only datasets D in which the second information indicates a success of the first operation, but also datasets D in which the second information indicates a failure of the first operation. The second control unit 211 may perform supervised learning in step S200 using datasets D in which the first image, the first information, and the success of the first operation are used, as well as datasets D in which the first image, the first information, and the failure of the first operation are used.

[0038] The effects of the above-described embodiment will be described.

[0039] One method for automating the pick-and-place operation of the workpiece W involves using a learning model to acquire the position of the workpiece W from a captured image and then using the acquired position of the workpiece W to cause the end effector E1 attached to the robot arm 20 to perform a gripping operation. This conventional learning model can be generated, for example, using a data set in which the captured image is assigned position information of the workpiece W. However, annotation of the position information of the workpiece W in the captured image is performed, for example, by displaying the captured image on a monitor and having the user input the workpiece position on the monitor, or by performing predetermined image processing on the captured image. This raises concerns about increased annotation costs and reduced annotation accuracy due to the user performing the annotation. In particular, reduced annotation accuracy can lead to failure of the end effector E1 to grip the workpiece W.

[0040] The container C shown in FIG. 1 contains multiple workpieces W. Workpiece W1 is a cable-like workpiece. Workpiece W2 is a workpiece having a transparent portion tr that transmits light. Workpiece W3 is a polygonal workpiece. Because the transparent portion tr is difficult to see in the captured image, assigning position information of workpiece W2 to the captured image can be difficult or inaccurate. This increases work costs and reduces annotation accuracy. Similar concerns exist not only for workpieces with transparent portions tr, but also for workpieces with translucent portions or workpieces made of reflective materials. Furthermore, even if a conventional learning model outputs workpiece position information in response to an image input, if the gripping position by the end effector is inappropriate, the gripping operation may fail. Grip operation failures may occur, particularly for workpieces with complex shapes such as workpieces W1 and W3.

[0041] In contrast, in the present embodiment, the first control unit 111 generates a learning dataset D from a first image including the workpiece W and first information including the positions and orientations of the end effectors E1 and E2 when the end effectors E1 and E2 actually grip the workpiece W. In other words, the first control unit 111 annotates the first image with the actual gripping positions and orientations of the end effectors E1 and E2, rather than the estimated gripping positions and orientations. Therefore, the learning dataset D can be generated and collected regardless of the material or shape of the workpiece W. Furthermore, because the first control unit 111 annotates the first image with the first information, the work cost for annotation is reduced compared to an aspect in which a user performs annotation.

[0042] Furthermore, the first control unit 111 determines whether the end effectors E1 and E2 have performed the first operation using the detection results of the first sensor 21 and the second sensor 22 serving as the detection unit 70. Therefore, according to this embodiment, it is possible to determine (detect) whether the first operation has been performed without interrupting the operation of the end effectors E1 and E2. This reduces the time required to collect the data set D.

[0043] The second control unit 211 also performs supervised learning using the collected data set D to generate a learning model. When a first image is input, the generated learning model can output first information including the positions at which the end effectors E1 and E2 perform the first operation on the workpiece W. The first information used for learning is not an estimated value, but includes the actual positions and orientations of the end effectors E1 and E2 when they performed the first operation on the workpiece W. Therefore, even if the workpiece W included in the first image includes a transparent portion tr or has a complex shape, a learning model can be generated that can output the positions and orientations of the end effectors E1 and E2 for performing the first operation in response to the input of the first image. Furthermore, when the first operation is performed using the positions and orientations output from the learning model, the success rate of the first operation can be increased.

[0044] Furthermore, the data set D includes second information regarding the success or failure of the first action, and the second control unit 211 selects, from the collected data sets D, a data set D in which the second information indicates the success of the first action, and performs learning. This improves the accuracy of the learning model. Therefore, when the first action is performed using the position and posture output from the generated learning model, the success rate of the first action can be further increased. Note that when supervised learning is performed using not only the data set D in which the second information indicates the success of the first action, but also the data set D in which the second information indicates the failure of the first action, the accuracy of the learning model is further improved.

[0045] In this embodiment, the first control unit 111 of the annotation device 100 can convert the movement of the end effector E2, which serves as an interlocking tool supported by the fingers of the user u, into the movement of the robot hand including the end effector E1. Therefore, collecting a data set D of the first image and first information including the position and posture of the end effector E2 when the end effector E2 performs the first movement, and performing a learning process, contributes to automating the gripping operation of the workpiece W. One mode for converting the movement of the end effector E2 into the movement of the robot hand is to previously store a correspondence between the position and posture of the end effector E2 and the position and posture of the robot hand in the memory 120, and then use the correspondence to convert the gripping position and gripping posture of the end effector E2 into the gripping position and gripping posture of the robot hand. This mode allows the input of a first image including the workpiece W to output the gripping position and gripping posture of the workpiece W for the robot hand to perform the gripping operation.

[0046] In step S30 (Figure 3) of the above annotation process, the first control unit 111 acquires the detection values ​​of the first and second sensors 21 and 22 as operation information and determines whether the first operation has been performed, but the first control unit 111 may also acquire a predetermined input operation as operation information.

[0047] FIG. 5 illustrates an example of a predetermined input action. The end effector E2a illustrated in FIG. 5 is a substantially V-shaped interlocking tool supported by one hand of the user u. The end effector E2a does not include a sensor. FIG. 5 illustrates a state in which the user u is making a hand sign s with the other hand toward the imaging device 30 as an input action. The predetermined input action may be an action in which the hand sign s is input to the imaging device 30, as illustrated in FIG. 5. In this case, the imaging device 30 may function as a detection unit 70 that detects the input action. Alternatively, the predetermined input action may be a button operation by the user on the input device 40, the operation unit 29, or another input device. In this case, the input device 40 or the operation unit 29 may function as the detection unit 70. When the first control unit 111 detects the predetermined input action via the detection unit 70 (FIG. 3, step S20, YES), the first control unit 111 may determine that the first action has been performed (FIG. 3, step S30, YES).

[0048] In the above embodiment, the annotation process is described using the gripping-type end effectors E1 and E2 that grip the workpiece W. However, the annotation process can also be applied to other end effectors. For example, the end effector may be a suction-type tool equipped with an adsorption unit that adsorbs the workpiece W and can adsorb and move the workpiece W. In this case, the first operation is the adsorption operation of the end effector with respect to the workpiece W, and the first information includes the adsorption position and adsorption posture. The adsorption unit provided on the end effector may be configured to adsorb the workpiece W using air, magnetic force, or electrostatic force. The end effector may also be equipped with a known sensor or mechanism that detects the adsorption force of the adsorption unit. The first control unit 111 may determine whether or not adsorption of the workpiece W has been performed ( FIG. 3 , step S40) using the detection results of the sensor or mechanism. Alternatively, the first control unit 111 may determine that the suction operation has been performed when a predetermined input operation is detected (FIG. 3, step S20, YES).

[0049] <Other Embodiments> The various embodiments described above can be modified in various ways, as long as the first control unit 111 determines whether or not a first action of the end effector on an object has been performed, acquires first information including the position and orientation of the end effector when it performed the first action, and generates a data set of the first image including the object and the first information.

[0050] For example, the end effector may be a user's fingers, and the first action may be a predetermined action of the fingers relative to an object. In this case, the first action may be an action of grasping the workpiece W with the fingers or an action of contacting the workpiece W with the fingers. In the above embodiment, just as the first control unit 111 can convert the action of the end effector E2 as a linked tool into the action of a robot hand such as the end effector E1, the action of the user's fingers can also be converted into the action of the robot hand using the correspondence between the action of the fingers and the action of the robot hand. Therefore, by collecting a dataset of a first image including an object and first information including the position and posture of the user's fingers when the user's fingers perform the first action and performing a learning process using the dataset, the position and posture for the robot hand to perform the first action relative to the first image can be output. Therefore, even when the end effector is a user's fingers or includes a user's fingers, the same effects as those of the above embodiment can be achieved.

[0051] When the annotation process is performed using the user's fingers as the end effector, the first control unit 111 may determine that the first action has been performed (FIG. 3, step S30, YES) when the first control unit 111 detects the user's operation of the operation unit 29, the input device 40, or another input device, or a predetermined input action such as a hand sign s (FIG. 3, step S20, YES). The first control unit 111 may also acquire, as the first information, the position and posture of the user's fingers corresponding to the first action via the imaging device 30 (FIG. 3, step S40), and generate a data set D including the first image and the position and posture of the user's fingers (FIG. 3, step S50). When the user's fingers are used as the end effector, the first control unit 111 may attach AR (Argumented Reality) markers to the user's fingers, acquire an image including the AR markers via the imaging device 30, and perform known position estimation using the AR markers to acquire the position and posture of the user's fingers corresponding to the first action. In the case where annotation processing is performed using the user's fingers as the end effector, further modifications such as those exemplified in the following other embodiments 1 to 3 are possible.

[0052] <Alternative Embodiment 1> FIG. 6 shows an interlocking tool 50 attached to a user's finger u1. In the embodiment shown in FIG. 6, the user's finger u1 and the interlocking tool 50 can be considered as an end effector E7 that performs a predetermined operation on an object. The end effector E7 performs a predetermined operation, such as a contact operation, on a workpiece W present in the direction of arrow D1. The interlocking tool 50 includes a limiting mechanism 51 that limits the degree of freedom of movement of the user's finger u1 and a box-shaped member 55 to which an AR marker 54 is attached. Note that the color of the portion of the box-shaped member 55 other than the AR marker 54 is actually colored black or the like so that the AR marker 54 can be identified by a camera such as the imaging device 30, 30a. However, this coloring of this portion is omitted in FIG. 6 and subsequent figures. The limiting mechanism 51 limits the degree of freedom of movement of the user's finger u1 to the same degree of freedom as a robotic hand. The limiting mechanism 51 includes a plate-shaped member 53 and an attachment unit 52 for attaching the plate-shaped member 53 to the finger u1. The attachment unit 52 is formed in a belt shape. The attachment unit 52 may be made of a rubber material. When the limiting mechanism 51 is attached to the finger u1, the movement of the joint of the finger u1 is restricted. Specifically, the joint of the finger u1 cannot bend. The plate-shaped member 53 is an angle restriction member that restricts the angle of the user's finger u1. According to this embodiment, the first control unit 111 performs position estimation using the AR marker 54 to obtain the position and posture (first information) of the finger u1 corresponding to the first movement, and the limiting mechanism 51 can make the movement of the finger u1 closer to the movement of the robotic hand. Therefore, when the position and posture of the finger u1 corresponding to the first movement are converted into the position and posture of the robotic hand, the success rate of the first movement by the robotic hand is improved. In addition, in this embodiment, the interlocking tool 50 is attached to the user's finger u1, so the first information can be acquired without interfering with the movement of the finger u1.

[0053] <Alternative Embodiment 2> FIG. 7 illustrates a situation in which a workpiece W is grasped by two fingers u1 and u2 equipped with interlocking tools 50. In FIG. 7, the user's fingers u1 and u2 and the interlocking tools 50 can be considered as an end effector E8 that performs a predetermined action on an object. In this embodiment, the first control unit 111 performs position estimation using the AR marker 54 to obtain the position and orientation (first information) of the fingers u1 and u2 corresponding to the action (first action) of grasping the workpiece W with the fingers u1 and u2. Furthermore, the limiting mechanisms 51 can more closely resemble the grasping action of a robotic hand. This improves the accuracy of converting the position and orientation of the fingers u1 and u2 corresponding to the grasping action into the position and orientation of the robotic hand. When using the wearable interlocking tool 50 shown in FIGS. 6 and 7, the user may calibrate the attachment state of the interlocking tool 50 by performing a predetermined action after wearing the interlocking tool 50. Furthermore, the limiting mechanism 51 of the interlocking tool 50 shown in FIGS. 6 and 7 may be configured to restrict the angle of the fingers relative to the direction in which the fingers extend, for example.

[0054] <Alternative Embodiment 3> FIG. 8 shows a pen-shaped interlocking tool 60. In FIG. 8, the interlocking tool 60 and the user's fingers (not shown) supporting the interlocking tool 60 can be considered as an end effector E9. The interlocking tool 60 includes a rod-shaped member 56 having a first end 57 and a second end 58, a box-shaped member 55 attached to the second end 58, and a cover 59 covering the outer periphery of the rod-shaped member 56. A user can support the interlocking tool 60 by placing their fingers along the cover 59. For example, the user can move the first end 57 in the direction of arrow D1 to perform a contact action (first action) of bringing the first end 57 into contact with an object. Even in this embodiment, the first control unit 111 can easily obtain first information including the position of the interlocking tool 60 corresponding to the contact action of the first end 57 by performing position estimation using the AR marker 54.

[0055] Alternative Embodiment 4 The first action applied to the annotation system 10 is not limited to an action in which the end effector grasps an object or an action in which the end effector contacts the object, and various actions can be applied. For example, the first image may be an image of an object grasped by the end effector, and the first action may be an action of throwing the object grasped by the end effector into a predetermined area (a release action). In this case, the first control unit 111 can determine whether the first action has been performed based on, for example, (i) a change in the detection value of a sensor provided at the tip of the end effector, (ii) image analysis results of the end effector, or (iii) detection of a user's operation of the operation unit 29, the input device 40, or other input devices, or a predetermined input action such as a hand sign S. Furthermore, the first information may include not only the position of the end effector when the end effector throws the workpiece, but also the posture and speed of the end effector when the end effector throws the workpiece. The first information may include time-series position information (operation sequence) before the end effector releases the workpiece, in addition to the time (moment) when the end effector releases the workpiece. The velocity information may be acquired by an inertial measurement unit (IMU) provided in the end effector. The second information may be information indicating whether the object is within a predetermined area, such as a container, as a result of the first operation. The end effector may be a user's finger or a robot hand. According to this embodiment, a first image of the end effector gripping an object is acquired, and first information regarding the first operation for releasing the object into the predetermined area is generated as a dataset D for the first image, and learning is performed. This allows outputting the release position, release posture, release speed, and other information required for the end effector to perform the release operation in response to an input of the first image including the object. Such annotations and learning regarding the release operation, when combined with the annotations and learning regarding the gripping operation of the above embodiment, contribute to the automation of tasks such as gripping waste and releasing the gripped waste into a predetermined area, for example, in waste disposal facilities.

[0056] Alternative Embodiment 5 The first control unit 111 may acquire a second image, which is an image of the workpiece W after the gripping operation has been performed, from the imaging device 30. The second image may be acquired, for example, together with the first operation determination process (step S30) and the first information acquisition process (step S40) in the annotation process of FIG. 3 . The first control unit 111 may generate the second image and the first information corresponding to the first operation as a data set. The first information including the position corresponding to the first operation can be used as information including the position for placing the workpiece W in the container C, which is the work area. The placing operation corresponds to the reverse playback of the gripping operation. Therefore, according to this embodiment, a data set of the work area image (second image) and the first information including the position for performing the second operation can be generated without actually performing the second operation corresponding to the reverse playback of the first operation. Therefore, the work cost for annotation is effectively reduced. This embodiment is not limited to picking (first operation) and placing (second operation) the workpiece W, but can also be applied to other first and second operations. For example, if the first operation is pulling a door, the first information including the position corresponding to the first operation can be used as information including the position for executing the second operation of pushing the door.

[0057] The above-described various embodiments can be further modified as follows.

[0058] For example, in the annotation process using the end effector E1, the end effector E1 may be equipped with a plurality of markers. The first control unit 111 may track the positions of the markers and acquire the position and orientation of the end effector E1 via the imaging device 30 using a motion capture technique. The imaging device 30 may be equipped with a plurality of calibrated multi-cameras.

[0059] The first information may include at least a position and may not include a posture. Note that the first information may include, in addition to the position when the first action is performed, at least one of information on the position of the end effector before the first action is performed and information on the position of the end effector after the first action is performed. In other words, the first information may include continuous information (time-series data) related to the first action.

[0060] When an object is not present in the working area, the first control unit 111 may generate a data set including an image of the working area and information indicating that the end effector will not perform an operation. The second control unit 211 may perform learning using the data set. This configuration makes it possible to generate a data set that can prevent the end effector from operating when an object is not present.

[0061] In the above embodiment, the first image including the target object is acquired by the imaging device 30, but the first image may be acquired by another visual sensor. Also, if the end effector is a robot hand, the imaging device that acquires the first image may be provided on the robot hand or a robot arm to which the robot hand is attached.

[0062] The operation information of the end effector may be detection values ​​of other sensors, such as a gyro sensor, an acceleration sensor, or an inertial measurement unit including these. The first control unit 111 may store the detection values ​​of the sensors provided in the end effector in the memory 120 and determine whether the first operation has been performed (step S30 in FIG. 3). The first control unit 111 may refer to the memory 120 and perform the determination in step S30 and the processes from step S40 to step S50 shown in FIG. 3 after a series of operations by the end effector has ended.

[0063] In step S30 ( FIG. 3 ) of the above embodiment, the first control unit 111 determines that the first action has been performed if the detection result of the first sensor 21 satisfies a predetermined determination condition. The determination condition is that the detection value of the first sensor 21 is equal to or greater than a first threshold value stored in the memory 120. Alternatively, the determination condition may be, for example, that the positional relationship between two fingers of the hand satisfies a predetermined relationship. The predetermined relationship may be predetermined based on the positions of the two fingers when the two fingers grasp an object. In this case, the first control unit 111 may acquire the positional relationship between the two fingers from the analysis results of the image acquired by the imaging device 30. Alternatively, the determination condition may be that "two or more fingers of the hand have opened and then closed." The first control unit 111 may determine that the fingers of the hand have closed based on the detection value of the first sensor 21. The determination condition may be determined by learning based on changes in the positions of the fingers, changes in the value of the first sensor 21, and the like. The first control unit 111 may determine whether the first action has been performed using the determination conditions determined by the learning. The determination conditions can be changed as appropriate depending on the type of end effector, sensor, or first action.

[0064] The order of the steps in the annotation process described above may be changed as appropriate, any of the steps may be performed simultaneously, or any of the steps may be omitted. For example, step S30 of acquiring the first information and step S45 of determining whether the annotation has been successful may be changed or performed simultaneously. Furthermore, step S45 of determining whether the annotation has been successful may be omitted in the annotation process.

[0065] In the above embodiment, the user operates the operation unit 29 to control the end effector E1 via the robot control device 25. Alternatively, the robot control unit 26 of the robot control device 25 may use an initial learning model to output a position and orientation for executing the first action for the acquired first image. The first control unit 111 may further generate a data set D including the first image, first information including the output position and orientation, and second information indicating the success or failure of the first action. The second control unit 211 may perform learning using the data set D indicating the second information indicating the success or failure of the first action, and update the initial learning model. This configuration can further improve the accuracy of learning, thereby increasing the success rate of the first action performed by the robot control device 25.

[0066] In the above embodiment, the annotation device 100 has the functions of the first control unit 111, and the learning device 200 has the functions of the second control unit 211. In contrast, the first control unit 111 and the second control unit 211 may be included in a single control device. Furthermore, at least some of the functions of the first control unit 111 may be executed by the second control unit 211, and at least some of the functions of the second control unit 211 may be executed by the first control unit 111.

[0067] Additionally, the functions of the elements disclosed herein can be performed using circuits or processing circuits, including general-purpose processors, special-purpose processors, integrated circuits, ASICs (Application Specific Integrated Circuits), conventional circuits, and / or combinations thereof, configured or programmed to perform the disclosed functions. A processor is considered a processing circuit or circuit because it includes transistors and other circuitry. In this disclosure, a circuit, unit, or means is hardware that performs the recited functions or hardware that is programmed to perform the recited functions. The hardware may be hardware disclosed herein or other known hardware that is programmed or configured to perform the recited functions. In the case of a processor, where the hardware is considered a type of circuit, the circuit, means, or unit is a combination of hardware and software, and the software is used to configure the hardware and / or processor.

[0068] The present disclosure is not limited to the above-described embodiments and can be realized in various forms without departing from the spirit thereof. For example, the present disclosure can also be realized in the following aspects. The technical features in the above embodiments corresponding to the technical features in each aspect described below can be appropriately replaced or combined to solve some or all of the problems of the present disclosure or to achieve some or all of the effects of the present disclosure. Furthermore, if a technical feature is not described as essential in this specification, it can be appropriately deleted.

[0069] <1> According to a first aspect of the present disclosure, there is provided an annotation device. The annotation device includes a first control unit that generates a learning dataset and a storage unit that stores the dataset. The first control unit acquires a first image including an object captured by an imaging device. The first control unit acquires operation information regarding an operation of an end effector with respect to the object and determines whether the end effector has performed a predetermined first operation with respect to the object. When the end effector has performed the first operation, the first control unit generates the first image and first information about the end effector as the dataset. The first information includes a position corresponding to the first operation of the end effector. The control unit stores the generated dataset in the storage unit. According to this aspect, the first control unit generates, as the learning dataset, the first image including the object and the position and orientation when the end effector actually performed the first operation with respect to the object. In other words, the first control unit annotates the first image with the actual position and posture at which the end effector performed the first action. Therefore, a learning dataset of the first image and the first information can be collected regardless of the material or shape of the object. Furthermore, because the first control unit annotates the first information on the first image, the work cost for annotation is reduced compared to an embodiment in which a user performs annotation. In this embodiment, the first action can be a variety of actions depending on the end effector, such as a grasping action to grasp the object, a suction action to suck the object, a contact action to contact the object, or a pre-grasp of these actions.

[0070] <2> In the above aspect, the operation information may include a detection value of a sensor provided in the end effector, the detection value changing depending on the operation of the end effector relative to the object. The first control unit may be configured to acquire the detection value of the sensor. The first control unit may determine that the end effector has performed the first operation relative to the object when the detection value satisfies a predetermined determination condition. According to this aspect, the first control unit determines whether the end effector has performed the first operation using the detection value of the sensor, and therefore can determine whether the first operation has been performed without interrupting the operation of the end effector. This reduces the time required for annotation.

[0071] <3> In the above aspect, the operation information may include a predetermined input operation to be performed on an input device. The first control unit may be configured to acquire the predetermined input operation via the input device. When the first control unit acquires the predetermined input operation, the first control unit may determine that the end effector has performed the first operation on the object. According to this aspect, the first control unit can acquire the predetermined input operation and determine whether the first operation has been performed. Note that the predetermined input operation may be a user gesture such as a hand sign, or may be a user operation of another input device provided separately from the annotation device.

[0072] <4> In the above aspect, the end effector may be a gripping tool that grips the object. The first action may be an action of the gripping tool gripping the object. The first information may include a position of the gripping tool when it grips the object. According to this aspect, it is possible to collect a data set of a first image including the object and a gripping position and gripping posture of the end effector for gripping the object, regardless of the object. In this aspect, the first information may further include an opening width of a plurality of tips of the gripping tool.

[0073] <5> In the above aspect, the gripping tool may be a robot hand attached to a robot arm and configured to grip the object. The first motion may be a motion of the robot hand gripping the object. The first information may include a position of the robot hand when gripping the object. The first control unit may acquire the first information of the robot hand via at least one of a control device that controls the robot hand and the imaging device. According to this aspect, the first control unit acquires the first information via the control device or the imaging device, thereby reducing the operational cost for generating a data set.

[0074] <6> In the above-described embodiment, the end effector may include a linked tool supported by a user's fingers and linked to the movement of the fingers. The first information may include a position of the linked tool when the end effector including the linked tool performs the first action on the object. According to this embodiment, a data set of a first image and a position of the linked tool can be generated. Note that the linked tool may operate in linkage with the user's fingers that perform the first action directly on the object.

[0075] <7> In the above aspect, the first control unit may acquire the first information via the imaging device. According to this aspect, the first control unit acquires the position of the interlocking tool via the imaging device, thereby reducing the operational cost for generating a data set. The interlocking tool may include a marker. The first control unit may acquire the position of the end effector corresponding to the first operation as the first information by acquiring and analyzing an image including the marker via the imaging device.

[0076] <8> In the above-described aspect, the interlocking tool may have a limiting mechanism. The limiting mechanism may include an attachment portion for attaching the interlocking tool to a user's finger. The limiting mechanism may be configured to limit the degree of freedom of movement of the user's finger. According to this aspect, the limiting mechanism can make the movement of the finger closer to, for example, the movement of a robot hand. Therefore, the position of the interlocking tool corresponding to the first movement can be accurately converted to the position of the robot hand.

[0077] <9> In the above-described embodiment, the limiting mechanism may include an angle limiting member that limits the angle of the user's fingers. According to this embodiment, the limiting mechanism can make the movement of the fingers more similar to the movement of the robot hand. Therefore, the position of the interlocking tool corresponding to the first movement can be converted with high accuracy based on the position of the robot hand. Note that the interlocking tool described in the above embodiments <6> to <9> is not limited to the user's fingers, but may be supported by the user's body and linked to the movement of at least a part of the body. For example, the interlocking tool may be attached to the user's leg and linked to the movement of the leg.

[0078] <10> In the above aspect, the first control unit may use the operation information to determine whether the first operation by the end effector is successful. The first control unit may generate second information indicating whether the first operation is successful, the first image, and the first information as the data set. According to this aspect, a data set that can further improve learning accuracy is generated.

[0079] <11> In the above aspect, the first image may include the object and a working area in which the object is placed. The first control unit may further acquire, from the imaging device, a second image corresponding to the first image, the second image being an image of the working area after the first action has been performed on the object. The first control unit may generate the second image and the first information corresponding to the first action as the dataset. According to this aspect, a dataset of the second image and the first information is generated, so that a dataset for a second action corresponding to a reverse playback of the first action can be generated without performing the second action. In this aspect, the first action may be a grasping action, and the first information may include a position of the end effector when grasping the object. The second control unit may perform learning using the dataset of the second image and the first information to generate a learning model that outputs a position for releasing the grasped object in response to an input of the second image. In this aspect, the second image may be an image of the working area after the object has been removed from the working area by performing the first action on the object.

[0080] <12> A second aspect of the present disclosure provides a learning system including the annotation device of the above aspect and a learning device having a second control unit that performs supervised learning. The second control unit may generate a learning model based on the dataset, which outputs the first information corresponding to the first action of the end effector in response to an input of an image including the object. According to this aspect, the first information used for learning is not an estimated value but includes the actual position of the end effector when it performed the first action on the object. Therefore, a learning model can be generated that can output first information including the position of the end effector for performing the first action in response to an input of a first image including the object, regardless of the material or shape of the object. Furthermore, when the first action is performed using the position output from the learning model, the success rate of the first action can be increased. In this aspect, the second control unit may use the dataset including the second information to generate a learning model that outputs the first information corresponding to the first action of the end effector in response to an input of an image including the object. According to this aspect, when the first action is performed using the position output from the generated learning model, the success rate of the first action can be further increased.

[0081] <13> According to a third aspect of the present disclosure, there is provided an annotation system. The annotation system includes an imaging device, a detection unit, and an annotation device. The imaging device captures a first image including an object. The detection unit detects motion information of an end effector that performs a predetermined motion on the object. The annotation device includes a first control unit that generates a learning dataset, and a storage unit. The first control unit acquires the motion information and determines whether the end effector has performed a predetermined first motion on the object. When the end effector has performed the first motion, the first control unit generates the first image and first information as the dataset. The first information includes a position corresponding to the first motion. The first control unit stores the generated dataset in the storage unit.

[0082] <14> In the above aspect, the detection unit may include a sensor provided in the end effector. The operation information may include a detection value of the sensor that changes depending on the operation of the end effector with respect to the object. The first control unit may determine that the end effector has performed the first operation with respect to the object when the detection value exceeds a predetermined threshold.

[0083] <15> In the above aspect, the detection unit may include an input device that detects a predetermined input motion. The motion information may include the predetermined input motion executed on the input device. The first control unit may determine that the end effector has executed the first motion when the predetermined input motion is received via the detection unit.

[0084] <16> According to a fourth aspect of the present disclosure, there is provided an annotation method, which includes: acquiring a first image captured by an imaging device and including an object; acquiring operation information related to an operation of an end effector on the object; determining whether the end effector has performed a predetermined first operation on the object; and, if the end effector has performed the first operation, generating the first image and first information about the end effector including a position corresponding to the first operation as a dataset for supervised learning.

[0085] <17> According to a fifth aspect of the present disclosure, there is provided an annotation program that causes a computer to acquire a first image including an object captured by an imaging device, acquire operation information regarding an operation of an end effector with respect to the object, determine whether the end effector has performed a predetermined first operation with respect to the object, and, when the end effector has performed the first operation, generate, as a dataset for supervised learning, the first image and first information about the end effector including a position corresponding to the first operation.

[0086] The present disclosure may be realized in various forms other than those described above, such as a non-transitory storage medium on which a computer program for implementing at least one function of the annotation device 100, the first control unit 111, and the second control unit 211 is recorded.

[0087] 10: Annotation system, 20: Robot arm, 21: First sensor, 22: Second sensor, 25: Robot control device, 26: Robot control unit, 29: Operation unit, 30: Imaging device, 40: Input device, 50: Interlocking tool, 51: Restriction mechanism, 52: Mounting unit, 53: Plate-shaped member, 54: AR marker, 55: Box-shaped member, 56: Rod-shaped member, 57: First end, 58: Second end, 59: Cover, 70: Detection unit, 100: Annotation device, 110: CPU, 111: First control unit, 120: Memory, 130 : Interface circuit, 200: Learning device, 210: CPU, 220: Memory, 211: Second control unit, 230: Interface circuit, 300: Learning system, C: Container, D: Data set, E1, E2, E2a, E7, E8, E9: End effector, E21: Base, E22: Tip, J: Joint, L: Link, P1: Annotation program, P2: Learning program, T: Tip, W, W1, W2, W3: Workpiece, s: Hand sign, tr: Transparent part, u: User, u1, u2: Fingers

Claims

1. An annotation device, comprising: a first control unit that generates a learning dataset; and a storage unit that stores the dataset, wherein the first control unit: obtains a first image including an object captured by an imaging device; obtains operation information regarding the operation of an end effector with respect to the object, and determines whether the end effector has executed a predetermined first operation on the object; when the end effector has executed the first operation, generates the first image and first information of the end effector including a position corresponding to the first operation as the dataset; and stores the generated dataset in the storage unit. Annotation device.

2. The annotation device according to claim 1, wherein the operation information is a detection value of a sensor provided on the end effector, and includes a detection value that changes due to the operation of the end effector with respect to the object, and the first control unit: obtains the detection value of the sensor; and determines that the end effector has executed the first operation on the object when the detection value satisfies a predetermined determination condition. Annotation device.

3. The annotation device according to claim 1, wherein the operation information includes a predetermined input operation executed on an input device, and the first control unit determines that the end effector has executed the first operation on the object when the predetermined input operation is obtained via the input device. Annotation device.

4. The annotation device according to claim 1, wherein the end effector is a gripping tool that grips the object, the first operation is an operation in which the gripping tool grips the object, and the first information includes the position of the gripping tool when the gripping tool grips the object. Annotation device.

5. The annotation device according to claim 4, wherein the gripping tool is a robot hand attached to a robot arm for gripping the object, the first operation is an operation of the robot hand for gripping the object, the first information includes a position when the robot hand grips the object, and the first control unit acquires the first information of the robot hand via at least one of a robot control device for controlling the robot hand and the imaging device.

6. The annotation device according to claim 1, wherein the end effector includes an interlocking tool supported by a user's finger and interlocking with the movement of the finger, and the first information includes a position of the interlocking tool when the end effector performs the first operation on the object.

7. The annotation device according to claim 6, wherein the first control unit acquires the first information via the imaging device.

8. The annotation device according to claim 6, wherein the interlocking tool is a restricting mechanism including a mounting portion for mounting the interlocking tool on the finger and having a restricting mechanism for restricting the degree of freedom of movement of the finger.

9. The annotation device according to claim 8, wherein the restricting mechanism includes an angle restricting member for restricting the angle of the finger.

10. The annotation device according to claim 1, wherein the first control unit determines whether the first operation by the end effector is successful using the operation information, and generates the second information indicating whether the first operation is successful, the first image, and the first information as the dataset.

11. The annotation device according to claim 1, wherein the first image includes the object and the work area where the object is disposed, and the first control unit further acquires, from the imaging device, a second image corresponding to the first image, which is an image of the work area after the first operation is performed on the object, and generates the second image and the first information corresponding to the first operation as the dataset.

12. A learning system, comprising: the annotation device according to claim 1; and a learning device having a second control unit that performs supervised learning, wherein the second control unit generates a learning model that outputs, for an input of an image including the object, the first information corresponding to the first operation of the end effector, using the dataset.

13. An annotation system, comprising: an imaging device that captures a first image including an object; a detection unit that detects operation information of an end effector that performs a predetermined operation on the object; a first control unit that generates a dataset for learning; and a storage unit, wherein the first control unit: acquires the operation information and determines whether the end effector has performed a predetermined first operation on the object; when the end effector has performed the first operation, generates the first image and the first information of the end effector including a position corresponding to the first operation as the dataset; and stores the generated dataset in the storage unit.

14. The annotation system according to claim 13, wherein the detection unit includes a sensor provided on the end effector, the operation information is a detection value of the sensor, and the detection value includes a detection value that changes according to the operation of the end effector on the object, and the first control unit determines that the end effector has performed the first operation on the object when the detection value satisfies a predetermined determination condition.

15. An annotation method, comprising: acquiring a first image including an object, captured by an imaging device; acquiring operation information regarding an operation of an end effector on the object and determining whether the end effector has performed a predetermined first operation on the object; and when the end effector has performed the first operation, generating the first image and the first information of the end effector including a position corresponding to the first operation as a dataset for supervised learning.

Citation Information

Patent Citations

  • Operation teaching method of parallel link robot and parallel link robot

    JP2014217913A

  • Robot teaching device and method, and robot system

    JP2016047591A

  • Machine learning apparatus, robot system, and machine learning method for learning workpiece take-out motion

    JP2017030135A

  • Machine learning device, machine learning system, data processing system and machine learning method

    JP2020082322A