Training data generation device, machine learning device, and robot joint angle estimation device

By generating training data and using machine learning to build a learned model, the problem of joint axis angles not being available for robots without logging functionality or dedicated I/F robots is solved, thus enabling accurate estimation of robot joint axis angles.

CN116615317BActive Publication Date: 2026-05-29FANUC LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FANUC LTD
Filing Date
2021-12-14
Publication Date
2026-05-29

Smart Images

  • Figure CN116615317B_ABST
    Figure CN116615317B_ABST
Patent Text Reader

Abstract

Even a robot not equipped with a log function or a dedicated I / F can easily obtain the angles of the joint axes of the robot. A training data generation device generates training data for generating a learned model that inputs a two-dimensional image of a robot taken by a camera, and a distance and an inclination between the camera and the robot, and estimates angles of a plurality of joint axes included in the robot at the time of taking the two-dimensional image and a two-dimensional pose representing center positions of the plurality of joint axes in the two-dimensional image, the training data generation device including: an input data acquisition unit that acquires the two-dimensional image of the robot taken by the camera, and the distance and the inclination between the camera and the robot; and a label acquisition unit that acquires the angles of the plurality of joint axes and the two-dimensional pose at the time of taking the two-dimensional image as label data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a training data generation device, a machine learning device, and a robot joint angle estimation device. Background Technology

[0002] One known method for setting the tool tip point of a robot involves teaching the robot to perform actions such as contacting a gripper in multiple postures, and then calculating the tool tip point based on the angles of the joint axes in each posture. For example, see Patent Document 1.

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 8-085083 Summary of the Invention

[0006] The problem that the invention aims to solve

[0007] However, in order to obtain the angles of each joint axis of the robot, a logging function needs to be installed in the robot program, or the robot's dedicated I / F needs to be used to acquire the data.

[0008] However, in robots without logging functionality or dedicated I / F (interconnection / free connection), it is impossible to obtain the angles of each joint axis.

[0009] Therefore, it is expected that even robots without logging functionality or dedicated I / F can easily obtain the angles of each joint axis.

[0010] Methods for solving problems

[0011] (1) One method of the training data generation apparatus of this disclosure generates training data for generating a learned model, wherein the learned model is input to a two-dimensional image of a robot captured by a camera, the distance and tilt between the camera and the robot, and estimates the angles of a plurality of joint axes contained in the robot when the two-dimensional image was captured, and a two-dimensional pose for representing the center position of the plurality of joint axes in the two-dimensional image. The training data generation apparatus comprises: an input data acquisition unit that acquires the two-dimensional image of the robot captured by the camera, the distance and tilt between the camera and the robot; and a label acquisition unit that acquires the angles of the plurality of joint axes and the two-dimensional pose when the two-dimensional image was captured as label data.

[0012] (2) One aspect of the machine learning apparatus of the present disclosure includes a learning unit that performs supervised learning based on training data generated by the training data generation device of (1) to generate a learned model.

[0013] (3) One embodiment of the robot joint angle estimation device of the present disclosure comprises: a learned model generated by the machine learning device of (2); an input unit that inputs a two-dimensional image of the robot captured by a camera, the distance between the camera and the robot, and the tilt; and an estimation unit that inputs the two-dimensional image and the distance between the camera and the robot input by the input unit into the learned model, and estimates the angles of a plurality of joint axes contained in the robot when the two-dimensional image was captured, and a two-dimensional pose representing the center position of the plurality of joint axes in the two-dimensional image.

[0014] Invention Effects

[0015] According to one method, even robots without logging functionality or dedicated I / F can easily obtain the angles of each joint axis of the robot. Attached Figure Description

[0016] Figure 1 This is a functional block diagram representing a functional structural example of a system in one implementation phase of the learning process.

[0017] Figure 2A An example of a frame image where the angle of joint axis J4 is 90 degrees.

[0018] Figure 2B An example of a frame image where the angle of joint axis J4 is -90 degrees.

[0019] Figure 3 This represents an example of using the amount of training data.

[0020] Figure 4 An example representing the coordinate values ​​of the joint axis in the standardized XY coordinate system.

[0021] Figure 5 An example illustrating the relationship between a two-dimensional skeleton estimation model and a joint angle estimation model.

[0022] Figure 6 An example of a feature map representing the joint axes of a robot.

[0023] Figure 7 This is an example of comparing the output of a frame image with that of a two-dimensional skeleton estimation model.

[0024] Figure 8 This represents an example of a model for estimating joint angles.

[0025] Figure 9 This is a functional block diagram representing a functional structural example of a system in the application phase.

[0026] Figure 10This is a flowchart illustrating the presumption processing of the terminal device during the application phase.

[0027] Figure 11 An example representing the structure of a system. Detailed Implementation

[0028] Hereinafter, an embodiment of the present disclosure will be described using the accompanying drawings.

[0029] <One Implementation Method>

[0030] First, an overview of this embodiment will be described. In this embodiment, a terminal device such as a smartphone operates as a training data generation device (annotation automation device). This training data generation device generates training data for generating a learned model during the learning phase. The learned model is input from a two-dimensional image of the robot captured by a camera included in the terminal device, as well as the distance and tilt between the camera and the robot. The angles of multiple joint axes contained in the robot and a two-dimensional pose representing the center position of the multiple joint axes are estimated when the two-dimensional image was captured.

[0031] The terminal device provides the generated training data to the machine learning device, which performs supervised learning based on the provided training data to generate a learned model. The machine learning device then provides the generated learned model to the portable terminal.

[0032] During the application phase, the terminal device acts as a robot joint angle estimation device. This robot joint angle estimation device inputs the two-dimensional image of the robot captured by the camera, as well as the distance and tilt between the camera and the robot, into the learned model to estimate the angles of multiple joint axes of the robot when the two-dimensional image was captured, as well as the two-dimensional pose representing the center position of multiple joint axes.

[0033] Therefore, according to this embodiment, the problem of "easily obtaining the angles of each joint axis of a robot even if the robot does not have a logging function or a dedicated I / F" can be solved.

[0034] The above is a summary of this implementation method.

[0035] Next, the structure of this embodiment will be described in detail with reference to the accompanying drawings.

[0036] <Systems in the Learning Phase>

[0037] Figure 1 This is a functional block diagram illustrating the functional structure of a system in one implementation phase of the learning process. For example... Figure 1 As shown, system 1 includes a robot 10, a terminal device 20 that serves as a training data generation device, and a machine learning device 30.

[0038] Robot 10, terminal device 20, and machine learning device 30 can be interconnected via a network not shown, such as a wireless LAN (Local Area Network), Wi-Fi (registered trademark), or a mobile phone network conforming to standards such as 4G or 5G. In this case, robot 10, terminal device 20, and machine learning device 30 have a communication unit (not shown) for communicating with each other through this connection. Furthermore, robot 10 and terminal device 20 can send and receive data via the communication unit (not shown), but data can also be sent and received via a robot control device (not shown) used to control the actions of robot 10.

[0039] Additionally, as described later, terminal device 20 may also include machine learning device 30. Furthermore, terminal device 20 and machine learning device 30 may also be included in a robot control device (not shown).

[0040] In the following description, the terminal device 20, which acts as a training data generation device, only acquires data at a timed synchronization with all data acquisition as training data. For example, if the camera included in the terminal device 20 captures frame images at 30 frames per second, and can acquire the angles of multiple joint axes included in the robot 10 at a period of 100 milliseconds, and can immediately acquire other data, the terminal device 20 outputs the training data as a file at a period of 100 milliseconds.

[0041] <Robot 10>

[0042] Robot 10, such as an industrial robot known to those skilled in the art, is equipped with a joint angle response server 101. Robot 10 drives movable parts (not shown) of the robot 10 by driving commands from a robot control device (not shown) to drive servo motors (not shown) configured on each of the multiple joint axes (not shown) included in the robot 10.

[0043] Furthermore, the robot 10 will be described below as a 6-axis vertical multi-joint robot with 6 joint axes J1 to J6, but it may also be a vertical multi-joint robot other than 6 axes, or a horizontal multi-joint robot or a parallel linkage robot, etc.

[0044] The joint angle response server 101, such as a computer, outputs joint angle data, including the angles of the joint axes J1 to J6 of the robot 10, at a predetermined period of 100 milliseconds or similar, based on a request from the terminal device 20 (described later as a training data generation device). Furthermore, as described above, the joint angle response server 101 can output directly to the terminal device 20, or via a robot control device (not shown), to the terminal device 20, which is also a training data generation device.

[0045] In addition, the joint angle response server 101 can be a device independent of the robot 10.

[0046] <Terminal Device 20>

[0047] Terminal device 20 is, for example, a smartphone, a tablet, augmented reality (AR) glasses, mixed reality (MR) glasses, etc.

[0048] like Figure 1 As shown, the terminal device 20, during the application phase, functions as a training data generation device and includes a control unit 21, a camera 22, a communication unit 23, and a storage unit 24. Furthermore, the control unit 21 includes a 3D object recognition unit 211, a self-position estimation unit 212, a joint angle acquisition unit 213, a forward kinematics calculation unit 214, a projection unit 215, an input data acquisition unit 216, and a tag acquisition unit 217.

[0049] The camera 22, such as a digital camera, captures images of the robot 10 at a predetermined frame rate (e.g., 30 frames per second) based on the operator's actions, generating a two-dimensional image, or frame image, projected onto a plane perpendicular to the optical axis of the camera 22. The camera 22 outputs the generated frame image to the control unit 21 (described later) at a predetermined period of 100 milliseconds or similar for synchronization. Furthermore, the frame image generated by the camera 22 can also be a visible light image, such as an RGB color image or a grayscale image.

[0050] The communication unit 23 is a communication control device that transmits and receives data with networks such as wireless LAN (Local Area Network), Wi-Fi (registered trademark), and mobile phone networks conforming to standards such as 4G or 5G. The communication unit 23 can communicate directly with the joint angle response server 101, or it can communicate with the joint angle response server 101 via a robot control device (not shown) used to control the movement of the robot 10.

[0051] The storage unit 24 may be, for example, a ROM (Read Only Memory) or an HDD (Hard Disk Drive), storing system programs executed by the control unit 21 (described later) and training data generation applications. Additionally, the storage unit 24 may also store input data 241, tag data 242, and 3D recognition model data 243.

[0052] Input data 241 stores the input data acquired by the input data acquisition unit 216, which will be described later.

[0053] Tag data 242 stores tag data acquired by tag acquisition unit 217, which will be described later.

[0054] The 3D recognition model data 243 stores, for example, feature quantities such as edge quantities extracted from multiple frame images as 3D recognition models. These multiple frame images are multiple frame images of the robot 10 captured by the camera 22 at various distances and angles (tilts) after the robot 10's posture and orientation have been changed in advance. Additionally, the 3D recognition model data 243 may also store the 3D coordinates of the origin of the robot coordinate system (hereinafter also referred to as the "robot origin") in the world coordinate system when each frame image of the 3D recognition model is captured, as well as information representing the directions of the X, Y, and Z axes of the robot coordinate system in the world coordinate system, corresponding to the 3D recognition model.

[0055] Furthermore, when the training data generation application is launched on the terminal device 20, a world coordinate system is defined, and the position of the origin of the camera coordinate system of the terminal device 20 (camera 22) is obtained as the coordinate value of the world coordinate system. Also, when the terminal device 20 (camera 22) moves after the training data generation application is launched, the origin of the camera coordinate system moves from the origin of the world coordinate system.

[0056] <Control Department 21>

[0057] The control unit 21 includes a CPU (Central Processing Unit), ROM, RAM, CMOS (Complementary Metal-Oxide-Semiconductor) memory, etc., which are configured to communicate with each other via a bus, as is known to those skilled in the art.

[0058] The CPU is the processor of the overall control terminal device 20. The CPU reads the system program and training data generation application stored in ROM via the bus, and controls the entire terminal device 20 according to the system program and training data generation application. Thus, as... Figure 1 As shown, the control unit 21 is configured to perform the functions of a 3D object recognition unit 211, a self-position estimation unit 212, a joint angle acquisition unit 213, a forward kinematics calculation unit 214, a projection unit 215, an input data acquisition unit 216, and a tag acquisition unit 217. Various data, such as temporary calculation data and display data, are stored in RAM. Furthermore, the CMOS memory 14 is supported by a battery (not shown) and is configured as a non-volatile memory that maintains its stored state even when the power supply to the control device 20 is turned off.

[0059] <3D Object Recognition Unit 211>

[0060] The 3D object recognition unit 211 acquires frame images of the robot 10 captured by the camera 22. The 3D object recognition unit 211 extracts feature quantities such as edge quantities from the frame images of the robot 10 captured by the camera 22, for example, using a known 3D coordinate recognition method for robots. The 3D object recognition unit 211 matches the extracted feature quantities with the feature quantities of the 3D recognition model stored in the 3D recognition model data 243. Based on the matching result, the 3D object recognition unit 211 obtains, for example, the 3D coordinate values ​​of the robot origin in the world coordinate system in the 3D recognition model with the highest consistency, as well as information representing the directions of the X-axis, Y-axis, and Z-axis of the robot coordinate system.

[0061] Furthermore, the 3D object recognition unit 211 uses a method for recognizing the robot's 3D coordinates to obtain the 3D coordinates of the robot's origin in the world coordinate system, as well as information indicating the directions of the X, Y, and Z axes of the robot's coordinate system, but is not limited to this. For example, the 3D object recognition unit 211 may also install markers such as a checkerboard pattern on the robot 10, and obtain the 3D coordinates of the robot's origin in the world coordinate system, as well as information indicating the directions of the X, Y, and Z axes of the robot's coordinate system, from an image of the marker captured by the camera 22 based on known marker recognition technology.

[0062] Alternatively, an indoor positioning device such as UWB (Ultra Wide Band) can be installed on the robot 10. The three-dimensional object recognition unit 211 obtains the three-dimensional coordinates of the robot's origin in the world coordinate system and the information representing the directions of the X-axis, Y-axis and Z-axis of the robot coordinate system from the indoor positioning device.

[0063] <Self-position estimation section 212>

[0064] The self-position estimation unit 212 uses a known self-position estimation method to obtain the three-dimensional coordinates of the origin of the camera coordinate system of camera 22 in the world coordinate system (hereinafter also referred to as "the three-dimensional coordinates of camera 22"). The self-position estimation unit 212 can calculate the distance and tilt between camera 22 and robot 10 based on the obtained three-dimensional coordinates of camera 22 and the three-dimensional coordinates obtained by the three-dimensional object recognition unit 211.

[0065] <Joint Angle Acquisition Unit 213>

[0066] The joint angle acquisition unit 213 sends a request to the joint angle response server 101 via the communication unit 23 at a predetermined period of 100 milliseconds or the above-mentioned synchronization, to obtain the angles of the joint axes J1 to J6 of the robot 10 when the frame image was captured.

[0067] <Forward Kinematics Calculation Section 214>

[0068] The forward kinematics calculation unit 214, for example, uses a predefined Denavit-Hartenberg (DH) parameter table to solve the forward kinematics based on the angles of joint axes J1 to J6 obtained by the joint angle acquisition unit 213, calculates the three-dimensional coordinate values ​​of the center positions of joint axes J1 to J6, and calculates the three-dimensional pose of the robot 10 in the world coordinate system. Furthermore, the DH parameter table is pre-generated based on the specifications of the robot 10 and stored in the storage unit 24.

[0069] <Projection Section 215>

[0070] The projection unit 215, for example, uses a known projection method onto a two-dimensional plane to position the center positions of the joint axes J1 to J6 of the robot 10, calculated by the forward kinematics calculation unit 214, in a three-dimensional space of the world coordinate system. The viewpoint of the camera 22, determined by the distance and tilt between the camera 22 and the robot 10 calculated by the self-position estimation unit 212, is projected onto a projection plane determined by the distance and tilt between the camera 22 and the robot 10, thereby generating two-dimensional coordinates (pixel coordinates) (x, y, y) of the center positions of the joint axes J1 to J6. i y i Let i be the two-dimensional pose of robot 10. In addition, i is an integer from 1 to 6.

[0071] In addition, such as Figure 2A as well as Figure 2B As shown, depending on the pose of robot 10 and the shooting direction, there are cases where joint axes are hidden in the frame image.

[0072] Figure 2A An example of a frame image where the angle of joint axis J4 is 90 degrees. Figure 2BAn example of a frame image where the angle of joint axis J4 is -90 degrees.

[0073] exist Figure 2A In the frame image, joint axis J6 is hidden and not visible. On the other hand, in Figure 2B The frame image shows joint axis J6.

[0074] Therefore, the projection unit 215 connects adjacent joint axes of the robot 10 to each other using line segments, and defines the thickness of each line segment with a preset link width of the robot 10. Based on the three-dimensional pose of the robot 10 calculated by the forward kinematics calculation unit 214, and the optical axis direction of the camera 22 determined by the distance and tilt between the camera 22 and the robot 10, the projection unit 215 determines whether other joint axes exist on the line segments. Other joint axes are located in a depth direction opposite to that of the camera 22 side relative to the line segment. Figure 2A In that case, the projection section 215 will project other joint axes Ji ( Figure 2A The confidence level of the joint axis J6) i Set to "0". On the other hand, when other joint axes Ji are located on the side of camera 22 relative to the line segment... Figure 2B In that case, the projection section 215 will project other joint axes Ji ( Figure 2B The confidence level of the joint axis J6) i Set to "1".

[0075] That is, the projection unit 215 can project the two-dimensional coordinates (pixel coordinates) (x, y, z) of the center position of the projected joint axes J1 to J6. i y i The two-dimensional pose of robot 10 includes a confidence level c indicating whether each joint axis J1 to J6 was captured in the frame image. i .

[0076] In addition, it is preferable to prepare multiple training data for supervised learning in the machine learning device 30 described later.

[0077] Figure 3 This represents an example of using the amount of training data.

[0078] like Figure 3 As shown, for example, to increase training data, the projection unit 215 randomly assigns a distance and tilt between the camera 22 and the robot 10, causing the three-dimensional pose of the robot 10 calculated by the forward kinematics calculation unit 214 to rotate. The projection unit 215 can also generate multiple two-dimensional poses of the robot 10 by projecting the rotated three-dimensional pose of the robot 10 onto a two-dimensional plane determined by the randomly assigned distance and tilt.

[0079] <Input Data Acquisition Unit 216>

[0080] The input data acquisition unit 216 acquires the frame image of the robot 10 captured by the camera 22, and the distance and tilt between the camera 22 that captured the frame image and the robot 10 as input data.

[0081] Specifically, the input data acquisition unit 216, for example, acquires the frame image from the camera 22 as input data. In addition, the input data acquisition unit 216 acquires the distance and tilt between the camera 22 and the robot 10 when the acquired frame image is captured from the self-position estimation unit 212. The input data acquisition unit 216 acquires the acquired frame image, and the distance and tilt between the camera 22 and the robot 10 as input data, and stores the acquired input data in the input data 241 of the storage unit 24.

[0082] In addition, when generating the joint angle estimation model 252 described below that is configured to complete learning the model, as Figure 4 shown, the input data acquisition unit 216 can convert the two-dimensional coordinates (pixel coordinates) (x i , y i ) of the center positions of the joint axes J1 to J6 included in the two-dimensional posture generated by the projection unit 215 into XY coordinate values that are respectively normalized with the base link of the robot 10, that is, the joint axis J1 as the origin, divided by the width of the frame image so that -1 < X < 1, and divided by the height of the frame image so that -1 < Y < 1.

[0083] <Label acquisition unit 217>

[0084] The label acquisition unit 217 acquires the angles of the joint axes J1 to J6 of the robot 10 when the frame image is captured at a predetermined period such as 100 milliseconds that can be synchronized, and the two-dimensional posture indicating the center positions of the joint axes J1 to J6 of the robot 10 in the frame image as label data (correct answer data).

[0085] Specifically, the label acquisition unit 217, for example, acquires the two-dimensional posture indicating the center positions of the joint axes J1 to J6 of the robot 10 and the angles of the joint axes J1 to J6 from the projection unit 215 and the joint angle acquisition unit 213 as label data (correct answer data). The label acquisition unit 217 stores the acquired label data in the label data 242 of the storage unit 24.

[0086] <Machine learning device 30>

[0087] The machine learning device 30, for example, acquires the frame image of the robot 10 captured by the camera 22 stored in the above input data 241 and the distance and tilt between the camera 22 that captured the frame image and the robot 10 from the terminal device 20 as input data.

[0088] In addition, the machine learning device 30 obtains from the terminal device 20 the angles of the joint axes J1 to J6 of the robot 10 when the frame image was captured by the camera 22 and the two-dimensional pose representing the center position of the joint axes J1 to J6, which are stored in the tag data 242, as labels (correct solution data).

[0089] The machine learning device 30 performs supervised learning using training data consisting of the combination of the acquired input data and labels, and constructs the fully learned model described later.

[0090] Thus, the machine learning device 30 is able to provide the constructed, learned model to the terminal device 20.

[0091] The machine learning device 30 will be described in detail.

[0092] like Figure 1 As shown, the machine learning device 30 has a learning unit 301 and a storage unit 302.

[0093] As described above, the learning unit 301 receives a combination of input data and labels from the terminal device 20 as training data. The learning unit 301 uses the received training data to perform supervised learning. Thus, when the terminal device 20 is operating as a robot joint angle estimation device as described later, it inputs frame images of the robot 10 captured by the camera 22, as well as the distance and tilt between the camera 22 and the robot 10, to construct a learned model for outputting the angles of the joint axes J1 to J6 of the robot 10 and a two-dimensional pose representing the center position of the joint axes J1 to J6.

[0094] Furthermore, in this invention, the learned model is constructed by consisting of a two-dimensional skeleton estimation model 251 and a joint angle estimation model 252.

[0095] Figure 5 An example illustrating the relationship between the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252.

[0096] like Figure 5 As shown, the 2D skeleton estimation model 251 takes a frame image of the robot 10 as input and outputs a 2D pose representing the pixel coordinates of the center positions of the joint axes J1 to J6 of the robot 10 in the frame image. On the other hand, the joint angle estimation model 252 takes the 2D pose output from the 2D skeleton estimation model 251, as well as the distance and tilt between the camera 22 and the robot 10, as input and outputs the angles of the joint axes J1 to J6 of the robot 10.

[0097] Then, the learning unit 301 provides the learned model, consisting of the constructed two-dimensional skeleton estimation model 251 and joint angle estimation model 252, to the terminal device 20.

[0098] The following section explains the construction of the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252.

[0099] <Two-dimensional skeleton estimation model 251>

[0100] The learning unit 301, for example, uses a deep learning model used in known tagless animal tracking tools (such as DeepLabCut) and trains it by performing machine learning based on input data of frame images of the robot 10 received from the terminal device 20 and training data consisting of labels representing the two-dimensional poses of the center positions of joint axes J1 to J6 when the frame image is captured. The two-dimensional skeleton estimation model 251 takes frame images of the robot 10 captured by the camera 22 of the terminal device 20 as input and outputs two-dimensional poses representing the pixel coordinates of the center positions of joint axes J1 to J6 of the robot 10 in the captured frame images.

[0101] Specifically, the two-dimensional skeleton estimation model 251 is constructed based on a convolutional neural network (CNN) as a neural network.

[0102] Convolutional neural networks are structures that consist of convolutional layers, pooling layers, fully connected layers, and output layers.

[0103] In convolutional layers, filters with predetermined parameters are applied to the input frame image to perform feature extraction, such as edge detection. These predetermined parameters are equivalent to the weights of the neural network and are learned through repeated forward and backward propagation.

[0104] In the pooling layer, the image output from the convolutional layer is blurred to allow for positional shifts in robot 10. Thus, even if the position of robot 10 changes, it can still be considered the same object.

[0105] By combining these convolutional and pooling layers, feature quantities can be extracted from frame images.

[0106] In the fully connected layer, the image data whose features have been extracted through the convolutional and pooling layers are combined into a node, and the output is the value transformed by the activation function, which is the feature map of confidence.

[0107] Figure 6 An example of a feature diagram showing the joint axes J1 to J6 of robot 10.

[0108] like Figure 6 As shown, in the feature diagrams of each joint axis J1 to J6, the confidence level c iThe value is represented by a range of 0 to 1. The closer the cell is to the center of the joint axis, the closer it is to the value of "1". As the cell moves away from the center of the joint axis, the closer it is to the value of "0".

[0109] In the output layer, the row, column, and maximum confidence of the cell with the highest confidence in the feature map of each joint axis J1 to J6 are output from the fully connected layer. Furthermore, when the frame image is convolved by 1 / N in the convolutional layer, the row and column of the cell in the output layer are multiplied by N to represent the pixel coordinates (N is an integer greater than or equal to 1) of the center position of each joint axis J1 to J6 in the frame image.

[0110] Figure 7 This is an example of comparing a frame image with the output of a two-dimensional skeleton estimation model 251.

[0111] <Joint Angle Estimation Model 252>

[0112] The learning unit 301 performs machine learning based on training data consisting of input data and label data to generate a joint angle estimation model 252. The distance and tilt between the camera 22 and the robot 10, as well as the standardized two-dimensional pose representing the center position of joint axes J1 to J6, are used as input data. The label data is the angle of joint axes J1 to J6 of the robot 10 when the frame image is captured.

[0113] Furthermore, the learning unit 301 standardizes the two-dimensional poses of joint axes J1 to J6 output from the two-dimensional skeleton estimation model 251, but it can also generate a two-dimensional skeleton estimation model 251 so that the standardized two-dimensional poses are output from the two-dimensional skeleton estimation model 251.

[0114] Figure 8 This represents an example of joint angle estimation model 252. Here, as... Figure 8 As shown, regarding the joint angle estimation model 252, a multi-layer neural network is illustrated. This multi-layer neural network takes the standardized two-dimensional pose representing the center positions of joint axes J1 to J6, output from the two-dimensional skeleton estimation model 251, as well as the distance and slope between the camera 22 and the robot 10, as input layers, and the angles of joint axes J1 to J6 as output layers. Furthermore, the two-dimensional pose is a standardized coordinate (x, y, y) containing the center positions of joint axes J1 to J6. i y i ) and confidence level c i of (x) i y i c i ).

[0115] In addition, the "X-axis tilt Rx", "Y-axis tilt Ry" and "Z-axis tilt Rz" are the rotation angles around the X-axis, Y-axis and Z-axis between the camera 22 and the robot 10 in the world coordinate system, calculated based on the three-dimensional coordinates of the camera 22 in the world coordinate system and the three-dimensional coordinates of the robot origin of the robot 10 in the world coordinate system.

[0116] In addition, after constructing a fully learned model consisting of a two-dimensional skeleton estimation model 251 and a joint angle estimation model 252, the learning unit 301 can further supervise the learning model consisting of the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252 when new training data is obtained, thereby updating the fully learned model consisting of the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252 constructed once.

[0117] Therefore, training data can be automatically obtained from the regular shooting of robot 10, thus improving the estimation accuracy of robot 10's two-dimensional posture and the angles of joint axes J1 to J6 on a daily basis.

[0118] The aforementioned supervised learning can be conducted through online learning, batch learning, or small-batch learning.

[0119] Online learning refers to a learning method where supervised learning is performed immediately whenever frame images of robot 10 are captured and training data is generated. Batch learning, on the other hand, involves collecting multiple training data sets corresponding to each repetition while repeatedly capturing frame images of robot 10 to generate training data, and then using all the collected training data for supervised learning. Mini-batch learning, then, refers to a learning method where supervised learning is performed whenever training data has accumulated to a level intermediate between online and batch learning.

[0120] The storage unit 302 is RAM (Random Access Memory) or the like, which stores input data and tag data obtained from the terminal device 20, as well as two-dimensional skeleton estimation model 251 and joint angle estimation model 252 constructed by the learning unit 301.

[0121] The above describes the machine learning used by the terminal device 20, which is used to generate the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252 when it performs actions as a robot joint angle estimation device.

[0122] Next, the terminal device 20, which acts as a robot joint angle estimation device during the application phase, will be described.

[0123] <System in the Application Phase>

[0124] Figure 9 This is a functional block diagram illustrating the functional structure of a system in one implementation phase during the application stage. For example... Figure 1 As shown, system 1 includes a robot 10 and a terminal device 20, which serves as a robot joint angle estimation device. Furthermore, for systems with... Figure 1 Elements in System 1 that have the same function are labeled with the same reference numerals, and detailed descriptions are omitted.

[0125] like Figure 1 As shown, the terminal device 20, which functions as a robot joint angle estimation device during the application phase, includes a control unit 21a, a camera 22, a communication unit 23, and a storage unit 24a. Furthermore, the control unit 21a includes a 3D object recognition unit 211, a self-position estimation unit 212, an input unit 220, and an estimation unit 221.

[0126] The camera 22 and communication unit 23 are the same as those in the learning phase.

[0127] The storage unit 24a may be, for example, a ROM (Read Only Memory) or an HDD (Hard Disk Drive), storing system programs executed by the control unit 21a (described later) and robot joint angle estimation applications. Alternatively, the storage unit 24a may also store the two-dimensional skeleton estimation model 251 and joint angle estimation model 252, as well as the three-dimensional recognition model data 243, which are learned models provided by the machine learning device 30 during the learning phase.

[0128] <Control Unit 21a>

[0129] The control unit 21a includes a CPU (Central Processing Unit), ROM, RAM, CMOS (Complementary Metal-Oxide-Semiconductor) memory, etc., which are configured to communicate with each other via a bus, as is known to those skilled in the art.

[0130] The CPU is the processor of the overall control terminal device 20. The CPU reads the system program and the robot joint angle estimation application stored in the ROM via the bus, and controls the entire terminal device 20, which serves as the robot joint angle estimation device, according to the system program and the robot joint angle estimation application. Thus, as... Figure 9 As shown, the control unit 21a is configured to perform the functions of the three-dimensional object recognition unit 211, the self-position estimation unit 212, the input unit 220, and the estimation unit 221.

[0131] The 3D object recognition unit 211 and the self-position estimation unit 212 are the same as those in the learning phase.

[0132] <Input Section 220>

[0133] The input unit 220 inputs frame images of the robot 10 captured by the camera 22, the distance L between the camera 22 and the robot 10 calculated by the self-position estimation unit 212, the tilt Rx of the X-axis, the tilt Ry of the Y-axis, and the tilt Rz of the Z-axis.

[0134] <Presumption Section 221>

[0135] The estimation unit 221 inputs the frame image of the robot 10, the distance L between the camera 22 and the robot 10, the tilt Rx of the X-axis, the tilt Ry of the Y-axis, and the tilt Rz of the Z-axis, which are input through the input unit 220, into the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252, which are the learned models. As a result, the estimation unit 221 can estimate the angles of the joint axes J1 to J6 of the robot 10 when the input frame image was captured, as well as the two-dimensional pose representing the center position of the joint axes J1 to J6, based on the output of the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252.

[0136] Furthermore, as described above, the estimation unit 221 normalizes the pixel coordinates of the center positions of joint axes J1 to J6 output from the two-dimensional skeleton estimation model 251 and inputs them into the joint angle estimation model 252. Additionally, the estimation unit 221 can also determine the certainty c of the two-dimensional pose output from the two-dimensional skeleton estimation model 251. i Set to "1" when the value is above 0.5, and set to "0" when the value is less than 0.5.

[0137] The terminal device 20 can display the estimated angles of the joint axes J1 to J6 of the robot 10 and the two-dimensional posture representing the center position of the joint axes J1 to J6 on a display unit (not shown) including a liquid crystal display included in the terminal device 20.

[0138] <Estimated processing of terminal device 20 during the application phase>

[0139] Next, the operations involved in the presumption processing of the terminal device 20 in this embodiment will be explained.

[0140] Figure 10 This is a flowchart illustrating the presumption processing of the terminal device 20 during the application phase. The process shown here is executed repeatedly each time a frame image of the robot 10 is input.

[0141] In step S1, the camera 22 takes pictures of the robot 10 based on the operator's instructions via an input device such as a touch panel (not shown) included in the terminal device 20.

[0142] In step S2, the 3D object recognition unit 211 obtains the 3D coordinates of the robot origin in the world coordinate system and information representing the directions of the X-axis, Y-axis and Z-axis of the robot coordinate system based on the frame images of the robot 10 captured in step S1 and the 3D recognition model data 243.

[0143] In step S3, the self-position estimation unit 212 obtains the three-dimensional coordinates of the camera 22 in the world coordinate system based on the frame images of the robot 10 captured in step S1.

[0144] In step S4, the self-position estimation unit 212 calculates the distance L between the camera 22 and the robot 10, the tilt Rx of the X-axis, the tilt Ry of the Y-axis, and the tilt Rz of the Z-axis based on the three-dimensional coordinate values ​​of the camera 22 obtained in step S3 and the three-dimensional coordinate values ​​of the robot origin of the robot 10 obtained in step S2.

[0145] In step S5, the input unit 220 inputs the frame image captured in step S1, the distance L between the camera 22 and the robot 10 calculated in step S3, the tilt Rx of the X-axis, the tilt Ry of the Y-axis, and the tilt Rz of the Z-axis.

[0146] In step S6, the estimation unit 221 inputs the frame image input in step S5, the distance L between the camera 22 and the robot 10, the tilt Rx of the X-axis, the tilt Ry of the Y-axis, and the tilt Rz of the Z-axis into the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252, which are the learned models, thereby estimating the angles of the joint axes J1 to J6 of the robot 10 when the input frame image was captured, and the two-dimensional pose representing the center position of the joint axes J1 to J6.

[0147] Based on the above, in one embodiment, the terminal device 20 can easily obtain the angles of each joint axis J1 to J6 of the robot 10 by inputting the frame images of the robot 10, as well as the distance and tilt between the camera 22 and the robot 10, into the two-dimensional skeleton estimation model 251 and the joint angle estimation model 252, which are learned models. Even if the robot 10 does not have a log function or a dedicated I / F installed, it can easily obtain the angles of each joint axis J1 to J6 of the robot 10.

[0148] The above describes one embodiment, but the terminal device 20 and the machine learning device 30 are not limited to the above embodiment, and include variations and improvements within the scope of achieving the purpose.

[0149] <Variation Example 1>

[0150] In the above embodiments, the machine learning device 30 is exemplified as a device different from the robot control device (not shown) and terminal device 20 of the robot 10, but the robot control device (not shown) or terminal device 20 may also have some or all of the functions of the machine learning device 30.

[0151] <Variation Example 2>

[0152] Additionally, for example, in the above embodiment, the terminal device 20, which acts as a robot joint angle estimation device, uses a two-dimensional skeleton estimation model 251 and a joint angle estimation model 252, which are learned models provided by the machine learning device 30, to estimate the angles of the robot 10's joint axes J1 to J6 and a two-dimensional pose representing the center position of the joint axes J1 to J6 based on the input frame images of the robot 10 and the distance and tilt between the camera 22 and the robot 10. However, it is not limited to this. For example, such as Figure 11 As shown, server 50 can store the two-dimensional skeleton estimation model 251 and joint angle estimation model 252 generated by machine learning device 30, and share the two-dimensional skeleton estimation model 251 and joint angle estimation model 252 with m terminal devices 20A(1) to 20A(m) (m is an integer of 2 or more) that are connected to server 50 via network 60 and act as robot joint angle estimation devices. Therefore, even if a new robot and terminal device are configured, the two-dimensional skeleton estimation model 251 and joint angle estimation model 252 can be applied.

[0153] Furthermore, robots 10A(1) to 10A(m) correspond to respectively Figure 9 Robot 10. Terminal devices 20A(1) to 20A(m) respectively correspond to Figure 9 Terminal device 20.

[0154] Furthermore, the functions included in the terminal device 20 and the machine learning device 30 in one embodiment can be implemented separately by hardware, software, or a combination thereof. Here, implementation by software means implementation by a computer reading and executing a program.

[0155] The various structural units included in the terminal device 20 and the machine learning device 30 can be implemented using hardware, software, or a combination thereof, including circuits. In the case of software implementation, the program constituting the software is installed on a computer. Alternatively, these programs can be recorded on removable media and distributed to users, or downloaded to a user's computer via a network. In the case of hardware implementation, for example, integrated circuits (ICs) such as ASICs (Application Specific Integrated Circuits), gate arrays, FPGAs (Field Programmable Gate Arrays), and CPLDs (Complex Programmable Logic Devices) can constitute part or all of the functionality of the various structural units included in the aforementioned devices.

[0156] Programs can be stored and provided to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media encompasses various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., floppy disks, magnetic tapes, hard disks), optical-magnetic recording media (e.g., optical discs), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash memory ROMs, and RAM). Alternatively, programs can also be provided to a computer using various types of transient computer-readable media. Examples of transient computer-readable media include electrical signals, optical signals, and electromagnetic waves. Transient computer-readable media can provide programs to a computer via wired communication paths such as wires and optical fibers, or via wireless communication paths.

[0157] Furthermore, the steps used to describe a program recorded in a recording medium certainly include processing performed sequentially in time, but not necessarily in time, and also include processing performed in parallel or individually.

[0158] In other words, the training data generation device, machine learning device, and robot joint angle estimation device disclosed herein can be implemented in various ways with the following structures.

[0159] (1) The training data generation apparatus of this disclosure generates training data for generating a learned model. The learned model takes into account a two-dimensional image of the robot 10 captured by the camera 22, as well as the distance and tilt between the camera 22 and the robot 10, and estimates the angles of multiple joint axes J1 to J6 contained in the robot 10 when the two-dimensional image was captured, and a two-dimensional pose representing the center position of the multiple joint axes J1 to J6 in the two-dimensional image. The training data generation apparatus includes: an input data acquisition unit 216, which acquires a two-dimensional image of the robot 10 captured by the camera, as well as the distance and tilt between the camera and the robot 10; and a label acquisition unit 217, which acquires the angles and two-dimensional poses of the multiple joint axes J1 to J6 when the two-dimensional image was captured as label data.

[0160] According to this training data generation device, even robots without log functions or dedicated I / F can generate training data that is most suitable for generating a learned model, wherein the learned model is used to easily obtain the angles of each joint axis of the robot.

[0161] (2) The machine learning apparatus 30 disclosed herein includes a learning unit 301, which performs supervised learning based on training data generated by the training data generation apparatus described in (1), thereby generating a learned model.

[0162] According to the machine learning device 30, even robots without logging functions or dedicated I / F can generate a fully learned model that is best suited for easily obtaining the angles of each joint axis of the robot.

[0163] (3) The machine learning device 30 described in (2) may include the training data generation device described in (1).

[0164] Therefore, the machine learning device 30 can easily obtain training data.

[0165] (4) The robot joint angle estimation device disclosed herein includes: a learned model generated by the machine learning device described in (2) or (3); an input unit 220, which inputs a two-dimensional image of the robot 10 captured by the camera 22, as well as the distance and tilt between the camera 22 and the robot 10; and an estimation unit 221, which inputs the two-dimensional image input by the input unit 220, as well as the distance and tilt between the camera 22 and the robot 10, into the learned model to estimate the angles of the multiple joint axes J1 to J6 contained in the robot 10 when the two-dimensional image was captured, and a two-dimensional pose representing the center position of the multiple joint axes J1 to J6 in the two-dimensional image.

[0166] Based on this robot joint angle estimation device, even robots without log functions or dedicated I / F can easily obtain the angles of each joint axis.

[0167] (5) In the robot joint angle estimation device described in (4), the learned model may include: a two-dimensional skeleton estimation model 251, which takes a two-dimensional image as input and outputs a two-dimensional pose; and a joint angle estimation model 252, which takes the two-dimensional pose output from the two-dimensional skeleton estimation model 251 as input and the distance and tilt between the camera 22 and the robot 10, and outputs the angles of multiple joint axes J1 to J6.

[0168] Therefore, even for robots without logging functionality or dedicated I / F, the robot joint angle estimation device can easily obtain the angles of each joint axis of the robot.

[0169] (6) In the robot joint angle estimation device described in (4) or (5), the learned model can be in the server 50 that can be accessed from the robot joint angle estimation device via the network 60.

[0170] Therefore, even with a new robot and robot joint angle estimation device configured, the robot joint angle estimation device can still apply the learned model.

[0171] (7) The robot joint angle estimation device described in any of (4) to (6) may include the machine learning device 30 described in (2) or (3).

[0172] Therefore, the robot joint angle estimation device can achieve the same effect as (1) to (6).

[0173] Explanation of reference numerals in the attached figures

[0174] 1 System

[0175] 10 robots

[0176] 101 Joint Angle Response Server

[0177] 20 Terminal devices

[0178] 21, 21a Control Section

[0179] 211 Three-dimensional object recognition unit

[0180] 212 Self-position estimation department

[0181] 213 Joint Angle Acquisition Section

[0182] 214 Forward Kinematics Calculation Department

[0183] 215 Projection Section

[0184] 216 Input Data Acquisition Department

[0185] 217 Tag Acquisition Department

[0186] 220 Input Section

[0187] 221 Presumption Department

[0188] 22 cameras

[0189] 23 Ministry of Communications

[0190] 24, 24a storage section

[0191] 241 Input Data

[0192] 242 tag data

[0193] 243 Three-dimensional recognition model data

[0194] 251 Two-dimensional skeleton estimation model

[0195] 252 Joint Angle Estimation Model

[0196] 30 machine learning devices

[0197] 301 Study Department

[0198] 302 Storage Section.

Claims

1. A training data generation apparatus, which generates training data for generating a learned model, wherein, The learned model includes a 2D skeleton estimation model and a joint angle estimation model. The 2D skeleton estimation model takes as input a 2D image of the robot captured by a camera and outputs a 2D pose representing the center positions of multiple joint axes of the robot at the time the 2D image was captured. The joint angle estimation model takes as input the 2D pose output from the 2D skeleton estimation model, the distance and tilt between the camera and the robot, and outputs the angles of the multiple joint axes. Its features are, The training data generation device includes: The input data acquisition unit acquires a two-dimensional image of the robot captured by the camera, as well as the distance and tilt between the camera and the robot; The tag acquisition unit acquires the angles of the multiple joint axes and the two-dimensional pose when the two-dimensional image was captured as tag data.

2. A machine learning device, characterized in that, The machine learning device includes a learning unit that performs supervised learning based on training data generated by the training data generation device of claim 1 to generate a learned model.

3. The machine learning apparatus according to claim 2, characterized in that, The machine learning device includes the training data generation device as described in claim 1.

4. A robot joint angle estimation device, characterized in that, have: The learned model generated by the machine learning apparatus of claim 2 or 3; The input unit receives a two-dimensional image of the robot captured by a camera, as well as the distance and tilt between the camera and the robot. The estimation unit inputs the two-dimensional image from the input unit, along with the distance and tilt between the camera and the robot, into the learned model to estimate the angles of multiple joint axes in the robot when the two-dimensional image was captured, and the two-dimensional pose representing the center position of the multiple joint axes in the two-dimensional image.

5. The robot joint angle estimation device according to claim 4, characterized in that, The learned model is contained in a server that can be accessed via a network from the robot joint angle estimation device.

6. The robot joint angle estimation device according to claim 4, characterized in that, The robot joint angle estimation device includes the machine learning device described in claim 2 or 3.