Operation command generation device

The operation command generation device addresses the issue of camera viewpoint deviation in autonomous learning type robot control systems by using machine learning to generate command models from randomly selected single-viewpoint data, ensuring efficient operation command generation without increased learning time.

JP2025093156APending Publication Date: 2025-06-23HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023208726
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-11
Publication Date
2025-06-23

AI Technical Summary

Technical Problem

Conventional autonomous learning type robot control systems face significant performance deterioration when the camera viewpoint deviates from the learned position, leading to increased learning time and inefficiency.

Method used

An operation command generation device that randomly extracts captured images and sensor data from a single viewpoint from a dataset including images from multiple viewpoints, performs machine learning to generate a command generation model, and uses this model to predict and generate operation commands for the robot without increasing learning time.

Benefits of technology

Enables the robot to perform desired operations even with shifted camera viewpoints without extending learning time, maintaining efficiency and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025093156000001_ABST
    Figure 2025093156000001_ABST
Patent Text Reader

Abstract

To generate an operation command so that a robot can carry out desired work without increasing learning time even if the imaging viewpoint of a work environment shifts.SOLUTION: An operation command generation device 30 comprises: a learning unit 312 that generates a command generation model 322 for randomly extracting a captured image of a single viewpoint and center data of the same time as the imaging time of the captured image from learning data 321 containing the center data in which the captured image of a work environment from a plurality of viewpoints and command data to a robot 10 are reflected to execute machine learning and autonomously operating the robot; and a command generation unit 313 that inputs the captured image and the center data at the current time to the command generation model 322 and generates an operation command of next time for the robot.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an operation command generation device.

Background Art

[0002] In order to improve production efficiency and reduce labor costs, there has been an increasing effort to replace tasks performed by humans, such as the assembly, welding, and transportation of industrial products, with robots. However, conventional robot systems have required enormous programming and high expertise, which has been a hindrance to the introduction of robots. Therefore, an autonomous learning type robot control system has been proposed that determines its own operation based on various sensor information attached to the robot device. Compared with conventional robot control systems, the autonomous learning type robot control system is expected to be able to easily introduce robots without the need for enormous programming and high expertise. Furthermore, the autonomous learning type robot control system is expected to be able to generate flexible operations in response to various environmental changes by memorizing and learning its own operation experience.

[0003] Examples of the operation experience of a robot include a method in which a user who operates the robot directly teaches and memorizes an operation to the robot, and a method of imitating the operations of a person or another robot. Also, generally, an autonomous learning type robot control system is equipped with a learning device, and stores sensor information during operation experience and adjusts parameters for generating operations. This stored sensor information during operation experience becomes learning data, and the adjustment of parameters for generating operations becomes learning. The learning device repeatedly performs learning using the learning data so that an expected output value is output for the input value (learning data) to the learning device.

[0004] The learning device stores, as learning data, for example, the joint angle information of the robot during a certain operation experience and the captured images captured in time series of the working environment. The learning device uses the stored learning data to perform machine learning so as to input the joint angle information and the captured image at time (t) and predict the joint angle information and the captured image at time (t+1). Then, the autonomous learning type robot control system can automatically generate an operation command according to its own state by sequentially inputting the joint angle information and the captured image of the robot to the learning device in which the learning has been completed.

[0005] A method of directly generating an operation command to a robot without going through object recognition or the like from the sensor information at a certain time as described above is called an autonomous learning type end-to-end operation generation method. As an autonomous learning type end-to-end operation generation technology, for example, the robot system described in Patent Document 1 is disclosed.

[0006] Patent Document 1 describes that a machine learning device recognizes a person's face during a period when a person and a robot cooperate to perform work, classifies the person's actions based on the weights of the neural network corresponding to the person, and learns the person's actions, and controls the robot's actions based on the result of classifying the person's actions.

Prior Art Documents

Patent Documents

[0007]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] As described above, conventionally, a technique (Patent Document 1) has been disclosed in which the operation of a robot is automatically controlled by machine learning using learning data such as captured images of the working environment. In the self-learning type end-to-end operation generation method described in Patent Document 1, basically, the viewpoint of the camera is fixed, and no deviation of the viewpoint is assumed. However, in practical use, there are many scenes in which the viewpoint of the camera deviates. In a mobile manipulator in which the robot itself moves, as the robot moves, the camera moves relatively with respect to the environment. Also, in a stationary robot, since the camera is also relocated when the robot is relocated, it is expected that the viewpoint of the camera will deviate from the viewpoint at the time of learning. When the viewpoint of the camera deviates, the position of the work object and the robot in the captured image also deviates. Since such a viewpoint deviation has not been learned at all, the performance of the self-learning type robot control system deteriorates significantly. In such a case, it is necessary to perform additional learning using the learning data obtained by shifting the viewpoint of the camera. Also, along with the additional learning, a problem of an increase in learning time occurs.

[0009] The present invention has been made to solve the above problems, and an object of the present invention is to generate an operation command so that the robot performs a desired operation without increasing the learning time even when the imaging viewpoint of the working environment deviates.

Means for Solving the Problems

[0010] To solve the above problems, an operation command generation device of the present invention randomly extracts a captured image of a single viewpoint and sensor data at the same time as the imaging time of the captured image from learning data including captured images of the working environment from a plurality of viewpoints and sensor data reflecting command data to the robot, performs machine learning, and generates a command generation model for autonomously operating the robot, and an imaging image and sensor data at the current time are input to the command generation model, and a command generation unit that generates an operation command for the next time for the robot.

Effects of the Invention

[0011] According to the present invention having the above configuration, even if the imaging viewpoint of the working environment is shifted, an operation command can be generated so that the robot performs a desired operation without increasing the learning time. Problems, configurations, and effects other than those described above will be clarified by the description of the following embodiments.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0013] Hereinafter, embodiments for carrying out the present invention will be described with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same function or configuration are denoted by the same reference numerals, and redundant description is omitted.

[0014] <First Embodiment> [Overall Configuration Example of Robot Control System] First, the overall configuration of the robot control system according to the first embodiment of the present invention will be described. FIG. 1 is a diagram showing an example of the overall configuration of the robot control system 1 according to this embodiment. As shown in FIG. 1, the robot control system 1 includes a robot 10, a camera 20 (imaging device), a sensor 21, an operation command generation device 30, and a control device 40. The camera 20, the operation command generation device 30, the control device 40, and the robot 10 are connected in this order. The sensor 21 is disposed on the arm of the robot 10.

[0015] The robot 10 can perform operations such as handling and moving a work object 50, assembling parts, acquiring images, and transporting. Regardless of its configuration, the robot 10 may be a single robot arm having an end effector 11 such as a gripper, a moving device having crawlers or wheels, or a combination of both. In this embodiment, it is assumed that the robot 10 is a single robot arm having an end effector 11.

[0016] The camera 20 is an example of an imaging device that images the working environment. The camera 20 may be attached to the robot 10, or may be attached to the working environment side so as to overlook the robot 10 as shown in FIG. 1. Also, in the working environment, instead of arranging a single camera 20, a plurality of cameras 20 may be arranged at different positions and postures.

[0017] The sensor 21 is a device that measures the state of the robot 10, the working environment, etc. The sensor 21 is, for example, a sensor that measures the current value of a motor provided in a joint of the robot 10, a tactile sensor externally attached to the robot 10, a force sensor, an inertial sensor, etc. Further, the sensor 21 may be a temperature sensor that measures the temperature of the working environment. Also, the sensor 21 may be not single but plural. The sensor 21 detects the state of the robot 10, the working environment, etc., and outputs a detection signal (sensor data) corresponding to the detection content to the operation command generation device 30 via the control device 40. In the present embodiment, it is assumed that the sensor 21 is a sensor that detects the joint angle of the robot 10, the posture (position) of the end effector 11, the force (torque), etc. Therefore, in the present embodiment, the sensor data measured by the sensor 21 reflects command data specifying the joint angle of the robot 10, the posture (position) of the end effector 11, the force (torque), etc.

[0018] The operation command generation device 30 is a terminal device that generates end-to-end operation commands, and has a function of learning a command generation model for autonomously operating the robot 10 using the output information of the sensor 21 and the camera 20. The operation command generation device 30 has a learning mode and an autonomous operation mode. In the learning mode, the operation command generation device 30 inputs the sensor data measured by the sensor 21 and the captured image of the working environment captured by the camera 20 as learning data. Then, the operation command generation device 30 performs machine learning using the learning data to generate a command generation model 322, which will be described later, for autonomously operating the robot 10.

[0019] On the one hand, in the autonomous operation mode, the operation command generation device 30 inputs the sensor data and captured images at the current time into the learned command generation model 322, and the command generation model 322 predicts the sensor data and captured images at the next time. Then, the operation command generation device 30 generates an operation command for autonomously operating the robot 10 based on the sensor data and captured images at the next time and outputs it to the control device 40. Here, the operation command is, for example, command data specifying the joint angles of the robot 10, the posture (position) of the end effector 11, the force (torque), etc. Also, the above-mentioned sensor data reflects the above-mentioned command data when the robot 10 operates. The above-mentioned captured image is an image captured in time series when the robot 10 performs a predetermined operation, and is an image captured so that the end effector 11 of the robot 10 and the work object 50 of the robot 10 are reflected.

[0020] The control device 40 outputs a control command to the robot 10 based on the operation command input from the operation command generation device 30 to control the operation of the robot 10. The control command is, for example, a signal indicating a current value, a voltage value, etc. for an actuator (such as a motor) provided in a joint of a robot arm or an end effector 11 provided in the robot 10. When the robot 10 receives the control command from the control device 40, the built-in drive circuit supplies a drive signal to the corresponding actuator.

[0021] [Functional Configuration Example of Operation Command Generation Device] Next, the functional configuration of the operation command generation device 30 will be described. FIG. 2 is a block diagram showing a functional configuration example of the operation command generation device 30 according to the present embodiment. As shown in FIG. 2, the operation command generation device 30 includes a control unit 31, a storage unit 32, and an input / output unit 33.

[0022] The control unit 31 is composed of, for example, a CPU (Central Processing Unit) not shown in the figure. As shown in FIG. 2, it includes an acquisition unit 311, a learning unit 312, and a command generation unit 313. The control unit 31 controls the operations of the components of the operation command generation device 30. Note that the control unit 31 may be composed of a GPU (Graphics Processing Unit) instead of a CPU, or may be composed of a CPU and a GPU.

[0023] The acquisition unit 311 acquires various instruction information input by the operator from the input / output unit 33. In addition, the acquisition unit 311 acquires the learning data 321 (described later) required for machine learning in the learning unit 312 from the storage unit 32. Further, the acquisition unit 311 acquires sensor data from the sensor 21 and acquires a captured image from the camera 20.

[0024] The learning unit 312 randomly extracts a captured image from a single viewpoint and sensor data at the same time as the capture time of the captured image from the learning data 321 (described later) stored in the storage unit 32, performs machine learning, and generates a command generation model 322 (described later) for autonomously operating the robot 10 that executes a predetermined operation.

[0025] The command generation unit 313 inputs a captured image of the working environment at the current time and sensor data in which command data for the robot 10 is reflected into the command generation model 322 learned by the learning unit 312, and generates an operation command for the next time for the robot 10.

[0026] The storage unit 32 is configured to include storage devices such as a ROM (Read Only Memory), a RAM (Random Access Memory), and an SSD (Solid State Drive). The storage unit 32 stores learning data 321, a command generation model 322, and a program 323.

[0027] The program 323 includes a description of the learning method of the command generation model 322 (see FIG. 3 described later). The command generation model 322 is a deep learning model that inputs the captured image of the working environment at the current time and the sensor data reflecting the command data to the robot 10, and predicts various parameters for controlling the working behavior of the robot 10, as well as the captured image and sensor data at the next time. The command generation model 322 is generated for each operation of the robot 10. When there are multiple operations, the command generation model 322 corresponding to each operation is generated and stored in the storage unit 32. Note that the present invention is not limited to generating only one command generation model 322 for each operation of the robot 10. For example, even if the robot 10 performs the same operation, if there is a change in the attached sensor, the command generation model also changes. Therefore, for the same operation, a plurality of command generation models 322 may be generated and stored.

[0028] The learning data 321 is learning data used for machine learning of the command generation model 322. The learning data 321 includes the captured images of the working environment from multiple viewpoints and the sensor data reflecting the command data to the robot 10, which are acquired in time series when the robot 10 executes a predetermined operation (see FIG. 3). In the present embodiment, it is assumed that the captured images from multiple viewpoints included in the learning data 321 are images when the robot 10 executes a predetermined operation, which are captured by a plurality of imaging devices (cameras 20) arranged at different positions and postures in the working environment.

[0029] The position and posture of the camera 20 are determined based on the upper limit position, lower limit position, and intermediate position of the deviation of the camera 20 assumed within the range where the camera 20 can acquire an image with respect to the work target. In machine learning, in many cases, the robustness with respect to the interpolation range of the learning data is confirmed. Therefore, it is preferable to acquire the learning data by changing the position and posture of the camera 20 so as to include the assumed upper limit, lower limit, and intermediate position postures of the deviation of the camera 20.

[0030] Also, when the robot 10 is working, it is the relative positional relationship rather than the absolute positional relationship between the robot 10 and the work object 50 that is important. For this reason, in this embodiment, the captured image is an image captured so that the end effector 11 of the robot 10 and the work object 50 are reflected. Therefore, it is assumed that the position and orientation of the camera 20 are set so that at least the end effector 11 of the robot 10 and the work object 50 are within the field of view. By including the end effector 11 and the work object 50 within the field of view, the relative positional relationship between the two can be grasped from the captured image of the camera 20.

[0031] Note that when acquiring the learning data 321, since the operation command generation device 30 is in the learning mode, the command generation function of the command generation unit 313 is in an invalid state. At this time, the robot 10 operates by the operation of an operator or automatically runs according to a pre-registered operation plan.

[0032] When the robot 10 operates by the operation of an operator, the control device 40 connected to an operation unit such as a joystick of the robot 10 transmits time-series command data for the robot 10 corresponding to the operation content of the operator to the operation command generation device 30. Also, when the robot 10 automatically runs according to a pre-registered operation plan, the control device 40 transmits time-series command data for the robot 10 registered in advance to the operation command generation device 30.

[0033] User interface devices such as a display, a keyboard, and a mouse are connected to the input / output unit 33 to input various instruction information input by the user. Also, the input / output unit 33 includes a communication device and performs data transmission and reception with the control device 40, the camera 20, and the sensor 21. The input / output unit 33 displays captured images and sensor data from a plurality of viewpoints when acquiring the learning data 321 (see FIG. 4A described later). Also, the input / output unit 33 displays the learning data and the progress of learning during the learning of the command generation model 322 (see FIG. 4B described later). Also, the input / output unit 33 displays the captured image and sensor data input when generating an operation command, and the predicted captured image and sensor data (see FIG. 4C described later).

[0034] [Learning Method by Instruction Generation Model] Next, the operation instruction generation process in the operation instruction generation device 30 will be described. FIG. 3 is a block diagram showing a learning method of the instruction generation model 322 in the operation instruction generation device 30 according to the present embodiment. Note that the learning method shown in FIG. 3 is executed by the learning unit 312 calling and executing the program 323 stored in the storage unit 32.

[0035] In the present embodiment, during learning, the position and orientation of the camera 20 are changed so as to include the upper limit, lower limit, and intermediate position and orientation of the deviation of the camera 20 assumed in the working environment, and the robot 10 is made to perform a predetermined operation a plurality of times (for example, K times). Each time the predetermined operation is performed, the acquisition unit 311 of the operation instruction generation device 30 obtains learning data 321. The learning data 321 includes data_1 to data_K, which are learning data obtained each time the predetermined operation is performed K (K is an integer greater than 1) times. Each of data_1 to data_K includes captured images of viewpoints_1 to viewpoints_N (N is an integer greater than 1) and sensor data in which command data to the robot 10 at the imaging time is reflected.

[0036] The learning unit 312 of the operation instruction generation device 30 first randomly extracts a captured image of a single viewpoint from the learning data 321 (data_1 to data_K) at the current time (t) or a plurality of times up to the current time (t). That is, the learning unit 312 performs the random extraction process for data_1 to data_K shown in FIG. 3. Further, the learning unit 312 extracts the sensor data at the same time as the imaging time of the extracted captured image. Next, the learning unit 312 inputs the extracted captured image of a single viewpoint and the sensor data into the instruction generation model 322, and uses the instruction generation model 322 to predict the captured image and sensor data of a single viewpoint at the next time (t + 1) (the same viewpoint as the imaging viewpoint of the input captured image). Note that the prediction of the captured image is not essential, but it is preferably included as a subtask for preventing overfitting.

[0037] Note that the deep learning model calculates the error between the output value and the correct value, and repeats the update of various parameters for controlling the operation behavior of the robot 10 using the error backpropagation method or the like to perform learning. In the present invention, the error used for updating the parameters of the command generation model 322 is calculated based on the predicted captured image and the predicted sensor data output by the command generation model 322, and the captured image and the sensor data which are the correct values at the same time. The captured image and the sensor data which are the correct values are, that is, the captured image and the sensor data included in the learning data 321 at the next time (t + 1).

[0038] In the present invention, randomly selecting a single-viewpoint captured image from the captured images of multiple viewpoints is for robustly predicting an operation command against changes in viewpoints. Further, although the learning data 321 includes captured images of multiple viewpoints, the number of the learning data 321 input to the command generation model 322 remains K regardless of the number of viewpoints. For this reason, the learning unit 312 of the present invention can perform learning in the same learning time as in the case of learning from a single viewpoint. Note that, as the error between the output value calculated by the learning unit 312 and the correct value, the mean squared error, the mean absolute error, or the like can be used. In particular, for the error of the captured image, for example, the structural similarity SSIM (Structural Similarity) for measuring the image quality can be used.

[0039] [Display example of operation command generation screen] Next, a display example of the operation command generation screen in the operation command generation device 30 will be described. FIG. 4 is a diagram showing a display example of each screen for generating an operation command in the operation command generation device 30 according to the present embodiment. FIG. 4A is a diagram showing an example of the display screen at the time of acquiring learning data. FIG. 4B is a diagram showing an example of the display screen during learning. FIG. 4C is a diagram showing an example of the display screen at the time of generating an operation command.

[0040] As shown in FIG. 4A, at the upper end of the screen SC1, which is the display screen at the time of learning data acquisition, a screen name "Data Acquisition" is displayed. "Learning", which is the name of the display screen during learning, and "Operation Generation", which is the name of the display screen at the time of operation command generation, are in a grayed-out state. Below the screen name, a display section P11 for real-time sensor information, a display section P12 for logging setting information, and a display section P13 for robot operation settings are displayed in order from top to bottom.

[0041] In the display section P11 for real-time sensor information, a display section for the captured image of the camera and a display section for sensor data are displayed in order from left to right. In the display section for the captured image, the captured images of viewpoints_1 to viewpoints_N are displayed. Also, in the display section for sensor data, a chart showing the time change characteristics of the measurement data of the sensor 21 is displayed. Note that the display mode of the sensor data is not limited to a chart and may be a table or the like.

[0042] In the display section P12 for logging setting information, as shown in FIG. 4A, various settings such as the storage time, sampling period, and acquired data type related to logging can be performed. Note that the settings in the display section P12 are not limited to the above-mentioned various settings and may include other settings according to the types of sensors installed in the robot 10. Also, as input methods for various settings, for example, direct input into a text box, selection from a selection list, etc. can be used.

[0043] On the display unit P13 for robot operation settings, each button for setting the operation method of the robot 10, for example, the "remote operation" button, the "direct teaching" button, the "reference trajectory_1" button, and the "reference trajectory_2" button, are displayed in order from the left when acquiring learning data. When any of the above buttons is pressed, the operation method corresponding to the pressed button is set for the robot 10. "Remote operation" corresponds to an operation method in which an operator uses a terminal device on the remote side to instruct designated data for operating the robot 10 to the control device 40 connected to the terminal device. "Direct teaching" corresponds to a method in which an operator directly operates the robot 10. "Reference trajectory_1" and "reference trajectory_2" correspond to methods in which the robot 10 automatically runs according to the corresponding operation plans of the registered reference trajectory_1 and reference trajectory_2 respectively.

[0044] In the display screen SC1 at the time of acquiring learning data shown in FIG. 4A, an operator can set the operation method of the robot 10 at the time of acquiring learning data, and can also check the acquired learning data and logging information.

[0045] As shown in FIG. 4B, at the upper end of the screen SC2 which is the display screen during learning, the screen name "Learning" is displayed. "Data acquisition", which is the name of the display screen at the time of data acquisition, and "Motion generation", which is the name of the display screen at the time of motion command generation, are in a grayed-out state. Below the screen name, a display section P21 for learning data and a display section P22 for the progress of learning are displayed in order from the top.

[0046] In the display section P21 for learning data, a display section for the captured image of the camera 20 and a display section for sensor data are displayed in order from the left. In the display section for the captured image, the captured images of viewpoints_1 to viewpoints_N included in the learning data from time 0 to time t are displayed. Also, in the display section for sensor data, a chart showing the time change characteristics of the measurement data of the sensor 21 is displayed. The display modes of the captured image and the sensor data displayed in the display section P21 for learning data are the same as the display modes of the captured image of the camera and the sensor data shown in FIG. 4A respectively.

[0047] On the learning progress display unit P22, as shown in FIG. 4B, information indicating the learning progress such as the total learning time and the remaining learning time is displayed in, for example, time units.

[0048] On the display screen SC2 shown in FIG. 4B, the operator can also check the learning data used during learning and the learning progress.

[0049] As shown in FIG. 4C, at the upper end of the screen SC3, which is the display screen when generating an operation command, a screen name "Operation Generation" is displayed. "Data Acquisition", which is the name of the display screen at the time of data acquisition, and "Learning", which is the name of the display screen at the time of learning, are in a grayed-out state. Below the screen name, a display unit P31 for the input value and predicted value of the model is displayed. At the lower left of the display unit P31, a setting unit P32 for generating an operation command is displayed. At the lower right of the display unit P31 for the input value and predicted value of the model, buttons for "Start Operation", "Temporary Stop", and "Abort Operation" are displayed in order from the top.

[0050] In the display unit P31 for the input value and predicted value of the model, a display unit for the captured image of the camera 20 and a display unit for sensor data are displayed in order from the left. In the display unit for the captured image, the captured image of the current time of the camera 20, which is the current value, and the captured image of the next time, which is the predicted value predicted by the command generation model 322, are displayed. In the display unit for sensor data, a chart showing the time change characteristics of the measured value (solid line) of the sensor at the current time and the predicted value (dashed line) of the sensor data at the next time predicted by the command generation model 322 is displayed.

[0051] In the operation command generation setting unit P32, as shown in FIG. 4C, a command generation model used for command generation and various parameters corresponding to the command generation model can be set. Further, as input methods for various settings, for example, direct input to a text box, selection from a selection list, etc. can be used. After various settings are made in the operation command generation setting unit P32, when the "Start Operation" button is pressed, the robot 10 starts autonomous operation. Also, during the autonomous operation of the robot 10, when the "Temporary Stop" button is pressed, the autonomous operation of the robot 10 is temporarily stopped. Also, during the autonomous operation of the robot 10, when the "Abort Operation" button is pressed, the autonomous operation of the robot 10 is aborted.

[0052] On the display screen SC3 when shown in FIG. 4C, the operator can also confirm the learning data input during operation command generation and the predicted data for the next time. Also, on the display screen SC3 during operation command generation shown in FIG. 4C, the operator can specify the command generation model and the parameters corresponding to the command generation model, and can also start, temporarily stop, or abort the autonomous operation of the robot 10.

[0053] [Effect] As described above, the operation command generation device 30 according to the present embodiment acquires learning data by changing the position and orientation of the camera 20 so as to include captured images from a plurality of viewpoints. The position and orientation of the camera 20 are determined based on the upper limit position, lower limit position, and intermediate position of the assumed deviation of the camera 20 within the range where the camera 20 can acquire an image with respect to the work object 50. Also, the operation command generation device 30 extracts a randomly single-viewpoint captured image from the captured images from a plurality of viewpoints that have been acquired. The operation command generation device 30 inputs the captured image and sensor data of the single viewpoint at the current time that have been extracted into the command generation model 322, and uses the command generation model 322 to predict the captured image and sensor data of the single viewpoint at the next time. Therefore, according to the operation command generation device 30 according to the present embodiment, even if the imaging viewpoint of the working environment is shifted, it is possible to generate an operation command so that the robot performs a desired operation without increasing the learning time.

[0054] <Second Embodiment> Next, the operation command generation device 60 according to the second embodiment of the present invention will be described. FIG. 5 is a block diagram showing a functional configuration example of the operation command generation device 60 according to this embodiment. As can be seen by comparing FIG. 5 with FIG. 2, each component of the operation command generation device 60 other than the control unit 61, the learning data 621, and the command generation model 622 is the same as each component of the operation command generation device 30 shown in FIG. 2. Here, duplicate descriptions of the components that are the same as those of the operation command generation device 30 shown in FIG. 2 in the operation command generation device 60 are omitted.

[0055] The learning data 621 includes sensor data, captured images from a plurality of viewpoints, and further viewpoint information of a plurality of viewpoints, which are acquired in time series when the robot 10 executes a predetermined operation (see FIG. 6). The viewpoint information of a plurality of viewpoints is the position and orientation information of the imaging device at the time of imaging of each captured image from a plurality of viewpoints. The reason for including the viewpoint information in the learning data 621 is to cope with a situation where the end effector 11 of the robot 10 and the work object 50 (see FIG. 1) do not fit within the field of view of the camera 20. When the end effector 11 of the robot 10 and the work object 50 do not fit within the field of view of the camera 20, captured images can be generated by a simulator using the viewpoint information of the camera 20. Further, by image processing using data augmentation, an image obtained by performing image processing including at least one of rotation, translation, and scaling on a captured image of a single viewpoint of the work environment may be included in the captured images from a plurality of viewpoints. Using the above-described various image processing methods, captured images from a plurality of viewpoints or new viewpoints can be generated (estimated).

[0056] The command generation model 622 is a deep learning model that inputs sensor data, captured images, and viewpoint information at the current time and predicts various parameters for controlling the working behavior of the robot 10, sensor data, and captured images at the next time.

[0057] As shown in FIG. 5, the control unit 61 includes an acquisition unit 611, a learning unit 612, and a command generation unit 313. Since the command generation unit 313 is the same as the command generation unit 313 described in FIG. 2, duplicate descriptions will be omitted. The acquisition unit 611 has the same function as the acquisition unit 311 shown in FIG. 2, and further acquires the viewpoint information of the camera 20 at the time of imaging the captured image.

[0058] The learning unit 612 randomly extracts a captured image of a single viewpoint, the position and orientation information of the imaging device, and sensor data at the same time as the imaging time of the captured image from the learning data 621, performs machine learning, and generates a command generation model 622.

[0059] Next, with reference to FIG. 6, a learning method of the command generation model 622 will be described. FIG. 6 is a block diagram showing a learning method of the command generation model in the operation command generation device 60 according to the present embodiment.

[0060] As can be seen by comparing FIG. 6 with FIG. 3, the learning method of the command generation model 622 in the learning unit 612 inputs the learning data 621 including the viewpoint information, and uses the command generation model 622 to predict the captured image, sensor data, and viewpoint information at the next time. Note that the random extraction process and error calculation in the learning method of the command generation model 622 are the same as the random extraction process and error calculation in the learning method of the command generation model 322 in the learning unit 312 according to the first embodiment, so duplicate descriptions will be omitted.

[0061] [Effect] As described above, the operation command generation device 60 according to the present embodiment extracts a random single-viewpoint captured image and viewpoint information from captured images from a plurality of viewpoints and viewpoint information (position and orientation information of the camera 20) of the plurality of viewpoints. Using the viewpoint information of the camera 20, a captured image can be generated by a simulator. Also, by image processing using data augmentation, the captured image generated by the simulator can be rotated, translated, enlarged or reduced, etc. to generate (estimate) captured images of a plurality of viewpoints or new viewpoints. Therefore, the operation command generation device 60 according to the present embodiment has the same effects as the operation command generation device 30 according to the first embodiment, and can generate an operation command so that the robot performs a desired operation even when the end effector of the robot and the work object do not fit within the angle of view of the camera.

[0062] Note that the present invention is not limited to the above-described embodiments, and it goes without saying that various other application examples and modification examples can be taken as long as they do not depart from the gist of the present invention described in the claims. For example, each of the above-described embodiments has described the configuration of the robot control system in detail and specifically for the purpose of explaining the present invention clearly, and is not necessarily limited to those having all the configurations described. Also, it is possible to replace a part of the configuration of the embodiment described here with the configuration of another embodiment, and further, it is possible to add the configuration of another embodiment to the configuration of a certain embodiment. Also, it is possible to add, delete, or replace a part of the configuration of each embodiment with another configuration. Also, the control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines on the product. In fact, it may be considered that almost all the components are interconnected.

[0063] In each of the above embodiments, when acquiring learning data, an example was described in which a plurality of cameras were used to image the working environment when a certain operation was performed. However, the present invention is not limited to this. When it is difficult to arrange a plurality of cameras due to restrictions on the camera placement space or the like, a working environment similar to the actual environment may be constructed on a simulator, and imaging by a plurality of imaging devices arranged in different positions and postures may be simulated. That is, the captured images from a plurality of viewpoints included in the learning data may be images generated by simulating imaging by a plurality of imaging devices arranged in different positions and postures in an environment similar to the working environment constructed on the simulator, when the robot performs a predetermined operation. Since it is on the simulator, it is possible to easily acquire captured images from a plurality of viewpoints without preparing a plurality of actual cameras.

[0064] Also, in the above modification, an example was described in which a working environment similar to the actual environment was constructed on a simulator. However, the present invention is not limited to this. In some cases, it may be difficult to construct a working environment on a simulator, such as when handling fluids such as liquids, flexible objects such as cables and cloths as work objects. In such a case, the captured images from a plurality of viewpoints included in the learning data may be captured images generated by a novel viewpoint image generation method using a captured image of a single viewpoint of the working environment. As a novel viewpoint image generation method, for example, image generation techniques such as NeRF (Neural Radiance Fields) and 3D Gaussian Splatting can be used.

[0065] Also, in each of the above embodiments, an example was described in which a plurality of viewpoints of captured images were pseudo-generated by applying image processing such as rotation, translation, and scaling to the captured image obtained from a single viewpoint of the working environment. The present invention is not limited to this. For example, in the captured image obtained from a single viewpoint, for an area that is not visible, such as the side of the work object, image processing such as rotation, translation, and scaling may be used assistively to generate a captured image of the invisible area.

Explanation of Reference Numerals

[0066] 1…Robot control system, 10…Robot, 11…End effector, 20…Camera, 21…Sensor, 30, 60…Operation command generation device, 31, 61…Control unit, 32…Memory unit, 33…Input / output unit, 40…Control device, 311, 611…Acquisition unit, 312, 612…Learning unit, 313…Command generation unit, 321, 621…Learning data, 322, 622…Command generation model, 323…Program

Claims

1. A learning unit that randomly extracts the captured image from a single viewpoint and the sensor data at the same time as the capture time of the captured image from learning data including captured images of the working environment from a plurality of viewpoints and sensor data reflecting command data to the robot, performs machine learning, and generates a command generation model for autonomously operating the robot; A command generation unit that inputs the captured image and the sensor data at the current time into the command generation model and generates an operation command for the next time for the robot, and An operation command generation device.

2. The command generation model predicts and outputs the captured image and the sensor data at the next time from the same viewpoint as the capture viewpoint of the captured image at the input current time using the captured image and the sensor data at the input current time. The operation command generation device according to claim 1.

3. The learning data includes position and orientation information of the imaging device at the time of capturing each of the captured images from a plurality of viewpoints, The learning unit randomly extracts the captured image from a single viewpoint, the position and orientation information at the same time as the capture time of the captured image, and the sensor data from the learning data, performs machine learning, and generates the command generation model. The operation command generation device according to claim 2.

4. The captured image is an image captured in time series when the robot performs a predetermined operation, and is an image captured so that the end effector of the robot and the work object of the robot are reflected. The operation command generation device according to claim 3.

5. The captured images from a plurality of viewpoints included in the learning data are captured images captured by a plurality of imaging devices arranged in different positions and orientations in the working environment. The operation command generation device according to claim 4.

6. The captured images from a plurality of viewpoints included in the learning data are the captured images generated by simulating imaging by a plurality of imaging devices arranged in different positions and postures in an environment similar to the work environment constructed on the simulator. The operation command generation device according to claim 4.

7. The captured images from a plurality of viewpoints included in the learning data are the captured images generated by a new viewpoint image generation method using the captured images of a single viewpoint of the work environment. The operation command generation device according to claim 4.

8. The captured images from a plurality of viewpoints included in the learning data include images obtained by performing image processing including at least one of rotation, translation, and scaling on the captured images of a single viewpoint of the work environment. The operation command generation device according to claim 4.

9. It includes an input / output unit that displays the learning data and the progress of learning during the learning of the command generation model. The operation command generation device according to claim 1.

10. The input / output unit displays the captured images and the sensor data from a plurality of viewpoints when acquiring the learning data, the captured images and the sensor data input when generating the operation command, and the predicted captured images and the sensor data. The operation command generation device according to claim 9.

Citation Information

Patent Citations

  • Control device for controlling robot by learning human action, robot system, and production system

    JP2018062016A