Operation command generation device

The operation command generation device addresses the challenge of camera viewpoint deviation in robot control systems by using machine learning to generate a command generation model from randomly extracted single-viewpoint data, allowing robots to perform desired operations efficiently without extended learning times.

WO2025126599A1PCT designated stage expired Publication Date: 2025-06-19HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/032263
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-09-09
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Conventional robot control systems face significant performance deterioration when the camera viewpoint deviates from the learned position, leading to increased learning time and reduced efficiency in generating operation commands for robots performing desired operations.

Method used

An operation command generation device that randomly extracts a captured image from a single viewpoint and corresponding sensor data from learning data, performs machine learning to generate a command generation model, and uses this model to predict and generate operation commands for the robot without increasing learning time, even when the camera viewpoint shifts.

Benefits of technology

Enables the generation of operation commands that allow robots to perform desired operations effectively without increasing learning time, even when the camera viewpoint deviates, thereby improving the robustness and efficiency of autonomous robot operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024032263_19062025_PF_FP_ABST
    Figure JP2024032263_19062025_PF_FP_ABST
Patent Text Reader

Abstract

This operation command generation device comprises: a learning unit that, from learning data including captured images of a work environment captured from a plurality of viewpoints and sensor data in which command data to a robot is reflected, randomly extracts a captured image of a single viewpoint and sensor data at the same time point as the time point when the captured image was captured, performs machine learning, and generates a command generation model for autonomously operating the robot; and a command generation unit that inputs a captured image and sensor data at the current time point to the command generation model, and generates an operation command at the next time point for the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Operation command generation device

[0001] The present invention relates to a motion command generating device.

[0002] In order to improve production efficiency and reduce labor costs, there are increasing efforts to use robots to perform tasks previously performed by humans, such as assembling, welding, and transporting industrial products. However, previous robot systems required extensive programming and a high level of specialized knowledge, which acted as an obstacle to the introduction of robots. Therefore, autonomous learning robot control systems have been proposed, in which robots determine their own behavior based on information from various sensors attached to the robot device. Compared to previous robot control systems, autonomous learning robot control systems do not require extensive programming or a high level of specialized knowledge, and are expected to make it easier to introduce robots. Furthermore, autonomous learning robot control systems are expected to enable robots to generate flexible behavior in response to diverse environmental changes by memorizing and learning from their own operating experiences.

[0003] Examples of the robot's behavioral experience include a method in which the user operating the robot directly teaches the robot the behavior and makes it memorize it, or a method in which the robot observes and imitates the behavior of people, other robots, etc. Furthermore, an autonomous learning robot control system is generally equipped with a learning device, which stores sensor information during behavioral experience and adjusts parameters for generating behavior. This stored sensor information during behavioral experience becomes learning data, and the adjustment of parameters for generating behavior becomes learning. The learning device uses the learning data to repeatedly perform learning so that an expected output value is output for an input value (learning data) to the learning device.

[0004] The learning device stores, as learning data, for example, joint angle information of the robot when experiencing a certain operation and captured images of the work environment taken over time. Using the stored learning data, the learning device inputs joint angle information and captured images at time (t) and performs time machine learning to predict joint angle information and captured images at time (t+1). The autonomous learning robot control system then sequentially inputs the robot's joint angle information and captured images to the learning device that has completed learning, enabling it to automatically generate operation commands according to its own state.

[0005] The above-described method of generating motion commands for a robot directly from sensor information at a certain time without going through object recognition or the like is called an autonomous learning type end-to-end motion generation method. For example, a robot system described in Patent Document 1 is disclosed as an autonomous learning type end-to-end motion generation technology.

[0006] Patent Document 1 describes a machine learning device that recognizes a person's face during a period when the person and the robot work together, classifies and learns the person's behavior based on the weights of a neural network corresponding to the person, and controls the robot's behavior based on the results of classifying the person's behavior.

[0007] Japanese Patent Application Laid-Open No. 2018-62016

[0008] As described above, a technology (Patent Document 1) has been disclosed that automatically controls robot behavior through machine learning using learning data such as captured images of the work environment. In the autonomous learning-based end-to-end behavior generation method described in Patent Document 1, the camera viewpoint is basically fixed and viewpoint deviation is not considered. However, in practice, camera viewpoint deviation often occurs. In a mobile manipulator where the robot itself moves, the camera moves relative to the environment as the robot moves. Even in a stationary robot, the camera is relocated when the robot is relocated, so the camera viewpoint is expected to deviate from the viewpoint at the time of learning. If the camera viewpoint deviates, the positions of the work object and the robot in the captured images also deviate. Because such viewpoint deviation is not learned at all, the performance of the autonomous learning-based robot control system significantly degrades. In such cases, additional learning is required using learning data acquired by shifting the camera viewpoint. Furthermore, additional learning increases the learning time.

[0009] The present invention has been made to solve the above problems, and an object of the present invention is to generate operation commands so that a robot can perform a desired task without increasing the learning time even if the imaging viewpoint of the work environment shifts.

[0010] In order to solve the above problems, the motion command generation device of the present invention comprises a learning unit that randomly extracts a single-viewpoint image and sensor data captured at the same time as the captured image from learning data including captured images of the work environment from multiple viewpoints and sensor data reflecting command data for the robot, performs machine learning, and generates a command generation model for operating the robot autonomously, and a command generation unit that inputs the captured image and sensor data at the current time into the command generation model, and generates a motion command for the robot at the next time.

[0011] According to the present invention having the above configuration, even if the imaging viewpoint of the work environment shifts, it is possible to generate operation commands so that the robot can perform the desired task without increasing the learning time. Other problems, configurations, and effects will become clear from the description of the following embodiments.

[0012] FIG. 1 is a diagram illustrating an example of the overall configuration of a robot control system according to a first embodiment of the present invention. FIG. 2 is a block diagram illustrating an example of the functional configuration of a motion command generating device according to a first embodiment of the present invention. FIG. 3 is a block diagram illustrating a method for learning a command generation model in the motion command generating device according to a first embodiment of the present invention. FIG. 4 is a diagram illustrating an example of a display screen for generating motion commands in the motion command generating device according to the first embodiment of the present invention, the diagram illustrating an example of a display screen when learning is performed. FIG. 5 is a diagram illustrating an example of a display screen for generating motion commands in the motion command generating device according to the first embodiment of the present invention, the diagram illustrating an example of a display screen when a motion command is generated. FIG. 6 is a block diagram illustrating an example of the functional configuration of a motion command generating device according to a second embodiment of the present invention. FIG. 7 is a block diagram illustrating a method for learning a command generation model in the motion command generating device according to the second embodiment of the present invention.

[0013] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted.

[0014] First Embodiment [Example of Overall Configuration of Robot Control System] First, the overall configuration of a robot control system according to a first embodiment of the present invention will be described. Fig. 1 is a diagram showing an example of the overall configuration of a robot control system 1 according to this embodiment. As shown in Fig. 1, the robot control system 1 includes a robot 10, a camera 20 (imaging device), a sensor 21, a motion command generation device 30, and a control device 40. The camera 20, the motion command generation device 30, the control device 40, and the robot 10 are connected in this order. The sensor 21 is disposed on the arm of the robot 10.

[0015] The robot 10 is capable of handling and moving a work object 50, assembling parts, acquiring images, transporting, and the like. The robot 10 may be configured as a robot arm alone equipped with an end effector 11 such as a gripper, or may be equipped with a moving device equipped with crawlers, wheels, or the like, or may be equipped with both. In this embodiment, it is assumed that the robot 10 is a robot arm alone equipped with the end effector 11.

[0016] The camera 20 is an example of an imaging device that captures images of the work environment. The camera 20 may be attached to the robot 10, or may be attached to the work environment side so as to overlook the robot 10, as shown in Fig. 1. Furthermore, instead of arranging a single camera 20 in the work environment, multiple cameras 20 may be arranged in different positions and orientations.

[0017] The sensor 21 is a device that measures the state of the robot 10, the working environment, and the like. The sensor 21 may be, for example, a sensor that measures the current value of a motor provided in a joint of the robot 10, or a tactile sensor, force sensor, or inertial sensor that is externally attached to the robot 10. The sensor 21 may also be a temperature sensor that measures the temperature of the working environment. The sensor 21 may not be a single sensor, but may be multiple sensors. The sensor 21 detects the state of the robot 10, the working environment, and the like, and outputs a detection signal (sensor data) corresponding to the detected state to the motion command generating device 30 via the control device 40. In this embodiment, the sensor 21 is assumed to be a sensor that detects the joint angles of the robot 10, the posture (position), and force (torque) of the end effector 11, and the like. Therefore, in this embodiment, the sensor data measured by the sensor 21 reflects command data that specifies the joint angles of the robot 10, the posture (position), and force (torque) of the end effector 11, and the like.

[0018] The motion command generation device 30 is a terminal device that generates motion commands end-to-end, and has a function of learning a command generation model for causing the robot 10 to operate autonomously, using output information from the sensor 21 and the camera 20. The motion command generation device 30 has a learning mode and an autonomous operation mode. In the learning mode, the motion command generation device 30 inputs sensor data measured by the sensor 21 and images of the work environment captured by the camera 20 as learning data. The motion command generation device 30 then performs machine learning using the learning data to generate a command generation model 322, described below, for causing the robot 10 to operate autonomously.

[0019] On the other hand, in the autonomous operation mode, the motion command generation device 30 inputs the sensor data and captured images at the current time into the trained command generation model 322, and the command generation model 322 predicts the sensor data and captured images at the next time. The motion command generation device 30 then generates motion commands for autonomously operating the robot 10 based on the sensor data and captured images at the next time, and outputs the generated motion commands to the control device 40. Here, the motion commands are command data that specify, for example, the joint angles of the robot 10, the posture (position) and force (torque) of the end effector 11, etc. The sensor data described above reflects the command data during the operation of the robot 10. The captured images are images captured in chronological order when the robot 10 performs a predetermined task, and are images captured so as to capture the end effector 11 of the robot 10 and the work target 50 of the robot 10.

[0020] The control device 40 outputs control commands to the robot 10 based on the operation commands input from the operation command generating device 30, thereby controlling the operation of the robot 10. The control commands are, for example, signals indicating current values, voltage values, etc. for actuators (motors, etc.) provided in the joints of the robot arm, the end effector 11, etc. of the robot 10. When the robot 10 receives a control command from the control device 40, a built-in drive circuit supplies a drive signal to the corresponding actuator.

[0021] [Example of functional configuration of action command generating device] Next, a description will be given of the functional configuration of the action command generating device 30. Fig. 2 is a block diagram showing an example of the functional configuration of the action command generating device 30 according to this embodiment. As shown in Fig. 2, the action command generating device 30 includes a control unit 31, a storage unit 32, and an input / output unit 33.

[0022] The control unit 31 is configured, for example, by a CPU (Central Processing Unit) not shown, and as shown in Fig. 2, has an acquisition unit 311, a learning unit 312, and a command generation unit 313. The control unit 31 controls the operation of each component of the action command generation device 30. Note that the control unit 31 may be configured by a GPU (Graphics Processing Unit) instead of a CPU, or may be configured by a CPU and a GPU.

[0023] The acquisition unit 311 acquires various pieces of instruction information input by an operator from the input / output unit 33. The acquisition unit 311 also acquires, from the storage unit 32, learning data 321 (described later) that is required for machine learning in the learning unit 312. The acquisition unit 311 also acquires sensor data from the sensor 21 and captured images from the camera 20.

[0024] The learning unit 312 randomly extracts a single-viewpoint image and sensor data captured at the same time as the captured image from the learning data 321 described below stored in the memory unit 32, performs machine learning, and generates a command generation model 322 described below for autonomously operating the robot 10 to perform a specified task.

[0025] The command generation unit 313 inputs the captured image of the work environment at the current time and sensor data reflecting command data for the robot 10 into the command generation model 322 learned by the learning unit 312, and generates an operation command for the robot 10 at the next time.

[0026] The storage unit 32 includes storage devices such as a read-only memory (ROM), a random access memory (RAM), a solid-state drive (SSD), etc. The storage unit 32 stores learning data 321, a command generation model 322, and a program 323.

[0027] The program 323 includes a description of a learning method for the command generation model 322 (see FIG. 3 , described later). The command generation model 322 is a deep learning model that inputs a captured image of the work environment at the current time and sensor data reflecting command data for the robot 10, and predicts various parameters for controlling the work behavior of the robot 10, as well as the captured image and sensor data at the next time. The command generation model 322 is generated for each task performed by the robot 10. When there are multiple tasks, a command generation model 322 corresponding to each task is generated and stored in the storage unit 32. Note that the present invention is not limited to generating only one command generation model 322 for each task performed by the robot 10. For example, even if the robot 10 performs the same task, if the attached sensors are changed, the command generation model also changes. Therefore, multiple command generation models 322 may be generated and stored for the same task.

[0028] The training data 321 is training data used for machine learning of the command generation model 322. The training data 321 includes captured images of the work environment from multiple viewpoints, acquired in time series when the robot 10 performs a predetermined task, and sensor data reflecting command data for the robot 10 (see FIG. 3 ). In this embodiment, it is assumed that the captured images from multiple viewpoints included in the training data 321 are images captured by multiple imaging devices (cameras 20) arranged in different positions and orientations in the work environment when the robot 10 performs the predetermined task.

[0029] The position and orientation of the camera 20 is determined based on the upper limit, lower limit, and intermediate positions of the expected deviation of the camera 20 from the work object within the range in which the camera 20 can acquire an image. In machine learning, robustness to the interpolation range of the training data is often confirmed. Therefore, it is preferable to acquire training data by changing the position and orientation of the camera 20 so that it includes the upper limit, lower limit, and intermediate positions and orientations of the expected deviation of the camera 20.

[0030] Furthermore, when the robot 10 is working, the relative positional relationship between the robot 10 and the work object 50 is more important than the absolute positional relationship between them. For this reason, in this embodiment, the captured image is an image that captures the end effector 11 of the robot 10 and the work object 50. Therefore, it is assumed that the position and orientation of the camera 20 is set so that at least the end effector 11 of the robot 10 and the work object 50 are within the angle of view. By including the end effector 11 and the work object 50 within the angle of view, the relative positional relationship between the two can be grasped from the image captured by the camera 20.

[0031] Note that when the learning data 321 is acquired, the motion command generation device 30 is in the learning mode, and therefore the command generation function of the command generation unit 313 is disabled. At this time, the robot 10 operates according to the operator's operation, or automatically operates according to a pre-registered motion plan.

[0032] When the robot 10 is operated by an operator, the control device 40 connected to an operation unit such as a joystick of the robot 10 transmits time-series command data for the robot 10 corresponding to the operation content of the operator to the operation command generation device 30. When the robot 10 is automatically operated according to a pre-registered operation plan, the control device 40 transmits pre-registered time-series command data for the robot 10 to the operation command generation device 30.

[0033] The input / output unit 33 is connected to user interface devices such as a display, keyboard, and mouse, and receives various instruction information input by the user. The input / output unit 33 also includes a communication device and transmits and receives data to and from the control device 40, the camera 20, and the sensor 21. The input / output unit 33 displays captured images and sensor data from multiple viewpoints when acquiring the learning data 321 (see FIG. 4A, described below). The input / output unit 33 also displays learning data and the learning progress status when learning the command generation model 322 (see FIG. 4B, described below). The input / output unit 33 also displays captured images and sensor data input when generating an action command, as well as predicted captured images and sensor data (see FIG. 4C, described below).

[0034] [Learning Method Using Command Generation Model] Next, a description will be given of the motion command generation process in the motion command generation device 30. Fig. 3 is a block diagram showing a learning method for the command generation model 322 in the motion command generation device 30 according to this embodiment. Note that the learning method shown in Fig. 3 is executed by the learning unit 312 calling and executing a program 323 stored in the storage unit 32.

[0035] In this embodiment, during learning, the position and orientation of the camera 20 are changed to include the upper limit, lower limit, and intermediate position and orientation of the camera 20 expected to be misaligned in the work environment, and the robot 10 is made to perform a predetermined task multiple times (for example, K times). Each time the predetermined task is performed, the acquisition unit 311 of the motion command generation device 30 acquires learning data 321. The learning data 321 includes data_1 to data_K, which are learning data acquired each time the predetermined task is performed K times (K is an integer greater than 1). Each of data_1 to data_K includes captured images from viewpoint_1 to viewpoint_N (N is an integer greater than 1) and sensor data reflecting command data for the robot 10 at the time of image capture.

[0036] The learning unit 312 of the motion command generation device 30 first randomly extracts a single-viewpoint captured image from learning data 321 (data_1 to data_K) at the current time (t) or at multiple times up to the current time (t). That is, the learning unit 312 performs a random extraction process on data_1 to data_K as shown in FIG. 3 . The learning unit 312 also extracts sensor data from the same time as the captured image of the extracted captured image. Next, the learning unit 312 inputs the extracted single-viewpoint captured image and sensor data to a command generation model 322, and uses the command generation model 322 to predict a single-viewpoint captured image and sensor data (the same viewpoint as the captured image of the input captured image) at the next time (t+1). Note that while the prediction of the captured image is not essential, it is preferable to include it as a subtask to prevent overlearning.

[0037] The deep learning model performs learning by calculating the error between the output value and the correct value and repeatedly updating various parameters for controlling the work behavior of the robot 10 using an error backpropagation method or the like. In the present invention, the error used to update the parameters of the command generation model 322 is calculated based on the predicted captured image and predicted sensor data output by the command generation model 322 and the captured image and sensor data that are correct values ​​at the same time, as shown in FIG. 3 . The captured image and sensor data that are correct values ​​are, in other words, the captured image and sensor data included in the learning data 321 at the next time point (t+1).

[0038] In the present invention, randomly selecting a single-viewpoint captured image from multiple-viewpoint captured images is intended to predict an action command robustly against changes in viewpoint. Furthermore, although the training data 321 includes multiple-viewpoint captured images, the number of training data 321 input to the command generation model 322 remains K, regardless of the number of viewpoints. Therefore, the training unit 312 of the present invention can perform training in the same amount of time as training with a single viewpoint. Note that the error between the output value calculated by the training unit 312 and the correct value can be calculated using a mean square error, a mean absolute error, or the like. In particular, the error of the captured image can be calculated using, for example, structural similarity (SSIM) for measuring image quality.

[0039] [Display example of action command generation screen] Next, a display example of the action command generation screen in the action command generation device 30 will be described. Fig. 4A to Fig. 4C are diagrams showing display examples of each screen for generating an action command in the action command generation device 30 according to this embodiment. Fig. 4A is a diagram showing an example of a display screen when learning data is acquired. Fig. 4B is a diagram showing an example of a display screen during learning. Fig. 4C is a diagram showing an example of a display screen when an action command is generated.

[0040] 4A, the screen name "Data Acquisition" is displayed at the top of screen SC1, which is the display screen when acquiring learning data. "Learning," the name of the display screen when learning, and "Movement Generation," the name of the display screen when generating motion commands, are grayed out. Below the screen name, a display section P11 for real-time sensor information, a display section P12 for logging setting information, and a display section P13 for robot operation settings are displayed in this order from top to bottom.

[0041] The real-time sensor information display section P11 displays, from left to right, a display section for the captured image from the camera and a display section for the sensor data. The captured images from viewpoints 1 to N are displayed in the captured image display section. The sensor data display section displays a chart showing the time-varying characteristics of the measurement data from the sensor 21. Note that the display format of the sensor data is not limited to a chart, and may be a table or the like.

[0042] 4A, the logging setting information display section P12 allows various settings related to logging, such as the storage time, sampling period, and type of acquired data. The settings on the display section P12 are not limited to the above-mentioned settings, and may include other settings depending on the type of sensor installed on the robot 10. The various settings can be input, for example, by directly entering them into a text box or by selecting them from a selection list.

[0043] The robot operation setting display section P13 displays, from left to right, buttons for setting the operation method of the robot 10 when learning data is acquired, such as a "remote operation" button, a "direct teaching" button, a "reference trajectory_1" button, and a "reference trajectory_2" button. When any of the above buttons is pressed, the robot 10 is set to the operation method corresponding to the pressed button. "Remote operation" corresponds to an operation method in which an operator uses a remote terminal device to instruct the control device 40 connected to the terminal device on specified data for operating the robot 10. "Direct teaching" corresponds to a method in which an operator directly operates the robot 10. "Reference trajectory_1" and "reference trajectory_2" correspond to methods in which the robot 10 automatically operates according to the operation plans corresponding to the registered reference trajectory_1 and reference trajectory_2, respectively.

[0044] On the display screen SC1 when acquiring learning data shown in Figure 4A, the worker can set the operation method of the robot 10 when acquiring learning data, and can also check the acquired learning data and logging information.

[0045] 4B, the screen name "Learning" is displayed at the top of screen SC2, which is the display screen during learning. "Data Acquisition," the name of the display screen during data acquisition, and "Movement Generation," the name of the display screen during motion command generation, are grayed out. Below the screen name, a display area P21 for learning data and a display area P22 for learning progress are displayed in order from top to bottom.

[0046] The learning data display section P21 displays, from left to right, a display section for images captured by the camera 20 and a display section for sensor data. The display section for captured images displays captured images from viewpoints 1 to N included in the learning data from time 0 to time t. The display section for sensor data displays a chart showing the time-varying characteristics of the measurement data of the sensor 21. The display modes of the captured images and sensor data displayed on the learning data display section P21 are the same as those of the captured images and sensor data of the camera shown in FIG. 4A .

[0047] As shown in FIG. 4B, the learning progress display section P22 displays information indicating the learning progress, such as total learning time and remaining learning time, for example, in units of time.

[0048] On the display screen SC2 shown in FIG. 4B, the worker can also check the learning data used during learning and the progress of learning.

[0049] As shown in FIG. 4C , the screen name "Movement Generation" is displayed at the top of screen SC3, which is the display screen when generating motion commands. "Data Acquisition," which is the name of the display screen when acquiring data, and "Learning," which is the name of the display screen when learning, are grayed out. Below the screen name, a display section P31 for the model's input values ​​and predicted values ​​is displayed. At the bottom left of display section P31, a setting section P32 for generating motion commands is displayed. At the bottom right of display section P31 for the model's input values ​​and predicted values, a "Start Motion" button, a "Pause" button, and a "Stop Motion" button are displayed, from top to bottom.

[0050] The display section P31 for the model input values ​​and predicted values ​​displays, from left to right, a display section for the captured image by the camera 20 and a display section for the sensor data. The display section for the captured image displays the image captured by the camera 20 at the current time, which is the current value, and the image captured at the next time, which is the predicted value predicted by the command generation model 322. The display section for the sensor data displays a chart showing the measured value of the sensor at the current time (solid line) and the time change characteristics of the sensor data at the next time, which is the predicted value predicted by the command generation model 322 (dashed line).

[0051] In the operation command generation setting section P32, as shown in FIG. 4C , a command generation model to be used for command generation and various parameters corresponding to the command generation model can be set. Furthermore, various settings can be input, for example, by directly inputting them into a text box or by selecting them from a selection list. After various settings have been configured in the operation command generation setting section P32, when the "Start Operation" button is pressed, the robot 10 starts autonomous operation. Furthermore, when the "Pause" button is pressed during the autonomous operation of the robot 10, the autonomous operation of the robot 10 is paused. Furthermore, when the "Stop Operation" button is pressed during the autonomous operation of the robot 10, the autonomous operation of the robot 10 is stopped.

[0052] 4C, the operator can check the learning data to be input when generating an operation command and the predicted data for the next time. Also, on the display screen SC3 when generating an operation command shown in FIG. 4C, the operator can specify a command generation model and parameters corresponding to the command generation model, and can also start, pause, or stop the autonomous operation of the robot 10.

[0053] [Effects] As described above, the motion command generation device 30 according to this embodiment changes the position and orientation of the camera 20 to acquire learning data that includes images captured from multiple viewpoints. The position and orientation of the camera 20 are determined based on the upper limit, lower limit, and intermediate positions of the expected deviation of the camera 20 relative to the work object 50 within the range in which the camera 20 can acquire images. The motion command generation device 30 also randomly extracts an image captured from a single viewpoint from the acquired images captured from multiple viewpoints. The motion command generation device 30 inputs the extracted single-viewpoint image and sensor data at the current time into the command generation model 322 and uses the command generation model 322 to predict the single-viewpoint image and sensor data at the next time. Therefore, the motion command generation device 30 according to this embodiment can generate motion commands that allow a robot to perform a desired task without increasing the learning time, even if the imaging viewpoint of the work environment is shifted.

[0054] Second Embodiment Next, a motion command generating device 60 according to a second embodiment of the present invention will be described. Fig. 5 is a block diagram showing an example of the functional configuration of the motion command generating device 60 according to this embodiment. As can be seen by comparing Fig. 5 with Fig. 2, the components of the motion command generating device 60 other than the control unit 61, learning data 621, and command generation model 622 are the same as the components of the motion command generating device 30 shown in Fig. 2. Here, a duplicated description of the components that are the same as those of the motion command generating device 60 and the motion command generating device 30 shown in Fig. 2 will be omitted.

[0055] The training data 621 includes sensor data and captured images from multiple viewpoints acquired in chronological order while the robot 10 performs a predetermined task, as well as viewpoint information for the multiple viewpoints (see FIG. 6 ). The viewpoint information for the multiple viewpoints is position and orientation information of the imaging device at the time of capturing each of the captured images from the multiple viewpoints. The viewpoint information is included in the training data 621 to accommodate situations in which the end effector 11 of the robot 10 and the work object 50 (see FIG. 1 ) cannot fit within the field of view of the camera 20. When the end effector 11 of the robot 10 and the work object 50 cannot fit within the field of view of the camera 20, the captured images can be generated by a simulator using the viewpoint information of the camera 20. Furthermore, image processing using data augmentation may be used to process a single-viewpoint captured image of the work environment, and the captured images from the multiple viewpoints may include images that have undergone image processing including at least one of rotation, translation, and scaling. Images captured from multiple viewpoints or new viewpoints can be generated (estimated) using the various image processing methods described above.

[0056] The command generation model 622 is a deep learning model that inputs sensor data, captured images, and viewpoint information at the current time and predicts various parameters for controlling the work behavior of the robot 10, as well as the sensor data and captured images at the next time.

[0057] As shown in Fig. 5, the control unit 61 has an acquisition unit 611, a learning unit 612, and a command generation unit 313. The command generation unit 313 is the same as the command generation unit 313 described in Fig. 2, so a duplicated description will be omitted. The acquisition unit 611 has the same functions as the acquisition unit 311 shown in Fig. 2, and further acquires viewpoint information of the camera 20 when capturing a captured image.

[0058] The learning unit 612 randomly extracts single-viewpoint captured images, position and orientation information of the imaging device at the same time as the captured images, and sensor data from the learning data 621, and performs machine learning to generate a command generation model 622.

[0059] Next, a learning method for the command generation model 622 will be described with reference to Fig. 6. Fig. 6 is a block diagram showing a learning method for the command generation model in the motion command generating device 60 according to this embodiment.

[0060] 6 and 3, the learning method of the command generation model 622 in the learning unit 612 inputs learning data 621 including viewpoint information and predicts the captured image, sensor data, and viewpoint information for the next time point using the command generation model 622. Note that the random sampling process and error calculation in the learning method of the command generation model 622 are the same as the random sampling process and error calculation in the learning method of the command generation model 322 in the learning unit 312 according to the first embodiment, respectively, and therefore will not be described again.

[0061] [Effects] As described above, the motion command generation device 60 according to this embodiment randomly extracts a captured image and viewpoint information from a single viewpoint from captured images from multiple viewpoints and viewpoint information from the multiple viewpoints (position and orientation information of the camera 20). Captured images can be generated by a simulator using the viewpoint information from the camera 20. Furthermore, image processing using data augmentation can be used to rotate, translate, scale, and otherwise modify the captured images generated by the simulator to generate (estimate) captured images from multiple viewpoints or new viewpoints. Therefore, the motion command generation device 60 according to this embodiment has the same effects as the motion command generation device 30 according to the first embodiment, and can generate motion commands to enable a robot to perform a desired task even when the robot's end effector and work object do not fit within the camera's angle of view.

[0062] The present invention is not limited to the above-described embodiments, and various other applications and modifications are possible without departing from the spirit and scope of the present invention as defined in the claims. For example, the above-described embodiments provide detailed and specific descriptions of the configuration of a robot control system in order to clearly explain the present invention, and the present invention is not necessarily limited to systems that include all of the described components. Furthermore, it is possible to replace some of the configurations of the embodiments described herein with configurations of other embodiments, and it is also possible to add configurations of other embodiments to the configurations of one embodiment. Furthermore, it is also possible to add, delete, or replace other configurations with respect to some of the configurations of each embodiment. Furthermore, the control lines and information lines shown are those considered necessary for the explanation, and not all control lines and information lines are necessarily shown in the actual product. In reality, it is reasonable to assume that almost all components are interconnected.

[0063] In the above embodiments, examples have been described in which multiple cameras are used to capture images of the work environment when a certain task is performed when acquiring training data. However, the present invention is not limited to this. When it is difficult to arrange multiple cameras due to space constraints for camera placement, a work environment similar to the real environment may be constructed on a simulator, and images captured by multiple image capture devices arranged at different positions and orientations may be simulated. In other words, the images captured from multiple viewpoints included in the training data may be images of a robot performing a predetermined task, generated by simulating images captured by multiple image capture devices arranged at different positions and orientations in an environment similar to the work environment constructed on the simulator. Because the images are captured on a simulator, it is possible to easily acquire images from multiple viewpoints without preparing multiple actual cameras.

[0064] Furthermore, in the above-described modified example, an example in which a work environment similar to a real environment is constructed on a simulator has been described, but the present invention is not limited to this. It may be difficult to construct a work environment on a simulator when working with fluids such as liquids, or flexible objects such as cables and cloth. In such cases, the images captured from multiple viewpoints included in the training data may be captured images generated by a new viewpoint image generation method using a single viewpoint image of the work environment. Examples of methods for generating new viewpoint images include image generation techniques such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting.

[0065] In addition, in the above-described embodiments, examples have been described in which pseudo-multiple-viewpoint captured images are generated by applying image processing such as rotation, translation, and scaling to a captured image acquired from a single viewpoint of the work environment. The present invention is not limited to this. For example, in a captured image acquired from a single viewpoint, an image of an unseen area, such as a side of a work object, may be generated by supplementarily using image processing such as rotation, translation, and scaling.

[0066] 1...Robot control system, 10...Robot, 11...End effector, 20...Camera, 21...Sensor, 30, 60...Motion command generation device, 31, 61...Control unit, 32...Memory unit, 33...Input / output unit, 40...Control device, 311, 611...Acquisition unit, 312, 612...Learning unit, 313...Command generation unit, 321, 621...Learning data, 322, 622...Command generation model, 323...Program

Claims

1. A motion command generation device comprising: a learning unit that randomly extracts a single-viewpoint image and the sensor data at the same time as the image was captured from learning data including captured images of a work environment from multiple viewpoints and sensor data reflecting command data for a robot, performs machine learning, and generates a command generation model for operating the robot autonomously; and a command generation unit that inputs the captured image and the sensor data at the current time into the command generation model, and generates a motion command for the robot at the next time.

2. The action command generating device according to claim 1, wherein the command generation model uses the captured image and sensor data at the input current time to predict and output the captured image and sensor data at the next time from the same viewpoint as the captured image at the input current time.

3. The action command generating device according to claim 2, wherein the learning data includes position and orientation information of the imaging device at the time of capturing each of the images from multiple viewpoints, and the learning unit randomly extracts the captured image from a single viewpoint, the position and orientation information at the same time as the captured image, and the sensor data from the learning data, performs machine learning, and generates the command generation model.

4. The motion command generating device according to claim 3, wherein the captured images are images captured in time series when the robot is performing a specified task, and the images are captured so as to capture the end effector of the robot and an object being worked on by the robot.

5. The motion command generating device according to claim 4, wherein the captured images from multiple viewpoints included in the learning data are captured by multiple imaging devices arranged in different positions and orientations in the work environment.

6. The motion command generating device according to claim 4, wherein the captured images from multiple viewpoints included in the learning data are captured images generated by simulating capture by multiple imaging devices arranged in different positions and orientations in an environment similar to the work environment constructed on a simulator.

7. The action command generating device according to claim 4, wherein the captured images from multiple viewpoints included in the learning data are captured images generated by a new viewpoint image generation method using the captured image from a single viewpoint of the work environment.

8. The action command generating device of claim 4, wherein the captured images from multiple viewpoints included in the learning data include images that have been subjected to image processing including at least one of rotation, translation, and zooming on the captured image of a single viewpoint of the work environment.

9. The motion command generating device according to claim 1, further comprising an input / output unit for displaying the learning data during learning of the command generation model and the progress of learning.

10. The motion command generating device of claim 9, wherein the input / output unit displays the captured images and the sensor data from multiple viewpoints when the learning data is acquired, and displays the captured images and the sensor data input when the motion command is generated, as well as the predicted captured images and the sensor data.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and computer program

    JP2013161391A

  • Robot, robot system, control unit and control method

    JP2015221485A

  • Operation command generation device and operation command generation method

    JP2023062361A