Robot control system, robot control method, and robot control program
Patent Information
- Application Number
- JP2024573239
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-01-26
AI Technical Summary
Existing robot control systems struggle to adapt robot operations to the dynamic and unpredictable nature of actual workspaces, particularly when dealing with workpieces that have irregular states or appearances, such as soft materials or fresh foods, as they are difficult to accurately predict and control.
A robot control system that includes a setting unit to initialize the next operation amount, a simulation unit to virtually execute tasks, an adjustment unit to adjust the operation based on prediction results, and a robot control unit to control the robot in the actual workspace, using machine learning models to predict and adjust for the current situation, allowing for real-time adaptation and optimal task execution.
Enables robots to operate appropriately in real-time according to the current workspace situation, ensuring accurate processing and task completion by predicting and adjusting for the state of workpieces, even those with irregular appearances or states, thereby improving operational efficiency and accuracy.
Abstract
Description
ROBOT CONTROL SYSTEM, ROBOT CONTROL METHOD, AND ROBOT CONTROL PROGRAM
[0001] One aspect of the present disclosure relates to a robot control system, a robot control method, and a robot control program.
[0002] Patent document 1 describes a robot system that includes an acquisition unit that acquires predetermined first input data that affects the operation of the robot, a calculation unit that calculates, based on the first input data, the computational cost of an inference process using a machine learning model that infers control data used to control the robot, an inference unit that infers the control data using a machine learning model set according to the computational cost, and a drive control unit that controls the robot using the inferred control data.
[0003] Patent No. 7021158
[0004] There is a need for a mechanism that allows a robot to operate appropriately according to the current situation in the real workspace.
[0005] A robot control system according to one aspect of the present disclosure includes a setting unit that is placed in a real workspace and that initially sets a next operation amount for a robot that executes a current task and processes a workpiece; a simulation unit that virtually executes, by simulation, the current task in which the robot operates with the next operation amount to process the workpiece; an adjustment unit that adjusts the next operation amount based on a predicted result obtained by the simulation; and a robot control unit that controls the robot in the real workspace based on the adjusted next operation amount.
[0006] A robot control method according to one aspect of the present disclosure is executed by a robot control system including at least one processor. The robot control method includes the steps of: initially setting a next manipulation variable for a current task of a robot placed in a real workspace and executing the current task to process a workpiece; virtually executing the current task by simulating the robot operating with the next manipulation variable to process the workpiece; adjusting the next manipulation variable based on a prediction result obtained by the simulation; and controlling the robot in the real workspace based on the adjusted next manipulation variable.
[0007] A robot control program according to one aspect of the present disclosure causes a computer to execute the following steps: initially setting a next operation amount for a robot placed in a real workspace to execute a current task and process a workpiece; virtually executing the current task by simulating the robot operating with the next operation amount to process the workpiece; adjusting the next operation amount based on a predicted result obtained by the simulation; and controlling the robot in the real workspace based on the adjusted next operation amount.
[0008] According to one aspect of the present disclosure, a robot can be made to operate appropriately according to the current situation of the actual workspace.
[0009] FIG. 1 is a diagram showing an example of an application of a robot control system; FIG. 2 is a diagram showing an example of a functional configuration of a robot control system; FIG. 3 is a diagram showing an example of a hardware configuration of a computer used for a robot control system; FIG. 4 is a flowchart showing an example of determining a next manipulated variable to control a robot; FIG. 5 is a diagram showing an architecture related to determining a next manipulated variable; FIG. 6 is a diagram showing an example of an architecture related to simulation; and FIG. 7 is a flowchart showing an example of task control.
[0010] Various examples of the present disclosure will be described in detail below with reference to the accompanying drawings. In the description of the drawings, the same or equivalent elements are designated by the same reference numerals, and redundant description will be omitted.
[0011] [System Overview] A robot control system according to the present disclosure is a computer system for autonomously operating a real robot in accordance with the current situation in a real workspace. In one example, the robot control system is arranged in a real workspace and determines a next manipulation amount for a robot executing a current task to process a workpiece, and causes the robot to continue the current task based on the next manipulation amount. In this disclosure, a "task" refers to an operation that a robot is made to perform to achieve a certain goal. For example, the task may be processing a workpiece. The robot's execution of the task results in a result desired by a user of the robot control system. A "current task" refers to a task currently being performed by the robot. In this disclosure, a "manipulated variable" (manipulated value) refers to information for generating the motion of a robot. Examples of a manipulation amount include the angle of each joint of the robot (joint angle) and the torque at each joint (joint torque). A "next manipulation amount" refers to a manipulation amount of the robot for a predetermined time period after the current time.
[0012] A robot control system does not determine the next operation amount of a robot according to a pre-planned target posture or path, but determines the next operation amount according to the current situation of the workspace, which is difficult to accurately predict in advance. For example, the robot control system determines the attributes (e.g., type, state, etc.) of the actual workpiece to be processed as the current situation of the workspace, and determines the next operation amount based on that determination. This type of control can realize robot operation according to the workpiece. For example, the robot control system determines the next operation amount of a robot processing a workpiece according to the current situation of a workpiece whose state transition is not repeatable. Alternatively, the robot control system determines the next operation amount of a robot processing a workpiece according to the current situation of a workpiece whose appearance is uncertain. The robot control system causes the robot to perform the current task based on the determined next operation amount.
[0013] In this disclosure, a workpiece refers to a tangible object that is directly or indirectly affected by the robot's motion. A workpiece may be a tangible object directly processed by a robot, or another tangible object located in the vicinity of a tangible object directly processed by a robot. For example, if the current task is to open packaging surrounding a product, the workpiece may be at least one of the packaging and the product. As another example, if the current task is to pack a product with an undefined appearance into a container, the workpiece may be at least one of the product and the container. "Workpieces with unrepeatable state transitions" refer to workpieces whose next or final state is difficult to predict. "Workpieces with unrepeatable state transitions" can also be described as workpieces whose states change irregularly. Examples of workpieces with unrepeatable state transitions include tangible objects whose external shapes change irregularly due to external forces (e.g., robot motion), such as soft plastic packaging or bags. "Workpieces with unrepeatable appearances" refer to workpieces whose appearances are not exactly the same across individual workpieces. Examples of tangible objects with indefinite appearance include fresh foods such as vegetables, fruits, fish, and meat.
[0014] To robustly control a robot according to the current situation, a robot control system initially sets a next manipulation variable and virtually executes a current task in which the robot operates with the next manipulation variable to process a workpiece through simulation. Simulation is a process of simulating the operation of a robot placed in a real workspace, rather than actually operating the robot. The robot control system adjusts the next manipulation variable based on the predicted results obtained by the simulation and controls the real robot based on the adjusted next manipulation variable. In other words, the robot control system predicts the state of the workpiece at a slightly future time and adjusts and determines the next manipulation variable taking into account the predicted results.
[0015] In one example, the robot control system controls, based on the execution status of the current task, whether to continue the current task without changing the action position, where the robot acts on the workpiece, or to change the action position and continue the current task. The action position is, for example, the position where the robot holds the workpiece with its end effector. In another example, the robot control system controls, based on the execution status of the current task, whether to continue the current task. The robot control system may plan a next task, which is a task that follows the current task, based on the execution status of the current task, and terminate the current task depending on the result of this plan. These controls are also examples of autonomously operating a real robot according to the current situation in the real workspace.
[0016] [System Configuration] Figure 1 is a diagram showing an example of an application of a robot control system. The robot control system 1 shown in this example autonomously operates a real robot 2, which is placed in a real workspace 9 and processes a real workpiece 8, in accordance with the current situation of the workspace 9. The robot control system 1 is connected to a robot controller 3 that controls the robot 2 and a camera 4 that captures images of the workspace 9 via a communication network. The communication network may be a wired network or a wireless network. The communication network may be configured to include at least one of the Internet and an intranet. Alternatively, the communication network may be realized simply by a single communication cable.
[0017] 1 shows the workpiece 8 as a product 81 and a sheet-like packaging material 82 that encases the product 81. In the current task, the robot 2 opens the packaging material 82 that encases the product 81 while changing the holding position of the packaging material 82. Therefore, in the current task, the packaging material 82 is a workpiece that is directly processed by the robot 2, and the product 81 is a workpiece that is indirectly affected by the motion of the robot 2 (i.e., the work performed by the robot 2). In the next task, the robot 2 may process the product 81 directly, or, for example, may move the product 81 away from the packaging material 82 and place it somewhere else.
[0018] The robot 2 is a device that receives power and performs a predetermined action according to a purpose to perform a useful task. In one example, the robot 2 includes multiple joints, an arm, and an end effector 2 a attached to the end of the arm. The robot 2 performs an unpacking task using the end effector 2 a, and in one example, may also perform additional tasks. Examples of the end effector 2 a include a gripper, a suction hand, and a magnetic hand. Each of the multiple joints has a joint axis. Some components of the robot 2, such as the arm and the rotating unit, rotate around the joint axis, allowing the robot 2 to change the position and orientation of the end effector 2 a within a predetermined range. In one example, the robot 2 is a multi-axis serial-link vertical articulated robot. The robot 2 may be a six-axis vertical articulated robot or a seven-axis vertical articulated robot with one additional redundant axis in addition to the six axes. The robot 2 may also be a self-propelled mobile robot, such as an autonomous mobile robot (AMR) or a robot supported by an automated guided vehicle (AGV). Alternatively, the robot 2 may be a stationary robot that is fixed in a predetermined location.
[0019] The robot controller 3 is a device that controls the robot 2 in accordance with a pre-generated operation program. In one example, the robot controller 3 receives from the robot control system 1 operation amounts for the robot to match the position and posture of the end effector with target values indicated in the operation program, and controls the robot 2 in accordance with the operation amounts. The robot controller 3 also transmits the operation amounts to the robot control system 1. As described above, examples of operation amounts include joint angles (angles of each joint) and joint torques (torque at each joint).
[0020] The camera 4 is a device that captures an image of at least a portion of the area within the workspace 9 and generates image data showing the situation within that area as a situation image. In one example, the camera 4 captures at least an image of the workpiece 8 being processed by the robot 2 and generates a situation image showing the current situation of the workpiece 8. The camera 4 transmits the situation image to the robot control system 1. The camera 4 may be fixed to a pillar, a ceiling, or the like, or may be attached near the tip of the arm of the robot 2.
[0021] In the present disclosure, the image data and various images may be still images or a collection of one or more frame images selected from a plurality of frame images that make up a video.
[0022] 2 is a diagram showing an example of the functional configuration of the robot control system 1. In this example, the robot control system 1 includes, as functional components, an acquisition unit 11, a setting unit 12, a simulation unit 13, a prediction / evaluation unit 14, an adjustment unit 15, a repetition control unit 16, a situation evaluation unit 17, a planning unit 18, a decision unit 19, a robot control unit 20, a data generation unit 21, a sample database 22, and a learning unit 23.
[0023] The acquisition unit 11 is a functional module that acquires data used to determine the next operation amount for the current task from the robot controller 3 and the camera 4. The setting unit 12 is a functional module that initially sets the next operation amount. The simulation unit 13 is a functional module that virtually executes the current task, in which the robot 2 operates with the next operation amount to process the workpiece 8, by simulation. The prediction / evaluation unit 14 is a functional module that calculates an evaluation value for the predicted result of the simulation based on a target value previously set in relation to the workpiece 8. In the present disclosure, this evaluation value is also referred to as a "predicted evaluation value." The adjustment unit 15 is a functional module that adjusts the next operation amount based on the predicted evaluation value. The repetition control unit 16 is a functional module that controls the simulation unit 13, the prediction / evaluation unit 14, and the adjustment unit 15 to repeat the simulation, calculation of the predicted evaluation value, and adjustment of the next operation amount. The situation evaluation unit 17 is a functional module that calculates an evaluation value for the execution status of the current task (e.g., the current state of the workpiece 8 being processed) based on a target value previously set in relation to the workpiece 8. In the present disclosure, this evaluation value is also referred to as a "situation evaluation value." The planner 18 is a functional module that plans the next task based on the execution status of the current task. The decision unit 19 is a functional module that determines the next operation of the robot 2 based on at least one of the adjusted next operation amount, the execution status of the current task, and the plan for the next task. The robot controller 20 is a functional module that controls the robot 2 based on the decision.
[0024] The data generation unit 21, the sample database 22, and the learning unit 23 are a group of functional modules for generating a trained model used to control the robot 2. The trained model is generated by machine learning, a method for autonomously finding laws or rules by iteratively learning based on given information. The data generation unit 21 is a functional module that generates at least a portion of the training data used in the machine learning based on the behavior of the robot 2 currently executing a task or the state of the workpiece 8 being processed in the current task. The sample database 22 is a functional module that stores the training data generated by the data generation unit 21 and training data collected in advance before the robot 2 executes the current task. In other words, the sample database 22 can store both the training data collected in advance and training data obtained while the robot 2 is executing the current task. The learning unit 23 is a functional module that generates a trained model by machine learning using the training data in the sample database 22. In one example, the learning unit 23 generates at least one of a control model used by the setting unit 12, a state prediction model used by the simulation unit 13, an evaluation model used by the prediction evaluation unit 14 and the situation evaluation unit 17, and a planning model used by the planner 18. These trained models are realized, for example, by a neural network such as a deep neural network (DNN). Generating the trained model through machine learning makes it possible to quantify the evaluation of the workpiece 8 or task based on tacit knowledge (knowledge based on human experience or intuition) and appropriately control the robot 2.
[0025] The robot control system 1 can be realized by any type of computer. The computer may be a general-purpose computer such as a personal computer or a business server, or may be incorporated into a dedicated device that executes a specific process.
[0026] 3 is a diagram showing an example of the hardware configuration of a computer 100 used for the robot control system 1. In this example, the computer 100 includes a main body 110, a monitor 120, and an input device .
[0027] The main body 110 is a device having a circuit 160. The circuit 160 has a processor 161, a memory 162, a storage 163, an input / output port 164, and a communication port 165. The number of each hardware component may be one or more. The storage 163 records programs for configuring each functional module of the main body 110. The storage 163 is a computer-readable recording medium such as a hard disk, a non-volatile semiconductor memory, a magnetic disk, or an optical disk. The memory 162 temporarily stores programs loaded from the storage 163, calculation results of the processor 161, and the like. The processor 161 configures each functional module by executing programs in cooperation with the memory 162. The input / output port 164 inputs and outputs electrical signals to and from the monitor 120 or the input device 130 in response to instructions from the processor 161. The communication port 165 performs data communication via a communication network N with other devices, such as the robot controller 3, in response to instructions from the processor 161.
[0028] The monitor 120 is a device for displaying information output from the main body 110. For example, the monitor 120 is a device capable of displaying graphics, such as a liquid crystal panel.
[0029] The input device 130 is a device for inputting information to the main body 110. Examples of the input device 130 include operation interfaces such as a keypad, a mouse, and an operation controller.
[0030] The monitor 120 and the input device 130 may be integrated as a touch panel. For example, the main body 110, the monitor 120, and the input device 130 may be integrated as a tablet computer.
[0031] Each functional module of the robot control system 1 is realized by loading a robot control program onto the processor 161 or memory 162 and having the processor 161 execute the program. The robot control program includes code for realizing each functional module of the robot control system 1. The processor 161 operates the input / output port 164 and the communication port 165 in accordance with the robot control program, and executes reading and writing of data from and to the memory 162 or the storage 163.
[0032] The robot control program may be provided in the form of a non-transitory recording medium such as a CD-ROM, a DVD-ROM, or a semiconductor memory. Alternatively, the robot control program may be provided via a communications network as a data signal superimposed on a carrier wave.
[0033] [Robot Control Method] (Robot Control Based on Next Manipulation Amount) As an example of a robot control method according to the present disclosure, an example of determining a next manipulation amount and controlling a robot will be described with reference to FIGS. 4 to 6. FIG. 4 is a flowchart showing the series of processes as processing flow S1. That is, the robot control system 1 executes processing flow S1. FIG. 5 is a diagram showing an architecture related to determining the next manipulation amount. In FIG. 5, time (t-1) is the current time point, and time t is the time point when robot control based on the next manipulation amount is executed, i.e., a time point slightly after the present. FIG. 6 is a diagram showing an example of an architecture related to a simulation.
[0034] In step S11, the acquisition unit 11 acquires observation data indicating the current situation of the workspace 9. For example, the acquisition unit 11 acquires the operation amount of the robot 2 processing the workpiece 8 as the current operation amount from the robot controller 3, and acquires a situation image indicating the workpiece 8 being processed by the robot 2 from the camera 4. That is, the observation data may include the current operation amount and the situation image.
[0035] In step S12, the setting unit 12 calculates the next operation amount OP of the robot 2 in the current task based on the observation data. initThe setting unit 12 inputs the situation image and the current manipulated variable into the control model 12a and sets the next manipulated variable OP. init The control model 12a is a trained model that has been trained to calculate a second operation amount of the robot 2 at a second time point after the first time point, based on a sample image showing a workpiece at the first time point and a first operation amount of the robot 2 at the first time point.
[0036] In step S13, the simulation unit 13 executes a simulation based on the set next operation amount. In the first loop process, the simulation unit 13 executes a simulation based on the set next operation amount OP initThe simulation unit 13 virtually executes a current task of processing a workpiece 8 by operating the robot 2. In one example, for the simulation, the simulation unit 13 uses a robot model representing the robot 2 and a context related to elements (hereinafter also referred to as "components") constituting the workspace 9. The robot model is electronic data indicating specifications of the robot 2 and the end effector 2 a. The specifications may include parameters related to the structure of the robot 2 and the end effector 2 a, such as shape and dimensions, and parameters related to the function of the robot 2 and the end effector 2 a, such as the range of motion of each joint and the performance of the end effector 2 a. The context is electronic data indicating various attributes of one or more components of the workspace 9 and may be expressed, for example, by text (i.e., natural language). The elements constituting the workspace 9 can also be considered tangible objects existing in the workspace 9. The context may include various attributes of the workpiece 8, such as the type, shape, physical properties, dimensions, and color of the workpiece 8. Alternatively, the context may include various attributes of the robot 2 or the end effector 2a, such as the type, shape, size, and color of the robot 2 or the end effector 2a. Alternatively, the context may include attributes of the surrounding environment of the robot 2 and the workpiece 8. Examples of surrounding environment attributes include the type, shape, and color of the workbench, the type and color of the floor, and the type and color of the wall. Thus, the context may include at least one of workpiece information about the workpiece 8, robot information (robot model) about the robot 2, and environmental information about the surrounding environment. The simulation unit 13 generates a prediction result including a predicted state of the workpiece 8 for a predetermined future time span including time t, based on the robot model, the context, and the set next operation amount. The prediction result may further include the operation of the robot 2 for that time span.
[0037] An example of the simulation will be described in detail with reference to FIG. 6 . In this example, the simulation unit 13 performs kinematics / dynamics calculations based on the next manipulation input to generate a virtual motion of the robot 2 operating with the next manipulation input. This process generates a motion that takes into account the geometric constraints (kinematics) and mechanical constraints (dynamics) of the robot 2. Next, the simulation unit 13 uses a renderer to generate a motion image Pm that shows the virtual motion of the robot 2. Because the virtual motion is generated based on the next manipulation input, the renderer that depicts the virtual motion can be said to be a process based on the next manipulation input. In one example, the simulation unit 13 generates the motion image Pm from the next manipulation input using differentiable kinematics / dynamics and a differentiable renderer. This example can be implemented to make a series of processes from input of the next manipulation input to output of the predicted evaluation value differentiable in order to use backpropagation (error backpropagation) to reduce the predicted evaluation value.
[0038] The simulation unit 13 inputs the virtual motion and context indicated by the motion image Pm into the state prediction model 13a and generates a predicted state of the workpiece 8 processed by the robot 2 operating with the next operation amount. The predicted state may indicate a change over time in the status of the workpiece 8 for a predetermined future time span including time t. The predicted state may further indicate the operation of the robot 2 for that time span. In one example, the state prediction model 13a generates a predicted image Pr indicating the predicted state. The state prediction model 13a is a trained model that has been trained to predict the state of the workpiece 8 based on the motion and context of the robot 2. The simulation unit 13 may generate a predicted state (predicted image Pr) that indicates a change over time in the virtual appearance state of the workpiece 8 due to the virtual motion of the robot 2. The appearance state of the workpiece refers to, for example, the external shape of the workpiece.
[0039] 4 and 5. In step S14, the prediction evaluation unit 14 evaluates the prediction result obtained by the simulation. In one example, the prediction evaluation unit 14 calculates a predicted evaluation value E pred is calculated. In one example, the target value is represented by a target image, which is an image showing a predetermined state of the workpiece 8 to be compared with the predicted state. The target value may be the final state of the workpiece 8 in the current task, in which case the target image shows that final state. Alternatively, the target value may be the state (intermediate state) of the workpiece 8 at a point in time during the current task, for example, the intermediate state of the workpiece 8 at the time when the next operation amount is actually applied (time t in the example of FIG. 5). In this case, the target image shows that intermediate state. Predicted evaluation value E pred is a value indicating how close the predicted state of the workpiece 8 is to the target value. pred The smaller the predicted state is, the closer it is to the target value. In one example, the prediction evaluation unit 14 inputs the predicted image Pr and the target image to the evaluation model 14a and calculates a predicted evaluation value E pred The evaluation model 14a is a trained model that has been trained to calculate an evaluation value based on the state of the workpiece 8 and a target value (for example, an image showing the state of the workpiece 8 and a target image showing the target value).
[0040] In step S15, the adjustment unit 15 adjusts the next manipulated variable based on the evaluation of the prediction result (predicted state). For example, the adjustment unit 15 adjusts the next manipulated variable based on the evaluation of the change over time in the virtual appearance state of the workpiece 8. The adjustment unit 15 adjusts the next manipulated variable so that the state of the workpiece 8 can be closer to the target value than the predicted state, and calculates the adjusted next manipulated variable OP adj The adjustment unit 15 may set the predicted evaluation value E pred The larger the value, that is, the further the predicted state deviates from the target value, the larger the adjustment amount of the next manipulated variable may be.
[0041] In step S16, the repetition control unit 16 determines whether or not to end the adjustment of the next manipulated variable based on a predetermined end condition. The end condition may be that the repetition process has been repeated a predetermined number of times, or that a predetermined calculation time has elapsed. Alternatively, the end condition may be that the previously obtained predicted evaluation value E pred and the predicted evaluation value E obtained this time pred The difference between the predicted evaluation value E pred It may also be that the trend has stagnated or converged.
[0042] If the next manipulated variable is to be further adjusted (NO in step S16), the process returns to step S13. In the repeated step S13, the simulation unit 13 adjusts the set next manipulated variable OP adj The simulation unit 13 executes a simulation based on the set next manipulated variable OP adj and the context, and generates at least a predicted state of the workpiece 8 for a predetermined time span in the future including the time t. adj is different from any of the next manipulated variables used in the past loop processes, the predicted state obtained in the current loop process may be different from any of the predicted states used in the past loop processes. As described above, the simulation unit 13 can generate a predicted image Pr indicating the predicted state. In the repeated step S14, the prediction evaluation unit 14 inputs the predicted state (predicted image Pr) obtained this time and the target value (target image) to the evaluation model 14a to calculate a predicted evaluation value E pred In the repeated step S15, the adjustment unit 15 calculates the predicted evaluation value E pred By repeating this process, a plurality of adjusted next manipulated variables OP adj is obtained.
[0043] If the adjustment is to be ended (YES in step S16), the process proceeds to step S17. adj to the final next operation amount OP finalFor example, the determination unit 19 determines the next manipulated variable OP finally obtained by the repetitive processing. adj The next operation amount OP final Alternatively, the determination unit 19 determines the next manipulated variable OP at which the state of the workpiece 8 is expected to converge to the target value related to the workpiece 8. adj The next operation amount OP final For example, the determination unit 19 may determine the next manipulated variable OP that is expected to be able to converge the workpiece 8 to its target value most quickly. adj The next operation amount OP final It is determined as follows.
[0044] In step S18, the robot control unit 20 calculates the next operation amount OP final The real robot 2 in the workspace 9 is controlled based on the next operation amount OP final is a plurality of next operation amounts OP adj Since the robot control unit 20 is one of the following, the adjusted next operation amount OP adj In other words, the robot control unit 20 controls the robot 2 based on the following operation amount OP final The robot controller 3 transmits the operation amount OP final The robot 2 continues to execute the current task according to the control and further processes the workpiece 8.
[0045] The robot control system 1 can repeatedly execute the process flow S1 at predetermined time intervals. In the example of FIG. 5 , the robot control system 1 executes the process flow S1 based on observation data at time (t-1) to determine the next operation amount at time t. The real robot 2 processes the real workpiece 8 based on that operation amount. The robot control system 1 acquires the operation amount at time t from the robot controller 3 as the current operation amount, and acquires from the camera 4 a situation image showing the state of the workpiece 8 at time t. The robot control system 1 executes the process flow S1 based on this observation data to determine the next operation amount at time (t+1). The real robot 2 further processes the real workpiece 8 based on that operation amount. The robot control system 1 repeats this process to sequentially generate the next operation amount while causing the robot 2 to execute the current task.
[0046] (Task Control) As an example of a robot control method according to the present disclosure, an example of task control will be described with reference to Fig. 7. Fig. 7 is a flowchart showing a series of steps in task control as process flow S2. That is, the robot control system 1 executes process flow S2. In one example, the robot control system 1 executes process flows S1 and S2 in parallel.
[0047] In step S21, the acquisition unit 11 acquires observation data indicating the current situation of the workspace 9. This process is the same as step S11. As described above, the acquisition unit 11 can acquire the current operation amount and a situation image as the observation data.
[0048] In step S22, the decision unit 19 determines whether to continue the current task. To make this determination, the situation evaluation unit 17 calculates a situation evaluation value, which is an evaluation value for the execution status of the current task, based on a target value previously set in relation to the workpiece 8. In one example, the target value is represented by a target image, which is an image showing a predetermined state of the workpiece 8, which is compared with the current state of the workpiece 8 represented by the situation image. The target value may be the final state of the workpiece 8 in the current task, and in this case, the target image indicates that final state. The situation evaluation value is a value indicating how close the execution status of the current task (e.g., the current state of the workpiece 8) is to the target value. In the present disclosure, the smaller the situation evaluation value, the closer the execution status of the current task (e.g., the current state of the workpiece 8) is to the target value. In one example, the situation evaluation unit 17 inputs the situation image and the target image into an evaluation model to calculate the situation evaluation value. The decision unit 19 switches whether to continue the current task based on the situation evaluation value. Therefore, the decision unit 19 also functions as a judgment unit. For example, the decision unit 19 determines to continue the current task if the situation evaluation value is equal to or greater than a predetermined threshold, and determines to end the current task if the situation evaluation value is less than the threshold. If the current task is to be continued (YES in step S22), the process proceeds to step S23, and if the current task is to be ended (NO in step S22), the process proceeds to step S26.
[0049] In step S23, the decision unit 19 determines whether to change the action position for the current task. To make this determination, the situation evaluation unit 17 calculates a situation evaluation value, which is an evaluation value for the execution status of the current task, based on a target value previously set in relation to the workpiece 8. As in step S22, the situation evaluation unit 17 may calculate an evaluation value for the current state of the workpiece 8 as the execution status of the current task. Unlike step S22, the target value in step S23 may be an ideal state (intermediate state) of the workpiece 8 at a point in the middle of the current task. In this case, the target image indicates that intermediate state. In one example, the situation evaluation unit 17 inputs the current image and the target image into an evaluation model to calculate the situation evaluation value. The decision unit 19 determines whether to change the action position from the current position based on the situation evaluation value. For example, if the situation evaluation value is equal to or greater than a predetermined threshold, the decision unit 19 determines to change the action position, and if the situation evaluation value is less than the threshold, the decision unit 19 determines not to change the action position. If the operating position is to be changed (YES in step S23), the process proceeds to step S24, and if the operating position is not to be changed (NO in step S24), the process proceeds to step S25.
[0050] In step S24, the robot control unit 20 controls the robot 2 to change the action position and continue the current task. For example, the robot control unit 20 analyzes the situation image to search for and determine a new action position. The robot control unit 20 then generates a command to change the action position from the current position to the new position and sends the command to the robot controller 3. The robot controller 3 controls the robot 2 in accordance with the command. The robot 2 changes the action position from the current position to the new position in accordance with the control and continues executing the current task.
[0051] In step S25, the robot control unit 20 controls the robot 2 so that the robot 2 continues the current task without changing the operating position. This process corresponds to step S18. The robot control unit 20 controls the robot 2 so that the robot 2 continues the current task without changing the operating position. This process corresponds to step S18. final The robot control unit 20 controls the robot 2 based on the following operation amount OP finalThe robot controller 3 transmits the operation amount OP final In accordance with this control, the robot 2 continues to execute the current task without changing the working position, and further processes the workpiece 8.
[0052] In step S26, the robot control unit 20 controls the robot 2 to complete the current task. In one example, for this process, the planner 18 inputs the situation image into a planning model to generate a plan for the next task following the current task. The planning model is a trained model that has been trained to plan the next task based on the current situation of the workpiece 8. The robot control unit 20 controls the robot 2 to complete the current task based on the results of the plan. For example, the plan for the next task may include a plan for the robot's operation in the next task, and the robot control unit 20 may control the posture of the robot 2 at the end of the current task so that the robot 2 can smoothly transition to that operation. The robot control unit 20 sends a command to the robot controller 3 to cause the actual robot 2 to complete the current task. The robot controller 3 causes the robot 2 to complete the current task in accordance with the command. In one example, the robot control unit 20 further sends a command for the next task to the robot controller. The robot controller 3 causes the robot 2 to start the next task in accordance with the command.
[0053] As shown in the process flow S2, the robot control unit 20 can control the robot 2 based on a switch (determination) as to whether or not to continue the current task, or a determination as to whether or not to change the operating position.
[0054] The robot control system 1 can repeatedly execute the process flow S2 at predetermined time intervals. As a result of this repetition, the robot 2 continues the current task while changing the operating position as needed, processing the workpiece 8, and finally completes the current task.
[0055] [Machine Learning] In one example, the learning unit 23 generates or updates at least one trained model used in the robot control system 1 through supervised learning. In supervised learning, training data (sample data) including multiple data records indicating combinations of input data to be processed by a machine learning model and correct answers in output data from the machine learning model is used. The learning unit 23 performs the following process for each data record of the training data. That is, the learning unit 23 inputs the input data indicated by the data record to the machine learning model. The learning unit 23 updates a set of parameters in the machine learning model by performing backpropagation (error backpropagation) based on the error between the output data estimated by the machine learning model and the correct answer indicated by the data record. The learning unit 23 generates or updates a trained model by repeating the process for each data record until a predetermined termination condition is met. The termination condition may be that all data records of the training data have been processed. Note that each trained model generated or updated is a computational model estimated to be optimal, and is not necessarily a "computational model that is actually optimal."
[0056] The generation or updating of a control model will now be described. In one example, the data generation unit 21 generates a data record including a combination of the current operation amount and situation image acquired by the acquisition unit 11, and the next operation amount adjusted based on the current operation amount (e.g., the finally determined next operation amount). The data generation unit 21 stores the data record in the sample database 22 as at least a part of the training data. The learning unit 23 updates the control model through machine learning using the data record. In this machine learning, the learning unit 23 uses the adjusted next operation amount (e.g., the finally determined next operation amount) as a correct answer.
[0057] As another example, the data generation unit 21 generates a teacher image from the predicted image Pr generated by the simulation unit 13 (state prediction model). The data generation unit 21 modifies the predicted image based on modification information for modifying the scene shown by the predicted image, i.e., the scene showing the predicted state, to obtain a teacher image showing a different state from the predicted state. The modification information may be information for changing the workpiece shown in the predicted image. For example, the modification information may be information for modifying a predicted image showing a scene in which plastic bags are being processed into a teacher image showing a scene in which burlap sacks are being processed. Alternatively, the modification information may be information for modifying the surrounding environment of the robot 2 and the workpiece 8. For example, the modification information may be information for modifying a predicted image showing a scene in which a workpiece placed on a workbench is being processed into a teacher image showing a scene in which a workpiece placed on the floor is being processed. The data generation unit 21 generates a data record including the current manipulation variable, the next manipulation variable (e.g., the finally determined next manipulation variable) adjusted based on the current manipulation variable, and the teacher image. The data generation unit 21 stores the data record in the sample database 22 as at least part of the training data. The learning unit 23 may update the control model through machine learning using the data record, or may newly generate another control model for initially setting the next manipulated variable. In either case, in such machine learning, the learning unit 23 uses the adjusted next manipulated variable (for example, the finally determined next manipulated variable) as the correct answer.
[0058] The generation or updating of a state prediction model will now be described. In one example, the data generation unit 21 generates a data record including a combination of an adjusted next operation amount (e.g., a finally determined next operation amount) and a real state, which is the state of the real workpiece 8 processed by the real robot 2 controlled by the robot control unit 20 based on the adjusted next operation amount. That is, the data generation unit 21 generates a data record including a combination of the adjusted next operation amount and a situation image obtained as a result of the adjusted next operation amount. The data generation unit 21 stores the data record in the sample database 22 as at least a part of the training data. The learning unit 23 may update the state prediction model or generate a new state prediction model through machine learning using the data record. In this machine learning, the learning unit 23 uses kinematics / dynamics and a renderer to generate a virtual motion of the robot 2 from the next operation amount indicated by the training data, and inputs the generated motion and a predetermined context into the machine learning model. The learning unit 23 uses the situation image as a correct answer.
[0059] As another example, when the context is expressed by text, the learning unit 23 may accept text indicating the context, compare the text with the predicted state generated by the state prediction model, and update the predicted state model through machine learning based on the results of the comparison. For example, the learning unit 23 may input a predicted image into an encoder model that converts a situation indicated by an image into text, thereby generating text indicating the predicted situation. The learning unit 23 may then compare the text indicating the context with the text indicating the predicted situation, and update the state prediction model through machine learning using the difference between the two texts (i.e., loss). Alternatively, the learning unit 23 may calculate latent variables from both the text indicating the context and the predicted state (predicted image), and update the state prediction model through machine learning using the difference between the two latent variables (loss). Alternatively, the learning unit 23 may use a predetermined comparison model that compares the text indicating the context with the predicted state (predicted image), and update the state prediction model through machine learning based on the comparison results obtained from the comparison model.
[0060] The generation of an evaluation model will now be described. In one example, the sample database 22 stores in advance, as training data, a plurality of data records each representing a combination of image data showing the state of a workpiece being processed at a certain point in the past, a target value previously set in relation to the workpiece, and an evaluation value set for the state of the workpiece. The learning unit 23 generates an evaluation model through machine learning using the training data. In this machine learning, the learning unit 23 uses the evaluation value indicated by the training data as a correct answer.
[0061] The generation of the planning model will now be described. In one example, the sample database 22 stores in advance, as training data, a plurality of data records indicating a combination of image data showing the state of a workpiece being processed at a certain point in the past and a plan for a next task related to the workpiece. The plan for the next task may include a plan for the operation of the robot 2 in the next task. The learning unit 23 generates the planning model through machine learning using the training data. In this machine learning, the learning unit 23 uses the plan for the next task indicated by the training data as a correct answer.
[0062] Generating a trained model corresponds to the learning phase of machine learning. Prediction or estimation using the generated trained model corresponds to the operation phase of machine learning. The above process flows S1 and S2 correspond to the operation phase.
[0063] The combination of the control model, state prediction model, and evaluation model in the above example can also be said to be a command generation model that has been trained to output command posture data that indicates the posture of the robot at a second time point after the first time point at which the image data (situation image) is acquired, when the image data (situation image) is input. The next manipulated variable can be said to be the command posture data.
[0064] [Modifications] The technology according to the present disclosure has been described in detail above based on various examples. However, the present disclosure is not limited to the above examples. The technology according to the present disclosure can be modified in various ways without departing from the spirit of the present disclosure.
[0065] The robot control system may control at least one of a plurality of real robots working together to process a workpiece according to a current situation in a real workspace where the real robots are located. For example, the robot control system may control each of two six-axis robots working together to open a package. The robot control system may execute the above process flows S1 and S2 for at least one of the plurality of robots, for example, for each robot.
[0066] The control model may be trained to calculate a second operation variable of the robot at a second time point based on one of a sample image showing a workpiece at a first time point and a first operation variable of the robot at the first time point. When this control model is used, the setting unit inputs one of the current operation variable and the situation image into the control model to initially set the next operation variable. Alternatively, the control model may be trained to calculate the second operation variable based on at least one of the sample image and the first operation variable, as well as at least one of a context, a target value indicating a final goal or intermediate goal related to the workpiece, and a teaching point. When this control model is used, the setting unit inputs at least one of the current operation variable and the situation image, and at least one of the context, the target value, and the teaching point into the control model to initially set the next operation variable.
[0067] The simulation method and the configuration of the state prediction model are not limited to the above examples. For example, the simulation unit may generate a predicted state of the workpiece by inputting a set next operation amount into a state prediction model that has been trained to predict the state of the workpiece based on the next operation amount. Therefore, the simulation unit may generate a predicted state without using kinematics / dynamics and a renderer.
[0068] The trained model is portable between computer systems. The robot control system may not include functional modules corresponding to the data generation unit 21, the sample database 22, and the training unit 23, and may use a trained model generated in another computer system.
[0069] The adjustment unit may adjust the initially set next manipulated variable, and the robot control unit may control the robot based on the adjusted next manipulated variable. Therefore, the robot control system does not need to include a functional module equivalent to the repetitive control unit 16.
[0070] The adjustment unit may adjust the next manipulated variable without using the predicted evaluation value. For example, the adjustment unit may calculate the difference between a target image indicating a target value and a predicted image, and adjust the next manipulated variable based on this difference. For example, the adjustment unit may increase the adjustment amount of the next manipulated variable as the difference increases. In such a modified example, the robot control system does not need to include a functional module equivalent to the prediction evaluation unit 14.
[0071] The robot control system may not execute a process of determining whether to end the current task and controlling the robot. Alternatively, the robot control system may not execute a process of determining whether to change the action position in the current task and controlling the robot. Alternatively, the robot control system may not execute a process of planning the next task and ending the current task depending on the results of the planning. Therefore, the robot control system may not include a functional module corresponding to at least one of the situation evaluation unit 17, the judgment unit (part of the decision unit 19), and the planner 18.
[0072] In the above example, the camera 4 captures the current situation in the workspace 9, but a different type of sensor than a camera, such as a laser sensor, may detect the current situation in the actual workspace.
[0073] The hardware configuration of the system is not limited to a configuration in which each functional module is realized by executing a program. For example, at least some of the functional modules may be configured by logic circuits specialized for the functions, or may be configured by an ASIC (Application Specific Integrated Circuit) that integrates the logic circuits.
[0074] The processing steps of the method executed by at least one processor are not limited to the above examples. For example, some of the steps or processes described above may be omitted, or the steps may be executed in a different order. Furthermore, any two or more of the steps described above may be combined, or some of the steps may be modified or deleted. Alternatively, other steps may be executed in addition to the steps described above.
[0075] When comparing the magnitude of two numbers within a computer system or computer, either of the two criteria "greater than or equal to" and "greater than" can be used, or either of the two criteria "less than or equal to" and "under".
[0076] [Supplementary Notes] As can be seen from the various examples above, the present disclosure includes the following aspects. (Supplementary Note 1) A robot control system comprising: a setting unit that initially sets a next manipulation amount for a robot that is placed in a real workspace and executes a current task to process a workpiece; a simulation unit that virtually executes the current task by simulating the robot operating with the next manipulation amount to process the workpiece; an adjustment unit that adjusts the next manipulation amount based on a prediction result obtained by the simulation; and a robot control unit that controls the robot in the real workspace based on the adjusted next manipulation amount. (Supplementary Note 2) The robot control system according to Supplementary Note 1, wherein the prediction result includes a predicted state that is a state of the workpiece processed by the robot operating with the next manipulation amount, and the adjustment unit adjusts the next manipulation amount based on at least the predicted state. (Supplementary Note 3) The robot control system according to Supplementary Note 2, further comprising: an evaluation unit that calculates an evaluation value for the predicted state of the workpiece based on a target value that is preset for the workpiece, and the adjustment unit adjusts the next manipulation amount based on the evaluation value. (Supplementary Note 4) The robot control system according to Supplementary Note 3, further comprising: a repetition control unit that controls the simulation unit, the evaluation unit, and the adjustment unit so as to repeat the simulation, calculation of the evaluation value, and adjustment of the next operation amount based on the evaluation value, and a determination unit that determines a final next operation amount from the plurality of adjusted next operation amounts obtained by the repetition, wherein the robot control unit controls the robot based on the final next operation amount. (Supplementary Note 5) The robot control system according to any one of Supplementary Notes 1 to 4, wherein the setting unit initially sets the next operation amount based on image data that shows the workpiece being processed by the robot in the actual workspace.(Supplementary Note 6) The robot control system according to any one of Supplements 1 to 5, wherein the setting unit inputs a current operation amount of the robot processing the workpiece into a control model trained to calculate a second operation amount at a second time point after a first time point based on a first operation amount of the robot at the first time point, to initially set the next operation amount. (Supplementary Note 7) The robot control system according to any one of Supplements 2 to 4, wherein the simulation unit generates a virtual motion of the robot operating with the next operation amount, and inputs the generated virtual motion into a state prediction model trained to predict a state of the workpiece based on the robot motion, to generate the predicted state. (Supplementary Note 8) The robot control system according to Supplementary Note 7, wherein the simulation unit generates a change over time in a virtual appearance state of the workpiece due to the virtual motion as the predicted state, and the adjustment unit adjusts the next operation amount based at least on the change over time in the virtual appearance state of the workpiece. (Supplementary Note 9) The robot control system according to Supplementary Note 7 or 8, wherein the simulation unit generates the predicted state by inputting the generated virtual motion and the context into a state prediction model that has been trained to predict the state of the workpiece further based on a context related to elements that constitute the workspace. (Supplementary Note 10) The robot control system according to any one of Supplements 7 to 9, further comprising a learning unit that updates the state prediction model by machine learning using teacher data including a combination of the adjusted next operation amount and an actual state that is the state of the workpiece processed by the robot controlled by the robot control unit. (Supplementary Note 11) The robot control system according to Supplementary Note 10, wherein the learning unit receives text as the context related to elements that constitute the workspace, compares the text with the predicted state, and updates the state prediction model by machine learning based on the result of the comparison. (Supplementary Note 12) The robot control system according to any one of Supplements 7 to 11, wherein the simulation unit generates an image showing the virtual motion using a renderer based on the next operation amount.(Supplementary Note 13) The robot control system according to any one of Supplements 1 to 12, further comprising: an evaluation unit that calculates an evaluation value regarding the execution status of the current task based on a target value that is set in advance in relation to the workpiece; and a determination unit that switches whether or not to continue the current task based on the evaluation value, wherein the robot control unit controls the robot based on the switching. (Supplementary Note 14) The robot control system according to any one of Supplements 1 to 13, further comprising: an evaluation unit that calculates an evaluation value regarding the execution status of the current task based on a target value that is set in advance in relation to the workpiece; and a determination unit that determines whether or not to change an action position, which is a position at which the robot acts on the workpiece in the current task, from its current position, based on the evaluation value, wherein the robot control unit, when it is determined that the action position should be changed from the current position, causes the robot to change the action position from the current position to a new position and continues the current task. (Supplementary Note 15) The robot control system according to any one of Supplementary Notes 1 to 14, further comprising: a planning unit that plans a next task following a current task based on a planning model that has been trained to output a plan for the next task following the current task when image data showing the workpiece being processed by the robot in the actual workspace is input, and the image data, and the robot control unit controls the robot in accordance with a result of the plan by the planning unit to end the current task. (Supplementary Note 16) The robot control system according to Supplementary Note 6, further comprising a learning unit that updates the control model by machine learning using teacher data including a combination of the current operation amount and the adjusted next operation amount.(Supplementary Note 17) The robot control system according to Supplementary Note 16, further comprising a data generation unit that generates the teacher data, wherein the simulation unit generates the predicted image based on the next operation amount and a state prediction model that has been trained to generate a predicted image showing a predicted state of the workpiece based on a motion of the robot operating with the next operation amount and a context related to elements that make up the workspace, the predicted image, the data generation unit modifies the predicted image based on modification information for modifying a scene showing the predicted state to generate a teacher image showing a state different from the predicted state, and generates the teacher data including a combination of the current operation amount, the adjusted next operation amount, and the teacher image, and the learning unit updates the control model by machine learning using the teacher data that further includes the teacher image, or generates another control model for initially setting the next operation amount. (Supplementary Note 18) A robot control method executed by a robot control system having at least one processor, comprising: a step of initially setting a next operation amount in a current task for a robot that is placed in a real workspace and executes a current task to process a workpiece, a step of virtually executing the current task by simulating the robot operating with the next operation amount to process the workpiece, a step of adjusting the next operation amount based on a prediction result obtained by the simulation, and a step of controlling the robot in the real workspace based on the adjusted next operation amount. (Supplementary Note 19) A robot control program that causes a computer to execute the steps of: initially setting a next operation amount in a current task for a robot that is placed in a real workspace and executes a current task to process a workpiece, a step of virtually executing the current task by simulating the robot operating with the next operation amount to process the workpiece, a step of adjusting the next operation amount based on the prediction result obtained by the simulation, and a step of controlling the robot in the real workspace based on the adjusted next operation amount.(Supplementary Note 20) A robot control system comprising: a robot that executes a current task on a workpiece; an acquisition unit that sequentially acquires image data showing the workpiece during execution of the current task; a command generation unit that sequentially generates command posture data corresponding to the sequentially acquired image data based on a command generation model that has been trained to output command posture data that indicates the posture of the robot at a second time point after a first time point at which the image data is acquired when at least the image data is input; and a robot control unit that controls the robot to execute the current task based on the sequentially generated command posture data. (Supplementary Note 21) The robot control system according to Supplementary Note 20 further comprises: an evaluation unit that evaluates an execution status of the current task at the time the image data is acquired based on an evaluation model that has been trained to output an evaluation value regarding the execution status of the current task when at least the image data is input; and a determination unit that switches whether to continue control of the robot based on the generated command posture data depending on a result of the evaluation by the evaluation unit. (Supplementary Note 22) The robot control system according to Supplementary Note 21, further comprising an action point extraction unit that extracts a new action point of the robot on the workpiece, wherein the robot control unit controls the robot to execute the current task while acting on the workpiece at the new action point when control of the robot is not to be continued. (Supplementary Note 23) The robot control system according to Supplementary Note 20, further comprising a planning unit that plans the next task based on the acquired image data and a planning model that has been trained to output a plan of a next task following the current task when at least the image data is input, wherein the robot control unit terminates execution of the current task by the robot according to a result of the planning by the planning unit.
[0077] According to Supplements 1, 18, and 19, how a robot will next process a workpiece in a current task currently being performed in reality is predicted by a simulation based on an initially set next operation variable. The next operation variable is then adjusted based on the prediction result, and the robot in the real workspace is controlled based on the adjusted next operation variable. Because the next operation variable for continuing to control the robot is adjusted based on a prediction from a simulation of the current task, the robot can be operated appropriately in accordance with the current situation in the real workspace. Furthermore, such appropriate robot control enables the current task and workpiece to converge to a desired target state.
[0078] According to Supplementary Note 2, the state of the workpiece in the current task is predicted by simulation, and the next operation amount is adjusted based on the prediction result. The state of the workpiece being processed by the robot is directly related to whether the current task will be successful. Therefore, by adjusting the next operation amount based on the state of the workpiece shortly after, it is possible to have the real robot appropriately process the real workpiece according to the current situation in the real workspace.
[0079] According to Supplementary Note 3, the state of the workpiece obtained by the simulation shortly after the simulation is evaluated based on a target value related to the workpiece, and the next manipulated variable is adjusted based on the evaluation. This target value can be said to indicate the desired state of the workpiece. Since the next manipulated variable is adjusted taking into account the target value, it is possible to have the real robot appropriately process the workpiece so as to bring the real workpiece into the desired state, depending on the current situation in the real workspace.
[0080] According to Supplementary Note 4, the next manipulation amount for controlling the robot is finally determined after repeated adjustment of the next manipulation amount based on the simulation and evaluation of the prediction results. By repeating the adjustment, the real robot can be controlled with a more appropriate next manipulation amount.
[0081] According to Supplementary Note 5, the next manipulated variable is initially set based on image data showing the actual workpiece being processed. By using image data that clearly shows the current status of the workpiece, the next manipulated variable can be appropriately initially set according to that status. Therefore, it can be expected that the next manipulated variable that is adjusted will also be a more appropriate value.
[0082] According to Supplementary Note 6, the next manipulation amount is initially set by the control model (trained model) based on the current manipulation amount of the real robot. This process is expected to more reliably obtain a next manipulation amount that is continuous with the current manipulation amount, i.e., a next manipulation amount that allows the real robot to operate smoothly. Therefore, the adjusted next manipulation amount is expected to be an appropriate value that achieves smooth robot control without abrupt changes in the posture of the real robot.
[0083] According to Supplementary Note 7, a virtual motion of the robot operating with the following operation amount is generated, and the motion is input into a state prediction model (trained model) to predict the state of the workpiece being processed by the robot. By generating a predicted state from the virtual motion using the state prediction model, the state of the workpiece can be accurately predicted.
[0084] According to Supplementary Note 8, a virtual change in the appearance state of the workpiece over time is generated as a predicted state, and the next operation amount is adjusted based on the change over time. Generally, for workpieces whose appearance changes, it is difficult to predict how the appearance will change in the near future. By predicting the change using simulation and then adjusting the next operation amount, the robot can appropriately process workpieces whose appearance changes irregularly according to the current situation.
[0085] According to Supplementary Note 9, the virtual motion of a robot operating with the next manipulation input and context related to the elements constituting the workspace are input into a state prediction model to predict the state of a workpiece being processed by the robot. The state prediction model generates a predicted state by accepting context input, so predicted states can be generated for various types of workpieces. By introducing a general-purpose state prediction model that can process multiple types of workpieces and separately generating the robot's motion and the predicted state of the workpiece in a simulation, general-purpose robot control that is independent of the components of the workspace becomes possible. Furthermore, since there is no need to prepare a state prediction model for each component of the workspace, the labor required to prepare the state prediction model can be reduced or minimized.
[0086] According to Supplementary Note 10, a state prediction model that predicts the state of a workpiece is updated by machine learning based on the state (actual state) of a workpiece processed by a robot that is actually controlled based on the adjusted next operation amount. Machine learning using new data obtained by actual robot control can further improve the accuracy of the state prediction model.
[0087] According to Supplementary Note 11, the state prediction model is updated by machine learning based on the comparison result between the text indicating the context and the predicted state of the work. This machine learning makes it possible to realize a state prediction model that generates a predicted state according to the context given in text format.
[0088] According to Supplementary Note 12, an image showing the virtual motion of the robot is generated by a renderer. By using the renderer, the three-dimensional structure and three-dimensional motion of the robot can be accurately represented in an image. As a result, it becomes possible to obtain more accurate prediction results from the simulation.
[0089] According to Supplementary Notes 13 and 21, the execution status of the current task is evaluated based on a target value related to the work, and whether or not to continue the current task is switched (i.e., determined) based on the evaluation. Since the determination regarding the continuation of the current task is made in consideration of the target value, which can be said to indicate the desired state of the work, the current task can be appropriately continued or terminated depending on the current situation of the actual workspace.
[0090] According to Supplementary Notes 14 and 22, the execution status of the current task is evaluated based on a target value related to the work, and whether or not to change the action position of the work is determined based on that evaluation. Since the action position in the current task is controlled taking into account the target value, which can be said to indicate the desired state of the work, the work can be appropriately processed in the current task according to the current situation in the actual workspace.
[0091] According to Supplementary Notes 15 and 23, image data representing the workpiece being processed by the current task is processed by a planning model (trained model), a next task following the current task is planned, and the current task is controlled according to the results of that plan. By controlling the current task in consideration of the plan for the next task rather than the current task itself, it is possible to smoothly carry out a series of processes continuing from the current task to the next task.
[0092] According to Supplementary Note 16, a control model for initially setting the next manipulated variable is updated by machine learning based on the current manipulated variable and the adjusted next manipulated variable. Machine learning using the next manipulated variable that was actually used for robot control can further improve the accuracy of the control model.
[0093] According to Supplementary Note 17, a predicted image showing a predicted state of the workpiece, generated by a state prediction model in a simulation, is used to generate a teacher image showing a different state from the predicted state. Then, a control model is updated or newly generated by machine learning based on a combination of the current manipulated variable, the adjusted next manipulated variable, and the teacher image. This machine learning using the teacher image generated using the predicted image can improve the accuracy of the control model or prepare a new control model according to variables in the workspace. Furthermore, the labor required to prepare the control model can be reduced or minimized.
[0094] According to Supplementary Note 20, image data representing a workpiece being processed by a current task at a first time point is generated based on a command generation model, and command posture data at a second time point after the first time point is generated. Then, based on the command posture data, the robot is controlled to further execute the current task. Since the command posture data for continuing to control the robot is generated according to the current situation of the current task, the robot can be operated appropriately according to the current situation of the actual workspace. Furthermore, such appropriate robot control enables the current task and the workpiece to converge to a desired target state.
[0095] 1...robot control system, 2...robot, 2a...end effector, 3...robot controller, 4...camera, 8...work, 9...work space, 11...acquisition unit, 12...setting unit, 12a...control model, 13...simulation unit, 13a...state prediction model, 14...prediction evaluation unit, 14a...evaluation model, 15...adjustment unit, 16...repetition control unit, 17...situation evaluation unit, 18...planning unit, 19...decision unit, 20...robot control unit, 21...data generation unit, 22...sample database, 23...learning unit, Pm...motion image, Pr...predicted image.
Claims
1. a setting unit that initially sets a next operation amount for a current task for a robot that is disposed in a real working space and executes a current task to process a workpiece; a simulation unit that generates a state of the workpiece processed by the robot operating with the next operation amount as a predicted state; A prediction evaluation unit that calculates an evaluation value of the predicted state of the workpiece based on a target value that is preset in relation to the workpiece; an adjustment unit that adjusts the next manipulated variable based on the evaluation value; a robot control unit that controls the robot in the real workspace based on the adjusted next operation amount; A robot control system comprising:
2. A robot control system as described in claim 1, which repeats a process including initial setting of the next operation amount by the setting unit, generation of the predicted state by the simulation unit, calculation of the evaluation value by the prediction evaluation unit, adjustment of the next operation amount by the adjustment unit, and control of the robot based on the adjusted next operation amount, thereby causing the robot to execute the current task while sequentially generating the next operation amount.
3. a repetitive control unit that controls the simulation unit, the prediction evaluation unit, and the adjustment unit so as to repeat generation of the predicted state, calculation of the evaluation value, and adjustment of the next manipulated variable based on the evaluation value; a determination unit that determines a final next manipulated variable from the plurality of adjusted next manipulated variables obtained by the repetition; Further comprising: The robot control unit controls the robot based on the final next operation amount. The robot control system of claim 2 .
4. The setting unit initially sets the next operation amount based on image data showing the workpiece being processed by the robot in the actual working space. The robot control system of claim 1 .
5. The simulation unit is performing kinematic / dynamic calculations based on the next manipulated variable to generate a virtual motion of the robot operating with the next manipulated variable; inputting the generated virtual motion into a state prediction model trained to predict a state of the workpiece based on the motion of the robot, thereby generating the predicted state; The robot control system according to any one of claims 1 to 4.
6. The simulation unit generates a motion image indicating the virtual motion by using a renderer based on the next operation amount. The robot control system of claim 5.
7. The simulation unit inputs the virtual motion represented by the motion image into the state prediction model to generate a prediction image representing the predicted state; the prediction evaluation unit calculates the evaluation value based on the predicted image and a target image expressing the target value; The robot control system of claim 6.
8. The simulation unit generates a change over time in a virtual appearance state of the workpiece due to the virtual motion as the predicted state, The adjustment unit adjusts the next operation amount based at least on a time-dependent change in the virtual appearance state of the workpiece. The robot control system of claim 5.
9. The simulation unit inputs the generated virtual motion and the context into the state prediction model, which has been trained to predict the state of the workpiece based on a context related to elements that configure the working space, to generate the predicted state. The robot control system of claim 5.
10. The robot control system according to claim 5, further comprising a learning unit that updates the state prediction model by machine learning using teacher data including a combination of the adjusted next operation amount and an actual state, which is a state of the workpiece processed by the robot controlled by the robot control unit.
11. The learning unit is accepting text indicating a context for elements that make up the workspace; comparing the text with the predicted state and updating the state prediction model by machine learning based on the results of the comparison; The robotic control system of claim 10.
12. The setting unit inputs a current operation amount of the robot that processes the workpiece into a control model that has been trained to calculate a second operation amount at a second time point after the first time point based on a first operation amount of the robot at the first time point, and initially sets the next operation amount. The robot control system according to any one of claims 1 to 4.
13. The robot control system according to claim 12 , further comprising a learning unit that updates the control model by machine learning using training data including a combination of the current operation amount and the adjusted next operation amount.
14. The data generating unit generates the teacher data. the simulation unit generates a predicted image based on the next operation amount and a state prediction model that has been trained to generate a predicted image indicating a predicted state of the workpiece based on a motion of the robot operating with the next operation amount and a context related to elements that configure the workspace; and The data generation unit modifying the predicted image based on modification information for modifying a scene showing the predicted state to generate a teacher image showing a different state different from the predicted state; generating the teacher data including a combination of the current operation amount, the adjusted next operation amount, and the teacher image; The learning unit updates the control model by the machine learning using the teacher data further including the teacher image, or generates another control model for initially setting the next manipulated variable. The robotic control system of claim 13.
15. a status evaluation unit that calculates an evaluation value regarding an execution status of the current task based on a target value that is preset in relation to the work; a determination unit that determines whether or not to continue the current task based on the evaluation value; Further comprising: The robot control unit controls the robot based on the switching. The robot control system according to any one of claims 1 to 4.
16. a status evaluation unit that calculates an evaluation value regarding an execution status of the current task based on a target value that is preset in relation to the work; a determination unit that determines whether or not to change an action position, which is a position where the robot acts on the workpiece in the current task, from a current position based on the evaluation value; Further comprising: when it is determined that the action position is to be changed from the current position, the robot control unit causes the robot to change the action position from the current position to a new position and continue the current task. The robot control system according to any one of claims 1 to 4.
17. a planning unit that plans the next task based on a planning model that has been trained to output a plan for a next task following the current task when image data showing the workpiece being processed by the robot in the actual working space is input, and the image data; the robot control unit controls the robot in accordance with a result of the plan by the planner to end the current task. The robot control system according to any one of claims 1 to 4.
18. 1. A robot control method executed by a robot control system having at least one processor, comprising: A step of initially setting a next operation amount in a current task for a robot that is arranged in a real workspace and executes a current task to process a workpiece; generating a state of the workpiece processed by the robot operating with the next manipulated variable as a predicted state; calculating an evaluation value of the predicted state of the workpiece based on a preset target value related to the workpiece; adjusting the next manipulated variable based on the evaluation value; controlling the robot in the real workspace based on the adjusted next operation amount; A robot control method comprising:
19. A step of initially setting a next operation amount in a current task for a robot that is arranged in a real workspace and executes a current task to process a workpiece; generating a state of the workpiece processed by the robot operating with the next manipulated variable as a predicted state; calculating an evaluation value of the predicted state of the workpiece based on a preset target value related to the workpiece; adjusting the next manipulated variable based on the evaluation value; controlling the robot in the real workspace based on the adjusted next operation amount; A robot control program that causes a computer to execute the above.