Information processing method, information processing device, and information processing program

By adjusting multiple evaluation indicators of robot movements, outputting images of multiple optimal solutions and obtaining the optimal solution selected by the user and its historical reference information, the problem of difficulty in optimizing multiple evaluation indicators in existing technologies is solved, and the safety of robot movements and task performance are improved.

CN120677039APending Publication Date: 2025-09-19PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202480014103.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-11
Filing Date
2024-02-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

It is difficult to optimize multiple evaluation indicators of robot movements in the existing technology, which makes it difficult for users to select an optimal solution from multiple optimal solutions.

Method used

By obtaining the trajectory information of the robot's movements, adjusting multiple evaluation indicators for optimization, outputting multiple solution display images of the optimal solutions, and obtaining the optimal solution selected by the user and its historical reference information, the user is assisted in selecting the optimal solution.

Benefits of technology

It optimizes multiple evaluation indicators, helps users select the best solution from multiple optimal solutions, and improves the safety of robot movements and task performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120677039A_ABST
    Figure CN120677039A_ABST
Patent Text Reader

Abstract

The information processing device acquires trajectory information pertaining to a trajectory of an operation of the robot, adjusts a parameter of the robot, which optimizes two or more evaluation indexes for evaluating the operation of the robot, on the basis of the trajectory information, and calculates a plurality of optimal solutions for the two or more evaluation indexes. Outputting a solution display image in which the plurality of calculated optimal solutions are drawn in a plane or a space having two or more evaluation indexes as coordinate axes, and acquiring at least one optimal solution selected by a user from among the plurality of optimal solutions displayed in the solution display image; and outputting reference information based on a history of actions of the robot corresponding to the acquired at least one optimal solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a technology for calculating a plurality of optimal solutions for two or more evaluation indicators for evaluating the behavior of a robot and presenting the calculated plurality of optimal solutions. Background Art

[0002] For example, the method shown in patent document 1 includes: a process of receiving trajectory information that specifies the trajectory of the object action of the robot; a process of obtaining the value of an evaluation index of the control result when the robot performs the object action using each initial parameter set based on an instruction to adjust the parameter set used to control the object action for one or more initial parameter sets; a process of displaying one or more reference displays based on the obtained evaluation index values ​​on a display unit; and a process of receiving condition information that determines the conditions for optimizing the parameter set, and optimizing the parameter set according to the condition information for the evaluation index to determine the value of the new parameter set.

[0003] However, in the above-mentioned existing technologies, the optimization of a single evaluation index is disclosed, but the optimization of more than two evaluation indexes is not considered. It is difficult to assist users in selecting one optimal solution from multiple optimal solutions, and further improvement is needed.

[0004] Prior art literature

[0005] Patent Literature

[0006] Patent Document 1: Japanese Patent Application Laid-Open No. 2022-70451 Summary of the Invention

[0007] The present disclosure is made to solve the above-mentioned problems, and its object is to provide a technology that can optimize two or more evaluation indicators and assist a user in selecting an optimal solution from among multiple optimal solutions.

[0008] The information processing method involved in the present disclosure is executed by a computer, and the information processing method includes: obtaining trajectory information related to the trajectory of the robot's movement; based on the trajectory information, adjusting the parameters of the robot for optimizing more than two evaluation indicators for evaluating the robot's movement, thereby calculating multiple optimal solutions for the more than two evaluation indicators; outputting a solution display image that depicts the calculated multiple optimal solutions in a plane or space with the more than two evaluation indicators as coordinate axes; obtaining at least one optimal solution selected by a user from the multiple optimal solutions displayed in the solution display image; and outputting reference information based on the history of the robot's movement corresponding to the at least one optimal solution obtained.

[0009] According to the present disclosure, two or more evaluation indicators can be optimized, and a user can be assisted in selecting one optimal solution from a plurality of optimal solutions. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a diagram showing the configuration of a teaching support system according to this embodiment.

[0011] Figure 2 This is a flowchart for explaining teaching support processing performed by the information processing device in the embodiment of the present disclosure.

[0012] Figure 3 This is a diagram showing an example of a display screen for accepting input of two or more evaluation indices and displaying a plurality of optimal solutions in the present embodiment.

[0013] Figure 4 Is used to illustrate Figure 2 Flowchart of the segmentation process in step S4.

[0014] Figure 5 FIG. 1 is a diagram showing an example of a reference information display screen displayed on the display unit in the present embodiment.

[0015] Figure 6 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 1 of the present embodiment.

[0016] Figure 7 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 2 of the present embodiment.

[0017] Figure 8 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 3 of the present embodiment.

[0018] Figure 9 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 4 of the present embodiment.

[0019] Figure 10 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 5 of the present embodiment.

[0020] Figure 11 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 6 of the present embodiment.

[0021] Figure 12 This is a diagram showing an example of a reference information display screen displayed on the display unit in Modification 7 of the present embodiment.

[0022] Figure 13 This is a diagram showing comparison results of learning curves indicating the growth of the area of ​​the Pareto optimal solution in the wiping operation.

[0023] Figure 14 This is a diagram showing the comparison results of learning curves indicating the growth of the area of ​​the Pareto optimal solution in the door opening operation.

[0024] Figure 15 This is a diagram showing the comparison results of the areas of the Pareto optimal solutions when prior knowledge is used in the IC-SLD method, the GMM method, and the SLD method, and the areas of the Pareto optimal solutions when no prior knowledge is used. DETAILED DESCRIPTION

[0025] (Understanding that forms the basis of this disclosure)

[0026] Robot control based on teaching recording and playback is a technology widely adopted in the industry due to its intuitiveness and ease of installation in embedded systems. Generally, position control is used for teaching playback, but there is a possibility of damage to the robot or its surroundings due to unexpected contact. Therefore, in order to achieve safe and desired task movements, it is indispensable to introduce impedance control with appropriately designed rigidity parameters. The rigidity parameters affect the safety of the movement and the reproducibility of the trajectory, and there is a trade-off relationship between the two. Therefore, the problem of determining the rigidity parameters requires the simultaneous optimization of the evaluation indicators of task performance and safety.

[0027] However, the above-mentioned prior art discloses the optimization of a single evaluation index, but does not consider the optimization of more than two evaluation indexes, making it difficult to assist users in selecting one optimal solution from multiple optimal solutions.

[0028] In order to solve the above problems, the following technology is disclosed.

[0029] (1) An information processing method according to one embodiment of the present disclosure is executed by a computer, the method comprising: obtaining trajectory information related to a trajectory of a robot's movement; adjusting parameters of the robot for optimizing two or more evaluation indicators for evaluating the robot's movement based on the trajectory information, thereby calculating a plurality of optimal solutions for the two or more evaluation indicators; outputting a solution display image in which the calculated plurality of optimal solutions are depicted on a plane or space having the two or more evaluation indicators as coordinate axes; obtaining at least one optimal solution selected by a user from among the plurality of optimal solutions displayed in the solution display image; and outputting reference information based on a history of the robot's movement corresponding to the at least one optimal solution obtained.

[0030] This configuration enables optimization of two or more evaluation indicators, and calculation of multiple optimal solutions for the two or more evaluation indicators. Furthermore, since reference information based on the robot's motion history corresponding to at least one optimal solution selected by the user from among the multiple optimal solutions is presented to the user, the user can be assisted in selecting the optimal solution from among the multiple optimal solutions.

[0031] (2) In the information processing method described in (1) above, the method may further include accepting input of the two or more evaluation indicators from the user.

[0032] According to this configuration, since the user input of two or more evaluation indices is accepted, it is possible to optimize the two or more evaluation indices desired by the user and calculate a plurality of optimal solutions for the two or more evaluation indices.

[0033] (3) In the information processing method described in (1) or (2) above, the calculation of the multiple optimal solutions may include: dividing the movement of the robot into multiple segments; and repeatedly searching for the optimal parameters of each segment through multi-objective Bayesian optimization, thereby optimizing the two or more evaluation indicators.

[0034] According to this configuration, two or more evaluation indices are optimized by dividing the robot's motion into a plurality of segments and repeatedly searching for optimal parameters for each segment through multi-objective Bayesian optimization. Therefore, two or more evaluation indices can be optimized.

[0035] (4) In the information processing method described in (3) above, the action may be represented by a plurality of combinations of impedance-controlled motion equations, the parameters include stiffness parameters of the impedance control, and the segmentation of the action includes estimating the stiffness parameters in each motion equation and the switching timing of each motion equation so that the error between the predicted trajectory and the taught trajectory in each motion equation is minimized.

[0036] According to this configuration, the motion can be divided into a plurality of segments by the switching timing of each motion equation, and the optimal stiffness parameter in each of the plurality of segments can be calculated by using the stiffness parameter in each motion equation.

[0037] (5) In the information processing method described in (4) above, the optimization of the two or more evaluation indicators may include: weighting the acquisition function in the multi-objective Bayesian optimization by the inferred rigid parameters, and repeatedly searching for the optimal rigid parameters in each segment by the acquisition function.

[0038] According to this configuration, since the rigidity parameters estimated in advance are used for the multi-objective Bayesian optimization, the efficiency of the optimization process can be improved.

[0039] (6) In the information processing method described in any one of (1) to (5) above, the reference information may include at least one time series data of the trajectory corresponding to the at least one optimal solution.

[0040] According to this configuration, the user can confirm at least one time-series data of a trajectory corresponding to at least one optimal solution, and the user can be assisted in selecting one optimal solution from among a plurality of optimal solutions.

[0041] (7) In the information processing method described in (6) above, when two or more optimal solutions are selected from the plurality of optimal solutions, two or more time series data corresponding to the two or more optimal solutions may be overlapped or displayed in a row.

[0042] According to this configuration, the user can easily compare two or more time-series data corresponding to two or more optimal solutions, and the user can be further assisted in selecting one optimal solution from among a plurality of optimal solutions.

[0043] (8) In the information processing method described in (6) above, the at least one time series data corresponding to the at least one optimal solution and the time series data of the trajectory of the target movement of the robot may be displayed superimposed or aligned.

[0044] According to this configuration, the user can easily compare at least one time series data corresponding to at least one optimal solution with the time series data of the trajectory of the robot's target motion, and can further assist the user in selecting one optimal solution from multiple optimal solutions.

[0045] (9) In the information processing method described in (6) above, the at least one time series data corresponding to the at least one optimal solution and the time series data of at least one parameter that changes according to the action corresponding to the at least one optimal solution may be overlapped or displayed in an aligned manner.

[0046] According to this structure, in addition to at least one time series data corresponding to at least one optimal solution, the user can also confirm the time series data of at least one parameter that changes according to the action corresponding to at least one optimal solution, which can further assist the user in selecting one optimal solution from multiple optimal solutions.

[0047] (10) In the information processing method described in any one of (1) to (5) above, the reference information may include at least one moving image obtained by recording the action corresponding to the at least one optimal solution.

[0048] According to this configuration, the user can check at least one moving image recorded with respect to the action corresponding to at least one optimal solution, and the user can be assisted in selecting one optimal solution from among a plurality of optimal solutions.

[0049] (11) In the information processing method described in (10) above, when two or more optimal solutions are selected from the plurality of optimal solutions, two or more dynamic images corresponding to the two or more optimal solutions may be displayed in an overlapping or aligned manner.

[0050] According to this configuration, when two or more optimal solutions are selected from a plurality of optimal solutions, two or more moving images corresponding to the two or more optimal solutions are superimposed or displayed side by side.

[0051] Therefore, the user can easily compare two or more moving images corresponding to two or more optimal solutions, and the user can be further assisted in selecting one optimal solution from among a plurality of optimal solutions.

[0052] (12) In the information processing method described in (10) above, the at least one dynamic image corresponding to the at least one optimal solution and a dynamic image obtained by recording the target action of the robot may be overlapped or displayed in parallel.

[0053] According to this configuration, the user can easily compare at least one moving image corresponding to at least one optimal solution with a moving image recorded of the target robot movement, thereby further assisting the user in selecting one optimal solution from among multiple optimal solutions.

[0054] (13) In the information processing method described in (10) above, the at least one dynamic image corresponding to the at least one optimal solution and time series data of at least one parameter that changes according to the action corresponding to the at least one optimal solution are arranged and displayed.

[0055] According to this structure, in addition to at least one dynamic image corresponding to at least one optimal solution, the user can also confirm the time series data of at least one parameter that changes according to the action corresponding to at least one optimal solution, which can further assist the user in selecting one optimal solution from multiple optimal solutions.

[0056] (14) In the information processing method described in any one of (1) to (13) above, the two or more evaluation indicators may include an evaluation indicator of task performance and an evaluation indicator of safety.

[0057] According to this configuration, it is possible to optimize the evaluation index of task performance and the evaluation index of safety, which are in a trade-off relationship.

[0058] Furthermore, the present disclosure can be implemented not only as an information processing method that performs the above-described characteristic processing, but also as an information processing device or the like having a characteristic structure corresponding to the characteristic processing performed by the information processing method. Furthermore, the present disclosure can also be implemented as a computer program that causes a computer to perform the characteristic processing included in the above-described information processing method. Therefore, the following other embodiments can also achieve the same effects as the above-described information processing method.

[0059] (15) The information processing device involved in other embodiments of the present disclosure comprises: a trajectory information acquisition unit for acquiring trajectory information related to the trajectory of the robot's movement; a calculation unit for adjusting the parameters of the robot for optimizing two or more evaluation indicators for evaluating the robot's movement based on the trajectory information, thereby calculating multiple optimal solutions for the two or more evaluation indicators; a first output unit for outputting a solution display image in which the calculated multiple optimal solutions are depicted in a plane or space with the two or more evaluation indicators as coordinate axes; an optimal solution acquisition unit for acquiring at least one optimal solution selected by a user from the multiple optimal solutions displayed in the solution display image; and a second output unit for outputting reference information based on the history of the robot's movement corresponding to the at least one optimal solution acquired.

[0060] (16) The information processing program involved in other aspects of the present disclosure enables the computer to perform the following functions: obtain trajectory information related to the trajectory of the robot's movement, adjust the parameters of the robot for optimizing two or more evaluation indicators used to evaluate the robot's movement based on the trajectory information, thereby calculating multiple optimal solutions for the two or more evaluation indicators, depicting the calculated multiple optimal solutions on a solution display image in a plane or space with the two or more evaluation indicators as coordinate axes, obtain at least one optimal solution selected by a user from the multiple optimal solutions displayed in the solution display image, and output reference information based on the history of the robot's movement corresponding to the at least one optimal solution obtained.

[0061] (17) Other aspects of the present disclosure relate to a non-temporary computer-readable recording medium recording an information processing program, wherein the information processing program enables a computer to perform the following functions: obtain trajectory information related to the trajectory of the robot's movement, adjust the parameters of the robot for optimizing two or more evaluation indicators for evaluating the robot's movement based on the trajectory information, thereby calculating multiple optimal solutions for the two or more evaluation indicators, depicting the calculated multiple optimal solutions on a solution display image in a plane or space with the two or more evaluation indicators as coordinate axes, obtain at least one optimal solution selected by a user from the multiple optimal solutions displayed in the solution display image, and output reference information based on the history of the robot's movement corresponding to the at least one optimal solution obtained.

[0062] Hereinafter, the embodiments of the present disclosure will be described with reference to the accompanying drawings. In addition, the embodiments described below all represent a specific example of the present disclosure. The numerical values, shapes, structural elements, steps, the order of steps, etc. shown in the following embodiments are an example and are not intended to limit the present disclosure. In addition, among the structural elements in the following embodiments, the structural elements that are not recorded in the independent technical solutions representing the highest concepts are described as arbitrary structural elements. In addition, in all embodiments, the various contents can also be combined.

[0063] (Implementation Method)

[0064] Figure 1 It is a diagram showing the configuration of a teaching support system according to this embodiment.

[0065] Figure 1 The teaching support system shown includes an information processing device 1 , a robot 2 , a display unit 3 , and an input unit 4 .

[0066] The robot 2 performs a given action. For example, the given action may involve contact between the robot 2 and a person or object. The person teaches the robot 2 the given action.

[0067] Robot 2 comprises a main body, an arm mounted on the main body, and an end effector mounted at the distal end of the arm. Robot 2 is a general-purpose robot capable of performing various tasks through instruction. More specifically, robot 2 is a single-arm robot with various end effectors mounted on its arm. For example, robot 2 is a six-axis robot with an arm equipped with six link members and six joints. The main body and the six link members are connected by six joints.

[0068] The end effector is attached to the link member at the front end of the arm. By driving the six-axis arm, Robot 2 can move the end effector to any position and assume any posture. The end effector changes depending on the robot's action. For example, if Robot 2 is opening a door, the end effector is a member that grips an object. Alternatively, if Robot 2 is wiping a workbench, the end effector is a member such as a sponge or cloth.

[0069] An acceleration sensor is attached to the end effector in the link member at the front end of the arm. The acceleration sensor can acquire information about acceleration in three mutually perpendicular axes, as well as angular velocity about each axis. Based on this information, the robot 2 identifies the end effector's inclination, its movement speed, including its velocity and orientation, and its current position.

[0070] Robot 2 outputs trajectory information related to the trajectory of robot 2's movement to information processing device 1. Robot 2 outputs trajectory information related to the trajectory of the movement taught to robot 2 by a human to information processing device 1. The trajectory information represents, for example, time-series data representing the coordinates in three-dimensional space of an end effector attached to the tip of robot 2's arm.

[0071] The robot 2 is connected to the information processing device 1 via a wired or wireless connection so as to be able to communicate with each other. Alternatively, the robot 2 may be connected to the information processing device 1 via a network so as to be able to communicate with each other. The network may be a local area network or a wide area network.

[0072] The display unit 3 is, for example, a liquid crystal display device, and displays information output from the information processing device 1. The display unit 3 is connected to the information processing device 1 via a wired or wireless connection so that they can communicate with each other. Alternatively, the display unit 3 can be connected to the information processing device 1 via a network so that they can communicate with each other. The network can be a local area network or a wide area network.

[0073] The input unit 4 is, for example, a keyboard, mouse, or touch panel, and receives user input. The input unit 4 is connected to the information processing device 1 via a wired or wireless connection for communication. Alternatively, the input unit 4 may be connected to the information processing device 1 for communication via a network. The network may be a local area network or a wide area network.

[0074] The information processing device 1 includes a processor 11 and a memory 12. The information processing device 1 is, for example, a personal computer, a tablet computer, or a server.

[0075] The processor 11 is, for example, a CPU (Central Processing Unit), which implements a trajectory information acquisition unit 111 , an evaluation index acquisition unit 112 , an optimal solution calculation unit 113 , an optimal solution display control unit 114 , an optimal solution acquisition unit 115 , and a reference information display control unit 116 .

[0076] The memory 12 is a storage device capable of storing various information, such as a RAM (Random Access Memory), an HDD (Hard Disk Drive), an SSD (Solid State Drive), or a flash memory. The memory 12 stores various information.

[0077] The memory 12 stores the trajectory information output by the robot 2. The memory 12 stores a plurality of trajectory information corresponding to a plurality of motions of the robot 2.

[0078] Furthermore, the memory 12 may store a moving image obtained by recording the movement of the robot 2 together with trajectory information related to the movement trajectory of the robot 2. The moving image is acquired from a camera that records the movement of the robot 2.

[0079] The trajectory information acquisition unit 111 acquires trajectory information related to the trajectory of the operation of the robot 2 . The trajectory information acquisition unit 111 reads the trajectory information from the memory 12 .

[0080] In addition, in the present embodiment, the trajectory information acquisition unit 111 acquires trajectory information actually created by a person teaching the robot 2 an action, but the present disclosure is not particularly limited to this. In the case where it is not a person teaching the robot 2 an action but a computer graphic simulating the action of teaching the robot 2 in a virtual three-dimensional space, the trajectory information acquisition unit 111 may also acquire trajectory information created by simulation.

[0081] The input unit 4 receives user input of two or more evaluation indices for evaluating the operation of the robot 2. The display unit 3 displays a screen for receiving user input of the two or more evaluation indices. The user enters the two or more evaluation indices on the displayed screen. The input unit 4 outputs the two or more evaluation indices entered by the user to the information processing device 1.

[0082] The evaluation index acquisition unit 112 acquires two or more evaluation indexes input by the user from the input unit 4 .

[0083] The optimal solution calculation unit 113 adjusts the parameters of the robot 2 that optimize two or more evaluation indicators used to evaluate the movement of the robot 2 based on the trajectory information acquired by the trajectory information acquisition unit 111, thereby calculating multiple optimal solutions for the two or more evaluation indicators. The parameters are parameters used to control the robot 2.

[0084] The optimal solution calculation unit 113 includes a motion division unit 1131 and an optimization unit 1132 .

[0085] The motion segmentation unit 1131 divides the robot 2's motion into multiple segments. The motion is represented by multiple combinations of impedance-controlled equations of motion. Parameters include stiffness parameters for impedance control. The motion segmentation unit 1131 estimates the stiffness parameters within each equation of motion and the switching timing (segmentation timing) for each equation of motion to minimize the error between the predicted trajectory and the taught trajectory in each equation of motion.

[0086] For example, if the robot 2's motion involves opening a door, the motion division unit 1131 divides the motion into a first segment in which the end effector approaches the door handle, a second segment in which the end effector grasps and rotates the door handle, and a third segment in which the end effector opens the door. Alternatively, if the robot 2's motion involves wiping a workbench, the motion division unit 1131 divides the motion into a first segment in which the end effector approaches the workbench, a second segment in which the end effector contacts and stops at the workbench, a third segment in which the end effector wipes the workbench, and a fourth segment in which the end effector moves away from the workbench.

[0087] Here, the division process performed by the motion division unit 1131 will be described in more detail.

[0088] In impedance control, the robot 2 operates in accordance with the motion equation of the virtual spring, mass, and damper system represented by the following equation (1).

[0089] [Mathematical formula 1]

[0090]

[0091] Here, x∈R 6 represents the position and posture of the end effector in the task space, x d ∈R 6 represents the attractor, F∈R 6 represents the external force acting on the end effector. In addition, Λ∈R 6*6 is the inertia matrix, D∈R 6*6 is the attenuation matrix, K∈R 6*6 is the stiffness matrix. Through the spring behavior realized by the stiffness K term, the robot can flexibly respond to the unexpected external force F while following the attractor xd In order for the system to adapt, the stiffness K needs to be as low as possible.

[0092] The action segmentation unit 1131 divides the task into multiple segments and assigns a rigid parameter of a certain value to each segment, thereby reducing the input dimension of the Bayesian optimization. In the past, such segmentation processing was implemented manually, by clustering based on GMM (Gaussian Mixture Model), or by system identification based on SLD (Switching Linear Dynamics). In contrast, in this embodiment, a segmentation processing technology called IC-SLD (Impedance Control Aware Switching Linear Dynamics) is newly introduced to achieve segmentation suitable for impedance control. SLD is suitable for settings such as assigning certain parameters to segments (or switching rigid control). In IC-SLD, the impedance model (1) is embedded in advance to formulate the SLD identification problem. It is expected that by being aware of the formulation of impedance control, the segment suitable for switching rigid control to be executed in the subsequent optimization can be determined.

[0093] In IC-SLD, the above equation of motion is expressed by the following dynamics: In this case, the trajectory is assumed to follow the probabilistic linear dynamics of the following equation (2).

[0094] [Mathematical formula 2]

[0095]

[0096] In the above formula (2), the state vector x t Defined by the following formula (3), the action vector u t It is defined by the following formula (4). Here, Δx t is the residual with the attractor.

[0097] [Mathematical formula 3]

[0098]

[0099] In addition, p(x t+1 |·) is a linear Gaussian model represented by the following equation (5).

[0100] [Formula 4]

[0101]

[0102] s t ∈{1, 2, ..., M} is a latent variable representing the dynamics pattern, and M is equivalent to the number of divisions.j 、B j and Σ j are the parameters of the linear dynamics, which depend on the latent variable s t =j. In impedance control, the Euler method is used to discretize Equation (1), so that A j and B j It is expressed by the following formula (6) and formula (7).

[0103] [Formula 5]

[0104]

[0105] Here, the attenuation D is generally a value proportional to the square root of the rigidity K. Therefore, the attenuation D is set to D=2K 1 / 2 In equations (6) and (7), Δt represents the sampling period, I represents the identity matrix, and O represents the zero matrix.

[0106] A j It is called the state matrix, B j It is called the input matrix or control matrix. In this embodiment, A j and B j This is the matrix obtained by discretizing the equation of motion (mass, spring, and damping system) in impedance control with a time width of Δt using the Euler method.

[0107] IC-SLD for teaching trajectory x 1:T , determine the dynamic parameter A j 、B j and Σ j , and for the latent variable s t =j to perform inference (segmentation). This can be achieved by optimizing the objective functions shown in the following equations (8), (9), and (10) using the EM (Expectation Maximization) algorithm.

[0108] [Formula 6]

[0109]

[0110] The EM (Expectation Maximization) algorithm repeatedly performs the E step and the M step until convergence, thereby numerically solving the optimization problem. In this case, the E step calculates W by using a fixed dynamic parameter θ. j t , M steps by fixing W jt The dynamic parameter θ is updated by maximizing the equation (6). The motion segmentation unit 1131 estimates the dynamic parameter θ (A j 、B j and Σ j ) and hidden state S (hidden variables s t ).

[0111] The optimizer 1132 uses multi-objective Bayesian optimization to repeatedly search for optimal parameters for each segment, thereby optimizing two or more evaluation indicators. The optimizer 1132 uses the estimated rigidity parameters to weight the acquisition function in the multi-objective Bayesian optimization, and then repeatedly searches for the optimal rigidity parameters for each segment using the acquisition function.

[0112] Here, the optimization process performed by the optimization unit 1132 is described in more detail.

[0113] In the multi-objective optimization of this embodiment, two objective functions are optimized. Generally, these objective functions are in a trade-off relationship. As shown in the following formula (11), Bayesian optimization is repeated to obtain the function α(θ; D n ) to implement optimization by searching for the proposed solution candidates.

[0114] [Formula 7]

[0115]

[0116] Here, n is the number of repetitions, D n is the history of past evaluation results. In multi-objective optimization, there are multiple optimal solutions (Pareto optimal solutions) that exhibit the best trade-offs. n ) is expressed by the following formula (12). Let the set of Pareto optimal solutions observed before the number of iterations n be Y * 1:n When, in multi-objective Bayesian optimization, the acquisition function is defined so that Y n 1:n Hypervolume indicator I H (Y * 1:n ) is improved.

[0117] [Formula 8]

[0118]

[0119] I H (Y * 1:n) represents the area of ​​the region formed by the Pareto optimal solution. In previous Bayesian optimization, prior knowledge cannot be introduced except to narrow the search space. However, such a method of imparting prior knowledge may overlook important areas and fail to obtain the optimal solution. For this reason, the π-BO proposed in recent years proposes to introduce prior knowledge into the acquisition function in the form of probability distribution π(θ). The acquisition function α(θ; D) with the prior knowledge introduced n ) is expressed by the following formula (13).

[0120] [Formula 9]

[0121]

[0122] Here, β∈R + is a hyperparameter that reflects the reliability of π(θ). In the early stage of optimization, the acquisition function gives a larger weight to the prior distribution, but when n becomes larger, the exponential of the prior distribution gradually decays, and α π Progressive towards α.

[0123] The motion segmentation unit 1131 divides the motion trajectory into M segments and estimates the rigidity parameter K of each segment j∈{1, 2, ..., M}. j The optimization unit 1132 uses the rigidity parameter θ of each segment estimated by the motion segmentation unit 1131 as prior knowledge.

[0124] In this embodiment, the optimization unit 1132 optimizes the task performance evaluation index that evaluates the degree of task achievement and the safety evaluation index that indicates the safety of impedance control using multi-objective Bayesian optimization. This multi-objective optimization problem is formulated as shown in the following equation (14).

[0125] [Formula 10]

[0126]

[0127] In formula (14), J T (θ|S) is the objective function of the evaluation index of task performance, J C (θ|S) is the objective function of the safety evaluation index. Here, θ is the rigidity parameter K in each segment. j The set of (θ={K j} M j=1 ), S is the latent variable s t The set (segmentation result) (S = {s t =j} T t=1 ).

[0128] Objective function of task performance J T(θ|S) is defined by the reward function, etc. The objective function J of task performance T (θ|S) is expressed by the following formula (15).

[0129] [Mathematical formula 11]

[0130]

[0131] In formula (15), R is used to evaluate each state x t The task-specific reward function, state transition is determined by K s1:T Dominate.

[0132] In addition, the security objective function J C (θ|S) is defined by the time integral of the rigidity parameter. The safety objective function J C (θ|S) is expressed by the following formula (16).

[0133] [Mathematical formula 12]

[0134]

[0135] The objective function shown in Equation (16) sums up the rigid parameters in each time step (each segment). The objective function J of task performance T (θ|S) and the security objective function J C (θ|S) is strongly affected by the set S.

[0136] Furthermore, in this embodiment, the optimization unit 1132 utilizes the rigid parameters included in the dynamic parameters identified by IC-SLD as prior knowledge. The optimization unit 1132 receives the set S of segments determined by the motion segmentation unit 1131 and uses state-of-the-art Bayesian optimization techniques to perform multi-objective optimization of Equation (14) in a realistic environment.

[0137] The optimal solution display control unit 114 outputs a solution display image that plots the multiple optimal solutions calculated by the optimal solution calculation unit 113 on a plane or space having two or more evaluation indicators as coordinate axes. The optimal solution display control unit 114 creates a solution display image and outputs the created solution display image to the display unit 3.

[0138] The display unit 3 displays the solution display image output by the optimal solution display control unit 114. The input unit 4 receives a user's selection of at least one optimal solution from among the multiple optimal solutions displayed in the solution display image. The user selects at least one optimal solution from among the multiple optimal solutions displayed in the solution display image. The input unit 4 outputs the at least one optimal solution selected by the user to the information processing device 1.

[0139] The optimal solution acquisition unit 115 acquires at least one optimal solution selected by the user from among the plurality of optimal solutions displayed on the solution display image from the input unit 4 .

[0140] The reference information display control unit 116 outputs reference information based on the history of robot 2's movements corresponding to at least one optimal solution acquired by the optimal solution acquisition unit 115. The reference information display control unit 116 creates a reference information display screen for presenting the reference information and outputs the created reference information display screen to the display unit 3. The reference information includes time-series data of the trajectory corresponding to the at least one optimal solution.

[0141] The display unit 3 displays the reference information output by the reference information display control unit 116. The display unit 3 displays a reference information display screen for presenting the reference information.

[0142] Next, a teaching support process performed by the information processing device 1 in the embodiment of the present disclosure will be described.

[0143] Figure 2 This is a flowchart for explaining the teaching support process performed by the information processing device 1 in the embodiment of the present disclosure.

[0144] First, in step S1 , the trajectory information acquisition unit 111 acquires trajectory information regarding the trajectory of the operation of the robot 2 .

[0145] Next, in step S2, the input unit 4 receives input from the user of two or more evaluation indices for evaluating the operation of the robot 2. The user selects two evaluation indices from among the plurality of evaluation indices.

[0146] Next, in step S3 , the evaluation index acquisition unit 112 acquires two or more evaluation indexes input by the user from the input unit 4 . The evaluation index acquisition unit 112 acquires two evaluation indexes from the input unit 4 .

[0147] Figure 3 This is a diagram showing an example of a display screen for accepting input of two or more evaluation indices and displaying a plurality of optimal solutions in the present embodiment.

[0148] The display unit 3 displays a display screen for accepting input of two or more evaluation indices and displaying a plurality of optimal solutions. Figure 3 The display screen shown includes an evaluation index selection area 31 and an optimal solution display area 32 .

[0149] The evaluation index selection area 31 displays a drop-down list showing a plurality of selectable evaluation indexes and accepts the user's selection of two evaluation indexes. The user selects the desired two evaluation indexes from the drop-down list displayed in the evaluation index selection area 31. The input unit 4 accepts input of the evaluation index corresponding to the X-axis and the evaluation index corresponding to the Y-axis. Figure 3 In the example, task performance is selected as the evaluation index corresponding to the X-axis, and safety is selected as the evaluation index corresponding to the Y-axis.

[0150] Furthermore, in the present embodiment, two or more evaluation indices are selected by the user, but the present disclosure is not particularly limited thereto, and two or more evaluation indices to be optimized may be determined in advance.

[0151] Next, in step S4 , the motion division unit 1131 performs a division process of dividing the motion of the robot 2 into a plurality of sections.

[0152] Figure 4 Is used to illustrate Figure 2 Flowchart of the segmentation process in step S4.

[0153] First, in step S21, the motion segmentation unit 1131 initializes multiple segments based on a given number of divisions. For example, if the number of divisions is N, the motion segmentation unit 1131 initializes N-1 segment moments. While N equal divisions of the taught motion are considered as an initialization method, the present disclosure is not limited to this. The motion segmentation unit 1131 may also cluster the taught motions beforehand and perform initialization based on the clustering results.

[0154] Next, in step S22, the motion segmentation unit 1131 optimizes the rigid parameters to minimize the error between the predicted start and the taught start for each segment. The optimization method used is Newton's method, but this disclosure is not particularly limited to this method. Furthermore, in addition to the error, the motion segmentation unit 1131 may also add a regularization term to the objective function to prevent the rigid parameters from becoming excessively large.

[0155] Next, in step S23, the motion segmentation unit 1131 determines whether a termination condition is satisfied. The termination condition is that the number of repetitions is greater than or equal to a threshold or the error is less than or equal to a threshold. If the termination condition is satisfied (yes in step S23), the segmentation process ends.

[0156] On the other hand, if the termination condition is determined not to be met (No in step S23), the action segmentation unit 1131 determines in step S24 whether a segment that further reduces the error can be inferred. If the optimization result in step S22 can be differentiated at the split time, the split time can be inferred based on gradient information. Furthermore, if the optimization result cannot be differentiated at the split time, inference based on evolutionary computation can also be performed. In this case, in step S22, the action segmentation unit 1131 may generate multiple segment candidates and calculate the error for each segment candidate. Furthermore, in step S24, the action segmentation unit 1131 may perform inference based on statistical information of the calculated errors. One method of generating multiple segment candidates is to apply a perturbation based on a normal random number to the split time.

[0157] Here, when it is determined that a segment that further reduces the error cannot be estimated (No in step S24 ), the division process ends.

[0158] On the other hand, when it is determined that a segment that further reduces the error can be estimated (YES in step S24 ), the process returns to step S22 .

[0159] Return to Figure 2 Next, in step S5, the optimization unit 1132 optimizes two or more evaluation indicators through multi-objective Bayesian optimization. The optimization unit 1132 calculates multiple optimal solutions by optimizing the two evaluation indicators.

[0160] Next, in step S6 , the optimal solution display control unit 114 generates a solution display image that plots the multiple optimal solutions calculated by the optimization unit 1132 on a plane or space having two or more evaluation indicators as coordinate axes, and outputs the generated solution display image to the display unit 3 .

[0161] Next, in step S7 , the display unit 3 displays the solution display image output by the optimal solution display control unit 114 .

[0162] like Figure 3 As shown, the display unit 3 displays a solution display image in the optimal solution display area 32. The X-axis of the solution display image represents the evaluation index of task performance, and the Y-axis represents the evaluation index of safety. Multiple optimal solutions 321 are solutions that are not surpassed by any feasible solution, also known as Pareto optimal solutions. Figure 3In the figure, multiple optimal solutions 321 are represented by shaded dots, and multiple feasible solutions other than the optimal solutions 321 are represented by white dots. When two evaluation indicators are maximized, the multiple optimal solutions 321 are arranged in the upper right corner of a plane whose coordinate axes are the two evaluation indicators. The Pareto optimal solution and the multiple feasible solutions other than the Pareto optimal solution are displayed in different ways. This allows the user to easily visually identify the Pareto optimal solution.

[0163] In addition, in this embodiment, two evaluation indicators are selected and multiple optimal solutions for the two evaluation indicators are calculated. However, the present disclosure is not particularly limited to this. Three or more evaluation indicators may be selected and multiple optimal solutions for the three or more evaluation indicators may be calculated. For example, when three evaluation indicators are selected and multiple optimal solutions for the three evaluation indicators are calculated, the optimal solution display control unit 114 may output a solution display image that depicts the multiple calculated optimal solutions in a three-dimensional space with the three evaluation indicators as coordinate axes.

[0164] In addition, in this embodiment, multiple optimal solutions are displayed after the optimization of two or more evaluation indicators is completed. However, the present disclosure is not particularly limited to this. Before the optimization of two or more evaluation indicators is completed, multiple optimal solutions can be displayed sequentially while optimizing two or more evaluation indicators. In this case, the user can confirm the progress of the optimization.

[0165] Return to Figure 2 Next, in step S8 , the input unit 4 receives the user's selection of one optimal solution from among the plurality of optimal solutions displayed on the display image.

[0166] like Figure 3 As shown, the display unit 3 displays a pointer 33 that can be moved with a mouse. The user moves the pointer 33 to one of the multiple optimal solutions 321 displayed and clicks the mouse button. This selects one of the multiple optimal solutions 321. The input unit 4 outputs information to the information processing device 1 indicating that the one optimal solution selected by the user is one of the multiple optimal solutions displayed in the solution display image.

[0167] Return to Figure 2 Next, in step S9 , the optimal solution acquisition unit 115 acquires one optimal solution selected by the user from among the multiple optimal solutions displayed on the solution display image from the input unit 4 .

[0168] Next, in step S10 , the reference information display control unit 116 outputs reference information based on the operation history of the robot 2 corresponding to one optimal solution acquired by the optimal solution acquisition unit 115 to the display unit 3 .

[0169] Next, in step S11, the display unit 3 displays the reference information output by the reference information display control unit 116. The display unit 3 displays a reference information display screen for presenting the reference information.

[0170] The user confirms the reference information of each of the plurality of optimal solutions and selects a desired optimal solution from the plurality of optimal solutions. Parameters corresponding to the optimal solution selected by the user are used to control the robot 2 .

[0171] Next, in step S12, the input unit 4 determines whether to accept the reselection of one optimal solution from among the multiple optimal solutions. The reference information display screen displayed on the display unit 3 may also include a reselection button for accepting the reselection of one optimal solution from among the multiple optimal solutions. When the reselection button displayed on the reference information display screen is pressed, the input unit 4 may also determine that the reselection of one optimal solution from among the multiple optimal solutions is accepted. In addition, the reference information display screen displayed on the display unit 3 may also include an end button for ending the selection of one optimal solution from among the multiple optimal solutions. When the end button displayed on the reference information display screen is pressed, the input unit 4 may also determine that the reselection of one optimal solution from among the multiple optimal solutions is not accepted.

[0172] Here, when it is determined that reselection of one optimal solution among the plurality of optimal solutions is accepted (YES in step S12 ), the process returns to step S7 .

[0173] On the other hand, when it is determined that reselection of one optimal solution among the plurality of optimal solutions is not accepted (No in step S12 ), the teaching support process ends.

[0174] According to this embodiment, two or more evaluation indicators can be optimized, and multiple optimal solutions for the two or more evaluation indicators can be calculated. In addition, since reference information based on the movement history of robot 2 corresponding to at least one optimal solution selected by the user from among the multiple optimal solutions is presented to the user, the user can be assisted in selecting the optimal solution from among the multiple optimal solutions.

[0175] Figure 5 FIG. 1 is a diagram showing an example of a reference information display screen displayed on the display unit 3 in this embodiment. Figure 5 In the figure, the horizontal axis represents time and the vertical axis represents coordinates.

[0176] Figure 5The reference information shown includes time-series data of the trajectory corresponding to one optimal solution. In this case, display unit 3 displays the time-series data of the trajectory corresponding to one optimal solution. Specifically, display unit 3 displays the time-series data of the x-coordinate, y-coordinate, and z-coordinate of robot 2's end effector in three-dimensional space corresponding to one optimal solution. The time-series trajectory data represents the trajectory of robot 2's door-opening action.

[0177] In addition, Figure 5 In the figure, the x-coordinate position is represented by a solid line, the y-coordinate position is represented by a dotted line, and the z-coordinate position is represented by a dashed line. Thus, the time series data of the x-coordinate, y-coordinate, and z-coordinate are represented by different types of lines, but the present disclosure is not particularly limited to this, and lines of different colors may also be used.

[0178] In this embodiment, one optimal solution is selected from among the multiple optimal solutions displayed in the solution display image. However, the present disclosure is not particularly limited to this. Alternatively, one feasible solution may be selected from among multiple feasible solutions other than the multiple optimal solutions displayed in the solution display image. In this case, reference information based on the history of robot 2's movements corresponding to the one feasible solution is output. This allows the user to view reference information not only for multiple optimal solutions but also for multiple feasible solutions other than the multiple optimal solutions.

[0179] In addition, in this embodiment, one optimal solution is selected from a plurality of optimal solutions. However, the present disclosure is not particularly limited to this. In Variation 1 of this embodiment, two or more optimal solutions may be selected from a plurality of optimal solutions. When two or more optimal solutions are selected from a plurality of optimal solutions, two or more time series data corresponding to the two or more optimal solutions may be displayed in an overlapping or aligned manner.

[0180] Figure 6 FIG. 1 is a diagram showing an example of a reference information display screen displayed on the display unit 3 in the first modification of the present embodiment. Figure 6 In the figure, the horizontal axis represents time and the vertical axis represents coordinates.

[0181] In Variation 1 of this embodiment, the input unit 4 receives a user's selection of two optimal solutions from among the multiple optimal solutions displayed on the solution display image. The optimal solution acquisition unit 115 acquires the two optimal solutions selected by the user from the multiple optimal solutions displayed on the solution display image from the input unit 4. The reference information display control unit 116 reads two time-series data corresponding to the two optimal solutions acquired by the optimal solution acquisition unit 115 from the memory 12 and outputs reference information superimposed on the two time-series data to the display unit 3. The display unit 3 displays the reference information output by the reference information display control unit 116.

[0182] Figure 6 The reference information display screen shown displays two overlapping time-series data for the two trajectories corresponding to the two optimal solutions. Specifically, the display unit 3 displays two overlapping time-series data for the x-, y-, and z-coordinates of the robot 2's end effector in three-dimensional space, corresponding to the two optimal solutions. In this case, the display unit 3 displays the first time-series data for the trajectory corresponding to the first optimal solution and the second time-series data for the trajectory corresponding to the second optimal solution, overlapping each other.

[0183] Furthermore, the two time series data corresponding to the two optimal solutions are represented by lines of different thicknesses, but the present disclosure is not particularly limited to this. Lines of different colors may also be used. For example, the time series data corresponding to one optimal solution may be represented by a red line, and the time series data corresponding to the other optimal solution may be represented by a blue line.

[0184] Alternatively, the reference information display screen may display two time series data items corresponding to the two optimal solutions in parallel. In this case, the display unit 3 may also display the first time series data item corresponding to the first optimal solution and the second time series data item corresponding to the second optimal solution in parallel.

[0185] In Variation 1 of this embodiment, at least one time series data item corresponding to at least one optimal solution and time series data item of the robot's target motion trajectory may be displayed superimposed or aligned. The robot's target motion may be taught by a user, for example. Memory 12 may also store time series data item of the robot's target motion trajectory.

[0186] In addition, the reference information in this embodiment includes time series data of the trajectory corresponding to at least one optimal solution, but the present disclosure is not particularly limited to this. The reference information in variant example 2 of this embodiment may also include a dynamic image obtained by recording the action corresponding to at least one optimal solution.

[0187] Figure 7 This is a diagram showing an example of a reference information display screen displayed on the display unit 3 in Modification 2 of the present embodiment.

[0188] Figure 7 The reference information shown includes a dynamic image 301 obtained by recording an action corresponding to one optimal solution. In this case, the display unit 3 displays the dynamic image 301 obtained by recording an action corresponding to one optimal solution. Figure 7, an image at which time 0 seconds have passed, an image at which time 10 seconds have passed, and an image at which time 20 seconds have passed are cut out and shown from the dynamic image 301. The dynamic image 301 shows the action of the robot 2 opening the door.

[0189] In the third modification of the present embodiment, when two or more optimal solutions are selected from a plurality of optimal solutions, two or more moving images obtained by recording actions corresponding to the two or more optimal solutions may be displayed side by side.

[0190] Figure 8 This is a diagram showing an example of a reference information display screen displayed on the display unit 3 in Modification 3 of the present embodiment.

[0191] Figure 8 The reference information shown includes a first moving image 302 obtained by recording an action corresponding to the first optimal solution selected by the user, and a second moving image 303 obtained by recording an action corresponding to the second optimal solution selected by the user. Figure 8 The reference information display screen shown shows two dynamic images obtained by recording two actions corresponding to two optimal solutions in a row. In this case, the display unit 3 displays a first dynamic image 302 obtained by recording the action corresponding to the first optimal solution and a second dynamic image 303 obtained by recording the action corresponding to the second optimal solution in a row. Figure 8 , an image at which time 0 seconds have passed, an image at which time 10 seconds have passed, and an image at which time 20 seconds have passed are cut out and shown from the first dynamic image 302 and the second dynamic image 303. The first dynamic image 302 and the second dynamic image 303 show the action of the robot 2 opening the door.

[0192] In Modification 4 of the present embodiment, when two or more optimal solutions are selected from a plurality of optimal solutions, two or more moving images obtained by recording actions corresponding to the two or more optimal solutions may be displayed superimposed on each other.

[0193] Figure 9 This is a diagram showing an example of a reference information display screen displayed on the display unit 3 in Modification 4 of the present embodiment.

[0194] Figure 9 The reference information shown includes a composite dynamic image 304, which is formed by overlapping a first dynamic image 302 obtained by recording the action corresponding to the first optimal solution selected by the user and a second dynamic image 303 obtained by recording the action corresponding to the second optimal solution selected by the user. Figure 9The reference information display screen shown will display two overlapping dynamic images obtained by recording two actions corresponding to the two optimal solutions. In this case, the display unit 3 displays a composite dynamic image 304 formed by overlapping a first dynamic image 302 obtained by recording the action corresponding to the first optimal solution and a second dynamic image 303 obtained by recording the action corresponding to the second optimal solution. For example, the reference information display control unit 116 can also create a composite dynamic image 304 by overlapping a semi-transparent second dynamic image 303 on an opaque first dynamic image 302. In addition, Figure 9 In FIG, the first dynamic image 302 is represented by a dotted line, and the second dynamic image 303 is represented by a solid line. Figure 9 , an image at which time 0 seconds have passed, an image at which time 10 seconds have passed, and an image at which time 20 seconds have passed are cut out and shown from the synthesized dynamic image 304. The first dynamic image 302 and the second dynamic image 303 show the action of the robot 2 opening the door.

[0195] By displaying two or more moving images in a superimposed manner, the user can intuitively grasp the difference between two or more robot movements.

[0196] Furthermore, in Variations 3 and 4 of this embodiment, at least one dynamic image corresponding to at least one optimal solution and a dynamic image obtained by recording the robot's target motion may be superimposed or displayed side by side. The robot's target motion may be taught by a user, for example. Memory 12 may also store the dynamic image obtained by recording the robot's target motion.

[0197] In a fifth variation of the present embodiment, at least one time series data of a trajectory corresponding to at least one optimal solution and time series data of at least one parameter that changes according to an action corresponding to the at least one optimal solution may be displayed superimposed on each other.

[0198] Figure 10 FIG. 1 is a diagram showing an example of a reference information display screen displayed on the display unit 3 in the fifth modification of the present embodiment. Figure 10 In the graph, the horizontal axis represents time, the vertical axis on the left represents coordinates, and the vertical axis on the right represents the value of the parameter.

[0199] In Variation 5 of this embodiment, the input unit 4 receives a user's selection of one optimal solution from among the multiple optimal solutions displayed in the solution display image. The optimal solution acquisition unit 115 acquires the optimal solution selected by the user from among the multiple optimal solutions displayed in the solution display image from the input unit 4. The reference information display control unit 116 reads from the memory 12 the time-series data of the trajectory corresponding to the acquired optimal solution and the time-series data of the parameters that change according to the action corresponding to the acquired optimal solution, and outputs reference information, which is a superposition of the time-series data of the trajectory and the time-series data of the parameters, to the display unit 3. The display unit 3 displays the reference information output by the reference information display control unit 116.

[0200] Figure 10 The reference information display screen shown superimposes time-series data of the trajectory corresponding to a single optimal solution and time-series data of parameters that change according to the action corresponding to the single optimal solution. Specifically, the display unit 3 superimposes the time-series data of the x-, y-, and z-coordinates of the robot 2's end effector in three-dimensional space corresponding to the single optimal solution and the time-series data of the parameters that change according to the action corresponding to the single optimal solution. The parameters are, for example, stiffness parameters in impedance control. The time-series trajectory data shows the trajectory of the robot 2's door opening action.

[0201] In addition, Figure 10 In the figure, the x-coordinate position is represented by a solid line, the y-coordinate position is represented by a dotted line, the z-coordinate position is represented by a dashed line, and the parameter value is represented by a double-dashed line. Thus, the time series data of the x-coordinate, y-coordinate, and z-coordinate and the time series data of the parameter are represented by different types of lines, but the present disclosure is not particularly limited to this, and lines of different colors may also be used.

[0202] Alternatively, the reference information display screen may present at least one time series data item of a trajectory corresponding to at least one optimal solution and at least one time series data item of a parameter that changes according to an action corresponding to the at least one optimal solution, arranged in a row. In this case, the display unit 3 may also display at least one time series data item of a trajectory corresponding to at least one optimal solution and at least one time series data item of a parameter that changes according to an action corresponding to the at least one optimal solution, arranged in a row.

[0203] In a sixth variation of the present embodiment, at least one moving image corresponding to at least one optimal solution and time-series data of at least one parameter that changes according to an action corresponding to the at least one optimal solution may be displayed side by side.

[0204] Figure 11This is a diagram showing an example of a reference information display screen displayed on the display unit 3 in Modification 6 of the present embodiment.

[0205] In Variation 6 of this embodiment, the input unit 4 receives a user's selection of one optimal solution from among the multiple optimal solutions displayed on the solution display image. The optimal solution acquisition unit 115 acquires the user-selected optimal solution from the multiple optimal solutions displayed on the solution display image from the input unit 4. The reference information display control unit 116 reads from the memory 12 a moving image corresponding to the acquired optimal solution and time-series data of parameters that change according to the action corresponding to the acquired optimal solution, and outputs reference information in which the moving image and the time-series data of the parameters are arranged to the display unit 3. The display unit 3 displays the reference information output by the reference information display control unit 116.

[0206] Figure 11 The reference information shown includes a moving image 301 obtained by recording a motion corresponding to one acquired optimal solution, and time series data 306 of parameters that change according to the motion corresponding to one acquired optimal solution. Figure 11 The reference information display screen shown shows a dynamic image corresponding to one optimal solution and time series data of parameters that change according to the action corresponding to one optimal solution, arranged in a row. That is, the display unit 3 shows a dynamic image 301 obtained by recording the action corresponding to one optimal solution and time series data 306 of parameters that change according to the action corresponding to one optimal solution in a row. The parameter is, for example, a stiffness parameter in impedance control. Figure 11 , an image at which time 0 seconds have passed, an image at which time 10 seconds have passed, and an image at which time 20 seconds have passed are cut out and shown from the dynamic image 301. The dynamic image 301 shows the action of the robot 2 opening the door.

[0207] The time series data 306 of the parameter is represented by an analog indicator with a needle moving on a circular dial. The value of the parameter is displayed around the dial, and the needle on the dial moves according to the change of the parameter value.

[0208] In addition, at least one dynamic image corresponding to at least one optimal solution and time series data of at least one parameter that changes according to the action corresponding to at least one optimal solution may be displayed in an overlapping manner. The reference information display screen may also display at least one dynamic image corresponding to at least one optimal solution and time series data of at least one parameter that changes according to the action corresponding to at least one optimal solution in an overlapping manner. In this case, the display unit 3 may also display at least one dynamic image corresponding to at least one optimal solution and time series data of at least one parameter that changes according to the action corresponding to at least one optimal solution in an overlapping manner. For example, the display unit 3 may also display a dynamic image and display the time series data of the parameter in the lower right portion of the dynamic image.

[0209] In addition, the time series data of the parameters can also be expressed by numerical values.

[0210] In addition, in variant example 7 of the present embodiment, a synthesized dynamic image and synthesized time series data can also be arranged and displayed. The synthesized dynamic image is formed by overlapping two or more dynamic images obtained by recording actions corresponding to two or more optimal solutions, and the synthesized time series data is formed by overlapping two or more time series data of two or more parameters that change according to the actions corresponding to two or more optimal solutions.

[0211] Figure 12 This is a diagram showing an example of a reference information display screen displayed on the display unit 3 in Modification 7 of the present embodiment.

[0212] In Variation 7 of this embodiment, the input unit 4 receives a user's selection of two optimal solutions from among the multiple optimal solutions displayed on the solution display image. The optimal solution acquisition unit 115 acquires the two optimal solutions selected by the user from among the multiple optimal solutions displayed on the solution display image from the input unit 4. The reference information display control unit 116 reads from the memory 12 two dynamic images corresponding to the two acquired optimal solutions and two time-series data of two parameters that change according to the actions corresponding to the two acquired optimal solutions. The reference information display control unit 116 then outputs reference information to the display unit 3, which arranges a composite dynamic image 304 formed by superimposing the two dynamic images and composite time-series data 308 formed by superimposing the two time-series data of the two parameters. The display unit 3 displays the reference information output by the reference information display control unit 116.

[0213] Figure 12The reference information shown includes: a synthesized dynamic image 304, which is formed by overlapping the first dynamic image 302 obtained by recording the action corresponding to the first optimal solution and the second dynamic image 303 obtained by recording the action corresponding to the second optimal solution; and synthesized time series data 308, which is formed by overlapping two time series data of two parameters that change according to the actions corresponding to the two optimal solutions obtained. Figure 12 The reference information display screen shown presents a composite dynamic image and composite time series data arranged in an array. The composite dynamic image is formed by overlapping two dynamic images obtained by recording two actions corresponding to the two optimal solutions, and the composite time series data is formed by overlapping two time series data of two parameters that change according to the actions corresponding to the two optimal solutions.

[0214] In this case, the display unit 3 displays a synthesized dynamic image 304 and synthesized time series data 308 in an arranged manner. The synthesized dynamic image 304 is formed by overlapping a first dynamic image 302 obtained by recording an action corresponding to the first optimal solution selected by the user and a second dynamic image 303 obtained by recording an action corresponding to the second optimal solution selected by the user. The synthesized time series data 308 is formed by overlapping two time series data of two parameters that change according to the actions corresponding to the two optimal solutions. For example, the reference information display control unit 116 may also create a synthesized dynamic image 304 by overlapping a semi-transparent second dynamic image 303 on an opaque first dynamic image 302. In addition, in Figure 12 In FIG, the first dynamic image 302 is represented by a dotted line, and the second dynamic image 303 is represented by a solid line. Figure 12 , an image at which time 0 seconds have passed, an image at which time 10 seconds have passed, and an image at which time 20 seconds have passed are cut out and shown from the synthesized dynamic image 304. The first dynamic image 302 and the second dynamic image 303 show the action of the robot 2 opening the door.

[0215] In addition, the parameter is, for example, a stiffness parameter in impedance control. The synthetic time series data 308 of the parameter is represented by an analog indicator in which a needle moves on a circular dial. The value of the parameter is displayed around the dial, and the needle on the dial moves according to the change of the parameter value. Figure 12 In FIG, the first time series data of the parameter that changes according to the action corresponding to the first optimal solution is represented by a dotted line, and the second time series data of the parameter that changes according to the action corresponding to the second optimal solution is represented by a solid line.

[0216] Alternatively, the synthesized dynamic image and the synthesized time-series data of the parameters may be displayed overlappingly. The reference information display screen may also display the synthesized dynamic image and the synthesized time-series data of the parameters overlappingly. In this case, the display unit 3 may also display the synthesized dynamic image and the synthesized time-series data of the parameters overlappingly. For example, the display unit 3 may display the synthesized dynamic image and display the synthesized time-series data of the parameters in the lower right portion of the synthesized dynamic image.

[0217] Alternatively, two time series data of parameters may be represented by numerical values.

[0218] Next, simulation results of the segmentation process and the optimization process in this embodiment will be described.

[0219] The simulation used two simulated tasks and a task using the actual robot. The simulated tasks included wiping the workbench and opening the door. In addition, the actual robot performed the wiping action in a real environment. The objective functions of these tasks were set as the task-specific reward function R(x t ). The reward for a simulated task is a binary variable representing success or failure, with "1" indicating that the current state is task completion (i.e., dirt is cleaned or the door is open). The actual task reward is represented by the negative mean squared error between the realized position trajectory and the verified position trajectory.

[0220] As the conventional segmentation methods for comparison, the conventional GMM method and the conventional SLD method which is not aware of impedance control are adopted. j It is expressed by the following formula (17): j It is expressed by the following equation (18). In this formulation, the dynamic parameters and segment identifiers are estimated by optimizing equation (8) using the EM algorithm.

[0221] A j =diag(a1,a2,…,a6)∈R 6*6 …(17)

[0222] B j =(diag(b1,b2,…,b6),diag(b’1,b’2,…,b’6))∈R 6*12 …(18)

[0223] In the conventional GMM method and the conventional SLD method, π-BO is not applied to the Bayesian optimization process unless otherwise specified.

[0224] The effectiveness of the proposed method in this embodiment was investigated through simulation. For each setting, 10 simulations were performed using different random seeds, and the statistical (average) results were compared.

[0225] Figure 13 : is a graph showing the comparison results of learning curves representing the growth of the area of ​​the Pareto optimal solution in the wiping operation. Figure 14 This is a diagram showing the comparison results of learning curves representing the growth of the area of ​​the Pareto optimal solution in the door opening action.

[0226] exist Figure 13 as well as Figure 14 In FIG, the solid line represents the simulation result based on the IC-SLD method of this embodiment, the dotted line represents the simulation result based on the existing GMM method, and the single-dot chain line represents the simulation result based on the existing SLD method. Figure 13 as well as Figure 14 In the example, the horizontal axis represents the number of trials, and the vertical axis represents the area of ​​the Pareto optimal solution (Hypervolumeindicator: I H (Y)).

[0227] Figure 13 as well as Figure 14 The progress of the optimization obtained by the IC-SLD method of this embodiment, the conventional GMM method, and the conventional SLD method is shown. In addition, the number of divisions M of the wiping action is set to 2, and the number of divisions of the door opening action is set to 3. In addition, in the optimization process of this embodiment, the hyperparameter β of π-BO is set to 1. Figure 13 as well as Figure 14 As shown, the IC-SLD method of this embodiment most effectively optimizes the two evaluation indicators and reaches convergence after about 100 trials.

[0228] Figure 15 This is a diagram showing the comparison results of the areas of the Pareto optimal solutions when prior knowledge is used in the IC-SLD method, the GMM method, and the SLD method, and the areas of the Pareto optimal solutions when no prior knowledge is used.

[0229] To determine whether prior knowledge of π-BO improves optimization, an ablation analysis was performed. The ablation analysis segmented the wiping action and the door opening action for the IC-SLD, GMM, and SLD methods. Furthermore, the two evaluation metrics were optimized using prior knowledge of π-BO for the IC-SLD, GMM, and SLD methods, respectively. The two evaluation metrics were optimized without prior knowledge of π-BO for the IC-SLD, GMM, and SLD methods.

[0230] also, Figure 15 The values ​​represent the average of the results of multiple experiments, and the values ​​after ± represent the error relative to the average.

[0231] like Figure 15 As shown, performance is best improved when the actions are segmented using the IC-SLD method and the two evaluation metrics are optimized using the prior knowledge of π-BO. This result indicates that both the IC-SLD method and the prior knowledge of π-BO contribute to performance improvement.

[0232] In addition, the parameter optimized in this embodiment is a stiffness parameter in impedance control, but the present disclosure is not particularly limited to this. It may also be a damping parameter in impedance control. Alternatively, the parameter may be the P gain, I gain, or D gain in PID (Proportional Integral Differential) control. Furthermore, the parameter may be the force control selection coefficient or the position control selection coefficient in hybrid control of force and position control.

[0233] Furthermore, in each of the above-described embodiments, each component may be configured by dedicated hardware and implemented by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Alternatively, the program may be implemented by another independent computer system by recording the program on a recording medium and transferring it, or by transferring the program via a network.

[0234] Part or all of the functions of the devices involved in the embodiments of the present disclosure are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These may be implemented individually or in part as a single chip. Furthermore, integrated circuitry is not limited to LSIs and can also be implemented using dedicated circuits or general-purpose processors. Alternatively, an FPGA (Field Programmable Gate Array) that can be programmed after LSI fabrication, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells within the LSI, can be utilized.

[0235] Furthermore, a part or all of the functions of the apparatus according to the embodiments of the present disclosure may be realized by executing a program by a processor such as a CPU.

[0236] In addition, the numbers used above are all exemplified to specifically describe the present disclosure, and the present disclosure is not limited to the exemplified numbers.

[0237] The order in which the steps are executed in the flowcharts is provided for the purpose of illustrating the present disclosure in detail, and any other order may be employed while achieving the same effect. Furthermore, some of the steps may be executed simultaneously (in parallel) with other steps.

[0238] Industrial applicability

[0239] The technology involved in the present disclosure can optimize more than two evaluation indicators and can assist users in selecting one optimal solution from multiple optimal solutions. Therefore, it is useful as a technology for calculating multiple optimal solutions for more than two evaluation indicators used to evaluate the actions of a robot and prompting the calculated multiple optimal solutions.

Claims

1. An information processing method, executed by a computer, The information processing method includes: Acquiring trajectory information related to the trajectory of the robot's movement; adjusting parameters of the robot that optimize two or more evaluation indicators for evaluating the movement of the robot based on the trajectory information, thereby calculating a plurality of optimal solutions for the two or more evaluation indicators; outputting a solution display image in which the calculated plurality of optimal solutions are plotted on a plane or space having the two or more evaluation indicators as coordinate axes; acquiring at least one optimal solution selected by a user from among the plurality of optimal solutions displayed in the solution display image; and Reference information based on the history of the movement of the robot corresponding to the at least one acquired optimal solution is output.

2. The information processing method according to claim 1, wherein: The information processing method further includes accepting input of the two or more evaluation indicators by the user.

3. The information processing method according to claim 1, wherein: The calculation of the multiple optimal solutions includes: dividing the motion of the robot into a plurality of segments; and The two or more evaluation indicators are optimized by repeatedly searching for the optimal parameters for each segment through multi-objective Bayesian optimization.

4. The information processing method according to claim 3, wherein: The motion is represented by multiple combinations of impedance-controlled motion equations, The parameters include stiffness parameters of the impedance control, The segmentation of the motion includes estimating the rigidity parameters in each motion equation and the switching timing of each motion equation so as to minimize the error between the predicted trajectory and the taught trajectory in each motion equation.

5. The information processing method according to claim 4, wherein: The optimization of the two or more evaluation indicators includes weighting an acquisition function in the multi-objective Bayesian optimization using the estimated rigid parameters, and repeatedly searching for the optimal rigid parameters for each segment using the acquisition function. The information processing method according to claim 1 , wherein: The reference information includes at least one time series data of the trajectory corresponding to the at least one optimal solution.

7. The information processing method according to claim 6, wherein: When two or more optimal solutions are selected from the plurality of optimal solutions, two or more time-series data corresponding to the two or more optimal solutions are overlapped or aligned and displayed.

8. The information processing method according to claim 6, wherein: The at least one time series data corresponding to the at least one optimal solution and the time series data of the trajectory of the target movement of the robot are overlapped or aligned and displayed.

9. The information processing method according to claim 6, wherein: The at least one time-series data corresponding to the at least one optimal solution and the time-series data of at least one parameter that changes according to the action corresponding to the at least one optimal solution are overlapped or displayed in parallel.

10. The information processing method according to claim 1, wherein: The reference information includes at least one dynamic image obtained by recording the action corresponding to the at least one optimal solution.

11. The information processing method according to claim 10, wherein: When two or more optimal solutions among the plurality of optimal solutions are selected, two or more moving images corresponding to the two or more optimal solutions are superimposed or displayed side by side.

12. The information processing method according to claim 10, wherein: The at least one dynamic image corresponding to the at least one optimal solution and a dynamic image obtained by recording the target movement of the robot are superimposed or displayed side by side.

13. The information processing method according to claim 10, wherein: The at least one dynamic image corresponding to the at least one optimal solution and time series data of at least one parameter that changes according to the action corresponding to the at least one optimal solution are arranged and displayed.

14. The information processing method according to claim 1, wherein: The two or more evaluation indicators include an evaluation indicator of task performance and an evaluation indicator of safety.

15. An information processing device comprising: a trajectory information acquisition unit that acquires trajectory information related to the trajectory of the robot's movement; a calculation unit that adjusts parameters of the robot for optimizing two or more evaluation indices for evaluating the motion of the robot based on the trajectory information, thereby calculating a plurality of optimal solutions for the two or more evaluation indices; a first output unit configured to output a solution display image in which the calculated plurality of optimal solutions are plotted on a plane or space having the two or more evaluation indicators as coordinate axes; an optimal solution acquiring unit configured to acquire at least one optimal solution selected by a user from among the plurality of optimal solutions displayed in the solution display image; and The second output unit outputs reference information based on the history of the movement of the robot corresponding to the at least one acquired optimal solution.

16. An information processing program that enables a computer to: Obtain trajectory information related to the trajectory of the robot's movement, Based on the trajectory information, the parameters of the robot are adjusted to optimize two or more evaluation indicators for evaluating the movement of the robot, thereby calculating a plurality of optimal solutions for the two or more evaluation indicators. outputting a solution display image in which the calculated plurality of optimal solutions are plotted on a plane or space having the two or more evaluation indices as coordinate axes, acquiring at least one optimal solution selected by the user from among the plurality of optimal solutions displayed in the solution display image, Reference information based on the history of the movement of the robot corresponding to the at least one acquired optimal solution is output.

Citation Information

Patent Citations

  • Method, program and information processing unit for assisting in adjusting parameter set of robot

    JP2022070451A