Method for setting force control parameters during operation of a robot and robot system

By setting limit values ​​and objective functions in robot operations, using the objective function to search for the optimal value of force control parameters, and introducing a penalty mechanism, it solves the problem of inaccurate force control parameters caused by factors such as sensor noise, and achieves a more stable and accurate robot operation.

CN115592661BActive Publication Date: 2025-06-17SEIKO EPSON CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210723911.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-06-28
Filing Date
2022-06-24
Publication Date
2025-06-17
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

In the prior art, due to factors such as sensor noise, workpiece shape and position deviation in robot operations, the force control parameters are inaccurately set, and the possibility of load exceeding the set range increases.

Method used

By setting the limit value and the objective function, the objective function is used to search for the optimal value of the force control parameter, and the set value of the force control parameter is determined based on the search results. A penalty mechanism is introduced in the objective function to punish small allowable values ​​whose force control characteristic values ​​exceed the limit value to ensure the reasonable setting of force control parameters.

Benefits of technology

Even if there is a sensor output deviation, it can effectively reduce the possibility that the force control characteristic value exceeds the limit value, ensure the appropriate setting of force control parameters, and improve the stability and accuracy of robot operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115592661B_ABST
    Figure CN115592661B_ABST
Patent Text Reader

Abstract

The present invention provides a method for setting force control parameters during the operation of a robot and a robot system, which can set the force control parameters to appropriate values even if there are deviations in sensor outputs. The method of the present disclosure includes: step (a) of setting a limit value and an objective function, where the limit value defines a limit condition for a specific force control characteristic value detected in force control, and the objective function is related to a specific evaluation item associated with the operation; step (b) of searching for an optimal value of the force control parameter using the objective function; and step (c) of determining a set value of the force control parameter based on the search result. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to the measured value of the evaluation item.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a method for setting force control parameters in the operation of a robot, a robot system, and a computer program. Background Art

[0002] Patent Document 1 discloses a system for adjusting parameters of a robot. Specifically, first, force data applied to a robot hand and control instruction adjustment data are generated as state data, and determination data showing a determination result of the operation state of the robot hand after the adjustment action is generated. Further, using these state data and determination data, a learning model obtained by performing reinforcement learning on an adjustment action of a control instruction for the state of the force applied to the robot hand is generated.

[0003] Patent Document 1: Japanese Unexamined Patent Application Publication No. 2020-55095

[0004] In the above prior art, as a reward when performing model learning, a good reward is given when the applied load falls within a specified setting range, and a bad reward is given otherwise. However, when conducting trials using an actual robot, due to the influence of sensor noise, the shape of the workpiece, position deviation, etc., the sensor output deviates. Therefore, there is a problem that even if learning can be performed, the possibility of the load exceeding the setting range increases due to the deviation of the sensor output. Therefore, a method that can set parameters to appropriate values even when the sensor output has a deviation is desired. Summary of the Invention

[0005] According to a first aspect of the present disclosure, there is provided a method for setting force control parameters in the operation of a robot. The method includes: step (a) of setting a limit value and an objective function, the limit value defining a limit condition regarding a specific force control characteristic value detected in force control, the objective function regarding a specific evaluation item associated with the operation; step (b) of searching for an optimal value of the force control parameter using the objective function; and step (c) of determining a set value of the force control parameter based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds a tolerance value smaller than the limit value is added to an actual measurement value of the evaluation item.

[0006] According to a second aspect of the present disclosure, a robot system is provided. The robot system includes: a robot; a sensor that detects a specific force control characteristic value during an operation based on force control of the robot; and a parameter setting unit that performs a process of setting force control parameters of the robot. The parameter setting unit performs: process (a) of setting a limit value and an objective function, the limit value defining a limitation condition regarding the force control characteristic value, and the objective function being related to a specific evaluation item associated with the operation; process (b) of searching for an optimal value of the force control parameters using the objective function; and process (c) of determining a set value of the force control parameters based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to an actual measured value of the evaluation item.

[0007] According to a third aspect of the present disclosure, a computer program is provided that causes a processor to perform a process of setting force control parameters in an operation of a robot. The computer program causes the processor to perform: process (a) of setting a limit value and an objective function, the limit value defining a limitation condition regarding a specific force control characteristic value detected in force control, and the objective function being related to a specific evaluation item associated with the operation; process (b) of searching for an optimal value of the force control parameters using the objective function; and process (c) of determining a set value of the force control parameters based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to an actual measured value of the evaluation item. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is an explanatory diagram showing the configuration of the robot system in the embodiment.

[0009] Figure 2 is a functional block diagram of the information processing device in the first embodiment.

[0010] Figure 3 is a flowchart showing the process of setting force control parameters.

[0011] Figure 4 is a flowchart showing the detailed process of the optimal value search process in the first embodiment.

[0012] Figure 5 is a flowchart showing the detailed process of implementing the operation in the first embodiment.

[0013] Figure 6 is a graph showing the shape of the penalty included in the optimized objective function.

[0014] Figure 7It is a functional block diagram of the information processing apparatus in the second embodiment.

[0015] Figure 8 It is an explanatory diagram showing a configuration example of the parameter determination function.

[0016] Figure 9 It is a flowchart showing the detailed process of the optimal value search process in the second embodiment.

[0017] Figure 10 It is a flowchart showing the detailed process of the execution of the operation in the second embodiment.

[0018] Explanation of Reference Numerals

[0019] 100... Robot, 110... Base, 120... Manipulator arm, 122... Arm encoder, 140... Force sensor, 150... End effector, 200... Control device, 300... Information processing apparatus, 310... Processor, 311... Parameter setting unit, 312... Action execution unit, 314... Parameter search unit, 320... Memory, 330... Interface circuit, 340... Input device, 350... Display unit. Detailed Embodiment

[0020] A. First Embodiment:

[0021] Figure 1 It is an explanatory diagram showing an example of a robot control system in the embodiment. The robot control system includes a robot 100, a control device 200 that controls the robot 100, and an information processing apparatus 300. The robot 100 is provided on a gantry. The information processing apparatus 300 is, for example, a personal computer.

[0022] The robot 100 includes a base 110 and a manipulator arm 120. The manipulator arm 120 is sequentially connected by four joints J1 to J4. A force sensor 140 and an end effector 150 are attached to the front end portion of the manipulator arm 120. A TCP (Tool Center Point) that is a control point of the robot 100 is set near the front end of the manipulator arm 120. In the present embodiment, a 4-axis robot having four joints J1 to J4 is illustrated, but a robot having an arbitrary arm mechanism including a plurality of joints can be used. In addition, the robot 100 of the present embodiment is a horizontal multi-joint robot, but a vertical multi-joint robot can also be used.

[0023] The force sensor 140 is a sensor that detects the force applied to the end effector 150. As the force sensor 140, a force sensor capable of detecting the force in a single-axis direction, a force sensor capable of detecting the force components in multiple axis directions, or a torque sensor can be used. In the present embodiment, as the force sensor 140, a 6-axis force sensor is used. The 6-axis force sensor detects the magnitude of the force parallel to three detection axes orthogonal to each other and the magnitude of the torque around the three detection axes in the inherent sensor coordinate system. It should be noted that the force sensor 140 can also be provided at a position other than the position of the end effector 150. For example, it can also be provided at one or more of the joints J1 to J4.

[0024] In the present embodiment, the robot 100 performs an operation of assembling two workpieces W1 and W2 by fitting the first workpiece W1 into the hole HL of the second workpiece W2. In this operation, force control using the force detected by the force sensor 140 is performed.

[0025] Figure 2 It is a block diagram showing the functions of the information processing device 300 in the first embodiment. The information processing device 300 can be implemented as an information processing device such as a personal computer. The information processing device 300 includes a processor 310, a memory 320, an interface circuit 330, an input device 340, and a display unit 350 connected to the interface circuit 330. A control device 200 is also connected to the interface circuit 330. The measured values of the arm encoder 122 of the robot 100 and the force sensor 140 are input to the interface circuit 330 via the control device 200. The arm encoder 122 is a position sensor that measures the position or displacement on multiple joints of the robotic arm 120.

[0026] The processor 310 has a function as a parameter setting unit 311 for setting the force control parameters of the robot 100. The parameter setting unit 311 includes the functions of an action execution unit 312 and a parameter search unit 314. The action execution unit 312 causes the robot 100 to perform an operation according to the robot control program. The parameter search unit 314 performs a process of searching for the force control parameters of the robot 100. The functions of the parameter setting unit 311 are implemented by the processor 310 executing the computer program stored in the memory 320. However, a part or all of the functions of the parameter setting unit 311 can also be implemented by a hardware circuit.

[0027] The memory 320 stores a parameter initial value IP, a search condition SC, and a robot control program RP. The parameter initial value IP is the initial value of the force control parameter. The search condition SC is the condition for performing the search process of the force control parameter. The robot control program RP is composed of a plurality of commands for causing the robot 100 to operate.

[0028] Figure 3This is a flowchart showing the process of setting force control parameters. In step S110, the operator determines the operation for which the force control parameters are to be set. In the present embodiment, as this operation, the fitting operation of the two workpieces W1 and W2 shown in Figure 1 is selected. The robot control program RP is a program that describes the actions of this operation. In the case where a plurality of robot control programs describing various operations have been created in advance, in step S110, the operator selects one of them. In the case where the robot control program has not been created, in step S110, the robot control program is created.

[0029] In step S120, the operator sets the constraints when searching for the optimal value of the force control parameter, and in step S130, sets the search range when searching for the optimal value of the force control parameter. As the constraints, for example, various force control characteristic value-related conditions such as the resultant force and resultant moment calculated from the measured values of the force sensor 140, the maximum value of the load of the arm motor, and the maximum value of vibration can be used. "Force control characteristic value" means the value detected in force control. In the case where the maximum value of vibration is set as the constraint, a vibration sensor for detecting vibration as the force control characteristic value is provided at the fingertip or the mount of the robot 100. In the present embodiment, the resultant force is used as the force control characteristic value for specifying the constraint. This will be further described later.

[0030] As the force control parameters to be optimized, for example, one or more of the virtual mass coefficient, virtual viscosity coefficient, and virtual elasticity coefficient in impedance control can be selected. Alternatively, they can also be expressed in the form of a function of the parameter α, that is, virtual mass coefficient = M(α), virtual viscosity coefficient = B(α), and virtual elasticity coefficient = K(α), and the parameter α is set as the object to be optimized. In addition, other items other than these can also be set as the object to be optimized. As the method of force control, it is not limited to impedance control, and force control based on position control such as rigid control and damping control, or force control based on torque control such as Active Stiffness Control can be adopted. These controls use different force control parameters.

[0031] In step S140, the parameter search unit 314 searches for the optimal value of the force control parameter according to the set constraints and search range.

[0032] Figure 4is a flowchart showing the detailed process of the optimal value search process in step S140. When the search starts in step S210, in step S220, the parameter search unit 314 searches for candidate values of the force control parameter. This search is performed using an optimization algorithm such as CMA-ES (Covariance Matrix Adaptation Evolution Strategy), for example. The objective function used in the optimization algorithm will be described later. It should be noted that when step S220 is first executed, a preset initial value is used as the candidate value of the force control parameter.

[0033] In step S230, the parameter search unit 314 sets the candidate value of the force control parameter selected in step S220 in the robot control program RP. In step S240, the action execution unit 312 performs the operation of the robot 100 according to the robot control program RP. During this operation, measurement values of sensors including the force sensor 140 are acquired.

[0034] Figure 5 is a flowchart showing the detailed process of the execution of the operation in step S240. In step S241, the action execution unit 312 performs an operation including a force control action. This operation continues until it is determined in step S242 that the operation has ended. When the operation ends, in step S243, the measurement results including the measurement values of the sensors and the cycle time of the operation are recorded.

[0035] In Figure 4 step S250, the parameter search unit 314 calculates the value of the objective function in the optimization process based on the measurement results obtained during the operation. In the present embodiment, as the objective function, for example, the following formula is used.

[0036] [Math.1]

[0037] y = t + G(f peak ) (1)

[0038] G(f peak ) = 0 (f peak ≤F1) (2a)

[0039] G(f peak ) = θ1(f peak -F1) (F1 < f peak ≤F2) (2b)

[0040] G(f peak ) = θ1(f peak -F1) + θ2(f peak -F2)(f peak >F2) (2c)

[0041] Here, y is the objective function, t is the cycle time as an evaluation item of the operation, and f peak is the maximum value of the resultant force of the forces measured during the operation, and G(f peak ) is the penalty based on the maximum resultant force f peak , F1 is the allowable value as the first threshold of the maximum resultant force f peak , F2 is the limit value as the second threshold of the maximum resultant force f peak , and θ1, θ2 are coefficients showing the increase rate of the penalty. The cycle time t is measured by the timer of the robot system.

[0042] In this objective function y, the cycle time t is used as an evaluation item of the operation, but evaluation items other than the cycle time t can also be used. For example, the number of operations per unit time, the error between the actual position / pose in the terminal and the target position / pose, etc. can also be used as evaluation items.

[0043] The maximum resultant force f peak is the force evaluation value that is the object of setting the limit condition in step S120, and is the maximum value measured during the operation among the resultant forces of the forces in the three-axis directions measured by the force sensor 140. As the force evaluation value that is the object of setting the limit condition, other force evaluation values such as the maximum value of the resultant torque can also be used.

[0044] The allowable value F1 and the limit value F2 are set to the following values respectively.

[0045] <Limit value F2>

[0046] The limit value F2 is the value that stipulates the upper limit of the maximum resultant force f peak as the limit condition. The limit value F2 is, for example, the threshold value that may damage the workpiece W1. The limit value F2 is specified by the user according to conditions such as the raw material of the workpiece W1.

[0047] <Allowable value F1>

[0048] The allowable value F1 is a value preset to be smaller than the limit value F2. Specifically, it is preferable to consider the measurement deviation of the maximum resultant force f peak as the force evaluation value and set the allowable value F1 to the value obtained by subtracting the quantified measurement deviation from the limit value F2. The difference between the two (F2 - F1) can be set to, for example, the value obtained by multiplying the standard deviation of the maximum resultant force f peak by a coefficient of about 1 to 3. Generally, it is preferable to set the difference between the allowable value F1 and the limit value F2 of the force evaluation value to a value that has a positive correlation with the standard deviation of the force evaluation value.

[0049] Figure 6is a graph showing the shape of the penalty G(f peak ) for the objective function y. This penalty G is zero when the maximum resultant force f peak is below the allowable value F1, and increases at a first increase rate θ1 according to the excess amount (f peak - F1) when it exceeds the allowable value F1. In addition, when the maximum resultant force f peak exceeds the limit value F2, a component that increases at a second increase rate θ2 corresponding to this excess amount (f peak - F2) is also added to the penalty G. That is, the penalty G when the maximum resultant force f peak exceeds the limit value F2 is the sum of the first term and the second term on the right side of the above (2c). The first term θ1(f peak - F1) on the right side of the above (2c) is called the "first penalty component", and the second term θ2(f peak - F2) is called the "second penalty component".

[0050] Thus, the penalty G for the objective function y when the maximum resultant force f peak exceeds the limit value F2 includes the first penalty component θ1(f peak - F1) and the second penalty component θ2(f peak - F2). Accordingly, when the maximum resultant force f peak exceeds the limit value F2, a greater penalty can be given, so the possibility that the maximum resultant force f peak exceeds the limit value F2 can be reduced.

[0051] It should be noted that as the objective function y, a function expressed in a form other than the above formula can also be used. For example, the penalty G for the objective function y when the maximum resultant force f peak exceeds the limit value F2 can be set to only the first penalty component θ1(f peak - F1). In this case, the objective function y also has a form in which the penalty G that increases according to the excess amount by which the maximum resultant force f peak exceeds the allowable value F1 smaller than the limit value F2 specified by the constraint condition is added to the measured value of the cycle time t as an evaluation item, which is common to the above objective function. If such an objective function is used, it has the advantage of being able to optimize the force control parameters while reducing the possibility of exceeding the limit value F2 as a constraint condition. In addition, even if there is an output deviation in the force sensor 140, the possibility that the maximum resultant force f peak exceeds the limit value F2 as a force control characteristic value can be reduced. In other examples, the first penalty component and the second penalty component can also be set to functions that increase in a curve form.

[0052] In Figure 4In step S260, the parameter search unit 314 confirms the optimal value of the force control parameter based on the value of the objective function y. That is, when the value of the objective function y is the minimum value so far, the current candidate value is updated to the new optimal value. On the other hand, when the value of the objective function y is not the minimum value so far, the previous optimal value is still maintained.

[0053] In step S270, the parameter search unit 314 determines whether the search end condition is satisfied. As the search end condition, for example, a condition that the cycle time t of the operation reaches a preset target value can be used. Or, when a better solution is not found in the most recent preset number of searches, for example, the most recent 20 searches, it can also be determined that the search end condition is satisfied.

[0054] In step S280, the parameter search unit 314 determines whether the optimal value of the force control parameter has been updated. When the optimal value has not been updated, it returns to step S220. On the other hand, when the optimal value of the force control parameter has been updated, it proceeds to step S290.

[0055] In step S290, the motion execution unit 312 tries the operation performed by the robot 100 a plurality of times using the updated optimal value of the force control parameter, and collects the force control characteristic values associated with the objective function y. In the present embodiment, as the force control characteristic value, the measured value of the maximum resultant force f peak is collected.

[0056] In step S300, the parameter search unit 314 updates the penalty setting in the objective function y. For example, the update can be performed as follows. In the following description, N is the number of trials of the operation in step S290 and is an integer of 2 or more.

[0057] <Update of the allowable value F1>

[0058] Preferably, the allowable value F1 is updated according to the following formula.

[0059] F1 = F2 - k1 × σ (3)

[0060] Here, σ is the standard deviation of the maximum resultant force f peak in N trials, and k1 is a coefficient. The coefficient k1 is a positive real number, and is set to a value in the range of 1 to 3, for example.

[0061] For example, when k1 = 3, the allowable value F1 is a value that can suppress the probability of exceeding the limit value F2 to 0.15% or less even considering the deviation when the average value of the measured maximum resultant force f peak is the allowable value. When the maximum resultant force f peakWhen the probability of exceeding the limit value F2 can be allowed even if it is slightly high, the coefficient k1 can be set to a value in the range of 1 < k1 ≤ 3. It should be noted that since the limit value F2 is provided as a limiting condition, it is preferably not updated.

[0062] The second term k1×σ on the right side of the above formula (3) is the difference between the allowable value F1 and the limit value F2 for the maximum resultant force f peak and is equal to the value obtained by multiplying the standard deviation σ of the maximum resultant force f peak by the coefficient k1. Generally, the difference between the allowable value F1 and the limit value F2 for the force control characteristic value is preferably set to be equal to the difference value that has a positive correlation with the standard deviation σ of the force control characteristic value. In this way, an appropriate allowable value F1 can be set according to the deviation of the force control characteristic value.

[0063] It should be noted that when the one-time operation in step S290 consists of multiple operation parts, the maximum resultant force f can also be recorded in advance peak for which operation part is likely to be the maximum, and only repeatedly perform a part of the operation including the operation part for which the maximum resultant force f peak is likely to be the maximum, and calculate the deviation of the maximum resultant force f peak For example, the operation of fitting the first workpiece W1 and the second workpiece W2 Figure 1 can be divided into the following operation parts.

[0064] (1) After moving the end effector 150 above the first workpiece W1, lower it to the position of the first workpiece W1.

[0065] (2) Grasp the first workpiece W1 with the end effector 150.

[0066] (3) After moving the grasped first workpiece W1 above the second workpiece W2, lower it directly above the second workpiece W2.

[0067] (4) Fit the first workpiece W1 and the second workpiece W2.

[0068] (5) Release the first workpiece W1 and move the end effector 150 upward to avoid.

[0069] In this case, a part of the operation including the operation part (4) of fitting the first workpiece W1 and the second workpiece W2 can also be repeatedly performed to calculate the deviation of the maximum resultant force f peak In this way, the time required for repeating the trial in step S290 can be reduced.

[0070] <Update of the first increase rate θ1>

[0071] The first increase rate θ1 preferably uses the maximum resultant force f in N trialspeak The average value f N is compared with the allowable value F1 and updated as described below.

[0072] (a) If f N ≥ F1 + m1, then set θ1 = θ1 + a to increase θ1.

[0073] (b) If f N ≤ F1 - m2, then set θ1 = θ1 - a to decrease θ1.

[0074] Here, m1 and m2 are the margins for whether the maximum resultant force f peak The average value f N is within a range approximately the same as the allowable value F1, and a is a constant for increasing or decreasing the increase rate θ1. m1, m2, and a are all positive real numbers. The margin m1 is called the "first determination value", and the margin m2 is also called the "second determination value". However, it is also possible that m1 = m2.

[0075] When it is assumed that the maximum resultant force f peak in N trials follows a normal distribution, the values corresponding to its deviation and standard deviation can be set as the margins m1 and m2. The average value f peak of the maximum resultant force f N is ideal near the allowable value F1 of the maximum resultant force f peak , but when the two are too far apart, by adjusting the increase rate θ1 of the penalty in this way, the average value f peak of the maximum resultant force f N can be made to fall within a range close to the allowable value F1.

[0076] In this way, it is preferable that when the average value f N of the force control characteristic value is greater than the allowable value F1 by more than the first determination value m1, the first increase rate θ1 is increased, and when the average value f N is less than the allowable value F1 by more than the second determination value m2, the first increase rate θ1 is decreased. In this case, there is an advantage that the first increase rate θ1 can be adjusted so that the force control characteristic value measured during the operation falls within a range close to the allowable value F1.

[0077] <Update of the second increase rate θ2>

[0078] The second increase rate θ2 is preferably updated as described below according to the ratio P = n / N of the number n of trials in which the maximum resultant force f peak exceeds the limit value F2 in N trials, that is, the exceeding ratio P.

[0079] (a) When P1 ≤ P, set θ2 = θ2 + b to increase θ2.

[0080] (b) When P ≤ P2, set θ2 = θ2 - b to decrease θ2.

[0081] Here, P1 and P2 are determination values, P2 < P1, and b is a constant used to increase or decrease the increase rate θ2. P1, P2, and b are all positive real numbers.

[0082] In this way, it is preferable to increase the second increase rate θ2 when the excess ratio P is above the first determination ratio P1, and to decrease the second increase rate θ2 when the excess ratio P is below the second determination ratio P2 which is smaller than the first determination ratio P1. If the second increase rate θ2 is updated in this way, the maximum resultant force f can be further reduced during the optimization process. peak The possibility of exceeding the limit value F2.

[0083] It is also possible not to perform part of the above-mentioned update of the allowable value F1, the update of the first increase rate θ1, and the update of the second increase rate θ2. In addition, steps S290 and S300 can be omitted without updating the penalty setting.

[0084] In addition, it is also possible not to perform the update of the penalty setting in step S290 before a preset number of searches are performed, but only to perform it when the optimal value of the force control parameter is updated later. The reason is that in the initial several searches, the possibility that the optimal value of the force control parameter at that time becomes the final optimal value is low. Therefore, when limiting the update of the penalty setting after a certain number of searches, the number of trials of the operation in step S290 can be reduced within a range with little influence on the final optimal value.

[0085] Moreover, in the case of using an algorithm such as CMA-ES that divides generations, it is also possible to confirm whether the optimal value of the force control parameter has been updated for each generation and update the penalty setting. In this way, when the optimal value is continuously updated within the same generation, the repeated trials of the operation can be prevented.

[0086] When the update of the penalty setting in step S300 ends, return to step S220 and repeat the above process again. It should be noted that the selection of the new candidate value in step S220 is performed according to the optimization algorithm using the objective function y and the history of the candidate values.

[0087] In this way, when the search for the optimal value of the force control parameter ends, proceed to Figure 3In step S150, the search results are displayed on the display unit 350. In step S160, the operator determines the force control parameter with reference to the search results. Specifically, for example, the operator may directly adopt the optimal value of the searched force control parameter as the final force control parameter. Alternatively, in the case where there is a candidate value that is not the optimal value but slightly deviates from the limit condition and has a tact time, which is a main evaluation item, superior to the optimal value, the operator may adopt this candidate value. In order to enable such a selection, as the search results, for a plurality of candidate values including the optimal value of the force control parameter, it is preferable to display the plurality of candidate values in a list form respectively including values of evaluation items such as tact time and the maximum resultant force f peak and other force control characteristic values.

[0088] In summary, in the first embodiment, as the objective function for optimizing the force control parameter, an objective function in a form where a penalty that increases according to the excess amount by which a specific force control characteristic value exceeds the allowable value is added to the measured value of the evaluation item is used. Therefore, it is possible to optimize the force control parameter while reducing the possibility of exceeding the limit value as a limit condition. In addition, even if there is an output deviation in the force sensor, it is possible to reduce the possibility of the force control characteristic value exceeding the limit value as a limit condition.

[0089] B. Second Embodiment:

[0090] Figure 7 is a block diagram showing the functions of the information processing apparatus 300 in the second embodiment. The difference from Figure 2 the first embodiment shown is only that the parameter determination function PF is stored in the memory 320 of the information processing apparatus 300 in the second embodiment, and other configurations are substantially the same as those in the first embodiment. In addition, the overall configuration of the robot system is also the same as Figure 1 the configuration shown. The parameter determination function PF is a function that determines the force control parameter according to the state of the robot 100 such as the fingertip position and speed of the robot 100.

[0091] Figure 8 is an explanatory diagram showing a configuration example of the parameter determination function PF. This parameter determination function PF is configured as a three-layer neural network. The input of the neural network is a value showing the state of the robot 100. In Figure 8 the example of, the fingertip position, the fingertip speed, the force, and the torque are used as the input. "Fingertip position" means Figure 1Regarding the position and orientation of the control point TCP shown, "tip speed" means the speed of the control point TCP. "Force" and "torque" are values measured by the force sensor 140. The input to the parameter determination function PF is also referred to as the "state observation value". The outputs of the parameter determination function PF are the virtual mass coefficient M, the virtual viscosity coefficient B, and the virtual elasticity coefficient K. It should be noted that, as inputs and outputs, items other than these can also be used. In addition, the number of neuron layers is not limited to 3 layers and can also be set to 4 layers or more. In addition, the parameter determination function PF can also be implemented by a configuration other than a neural network.

[0092] As is well known, the operation at each node of a neural network uses weights w L ij for weighted addition. Here, L is an ordinal number indicating the order of the layer, i is an ordinal number indicating the order of the node in the lower layer, and j is an ordinal number indicating the order of the node in the upper layer. Here, all the weights w L ij in the parameter determination function PF are arranged in sequence, and the resulting vector is defined as the weight vector W = (w 1 11 , w 1 12 , …, w 2 33 ). In the second embodiment, the parameter search unit 314 uses an optimization algorithm to perform optimization with this weight vector W as the search target. In the following description, the weight vector W is referred to as the "internal parameter W".

[0093] Figure 9 is a flowchart showing the detailed process of the optimal value search process in the second embodiment. The difference from the Figure 4 process of the first embodiment shown is only that Figure 4 the steps S220, S230, S240, S260 are replaced with steps S225, S235, S245, S265, and the other steps are substantially the same as those of the first embodiment. In addition, Figure 3 the overall process of the setting process of the force control parameters shown is the same as that of the first embodiment.

[0094] When the search starts in step S210, in step S225, the parameter search unit 314 searches for candidate values of the internal parameter W of the parameter determination function PF. This search is also performed using an optimization algorithm such as CMA-ES. It should be noted that when step S225 is first executed, a preset initial value is used as the candidate value of the internal parameter W.

[0095] In step S235, the parameter search unit 314 sets the candidate value of the internal parameter W selected in step S225 to the parameter determination function PF. In step S245, the action execution unit 312 implements the operation of the robot 100 according to the robot control program RP.

[0096] Figure 10 is a flowchart showing the detailed process of implementing the operation in step S245. In step S410, the parameter search unit 314 acquires the state observation value of the robot 100. The state observation value of the robot 100 is a value that becomes Figure 8 the input value of the parameter determination function PF shown. The state observation value is obtained by observing the state of the robot 100 using sensors such as the force sensor 140 and the arm encoder 122. In step S420, the parameter search unit 314 uses the parameter determination function PF to determine the candidate value of the force control parameter. At this time, the state observation value acquired in step S410 is used as the input of the parameter determination function PF. The internal parameter W of the parameter determination function PF is set in Figure 9 step S235.

[0097] In step S430, the action execution unit 312 executes an operation including a force control action. This operation continues until it is determined in step S440 that the operation is completed. If the operation is not completed, it proceeds to step S450 described later. On the other hand, if the operation is completed, it proceeds to step S460, and records the measurement results including the measured values of the sensors and the cycle time of the operation.

[0098] In step S450, the parameter search unit 314 determines whether a preset update time has elapsed. The "update time" is a time suitable for updating the candidate value of the force control parameter. If the update time has not elapsed, the operation in step S440 continues. On the other hand, if the update time has elapsed, it returns to step S410, and the above-described steps S410 and subsequent processes are executed again.

[0099] As the judgment criterion in step S450, generally, the judgment criterion of whether a preset update condition is satisfied can be used. At this time, in the middle of one operation, steps S410 to S430 are repeated every time the update condition is satisfied. As the update condition in step S450, as Figure 10 shown, a certain update time can be used, or instead, the condition of whether one operation part as a part of the operation is completed can be used. For example, when the entire operation includes an operation part using force control and an operation part not using force control, the completion of each of these operation parts can also be set as the judgment criterion in step S450. For example, regarding Figure 1For the operation of fitting the first workpiece W1 and the second workpiece W2, it is also possible to set whether each of the following five operation parts is completed as the judgment criterion.

[0100] (1) After moving the end effector 150 above the first workpiece W1, lower it to the position of the first workpiece W1.

[0101] (2) Grasp the first workpiece W1 with the end effector 150.

[0102] (3) After moving the grasped first workpiece W1 above the second workpiece W2, lower it to directly above the second workpiece W2.

[0103] (4) Fit the first workpiece W1 and the second workpiece W2 together.

[0104] (5) Release the first workpiece W1 and move the end effector 150 upward to avoid it.

[0105] When updating the candidate values of the force control parameters for each different operation part as described above, as the state observation value obtained in step S410, it is also possible to use the state observation value at a specific timing in the history of the state observation values during the trial of the past operation parts. For example, regarding the operation part to be executed next in step S430, it is also possible to obtain the maximum resultant force f in the past trial of obtaining the optimal value of the force control parameter from the memory 320 in step S410 peak at the timing of. In this way, it is possible to use the state observation value suitable for this operation part to determine the candidate value of the force control parameter through the parameter determination function PF. It should be noted that when step S410 is first executed, a prescribed initial value is used as the state observation value.

[0106] It should be noted that step S450 can also be omitted, and when it is determined in step S440 that the operation is not completed, return to step S430. In this case, the candidate value of the force control parameter is maintained as the same value throughout the operation. In addition, in this case, as the state observation value in step S410, it is also possible to use the state observation value at a specific timing in the history of the state observation values during the trial of the past operation. For example, it is also possible to obtain the maximum resultant force f in the past trial of obtaining the optimal value of the force control parameter from the memory 320 in step S410 peak at the timing of. However, when step S410 is first executed, a prescribed initial value is used as the state observation value.

[0107] Thus, when the operation ends, in step S460, the measurement results including the measured values of the sensor type and the cycle time of the operation are recorded. It should be noted that when the candidate value of the force control parameter is updated during the operation, the history of the candidate value of the force control parameter is also recorded. At this time, the history of the state observation value can also be recorded.

[0108] Figure 9 The processing of step S250 of is the same as that of the first embodiment, and the parameter search unit 314 calculates the value of the objective function in the optimization process based on the measurement results obtained during the operation. In step S265, the parameter search unit 314 confirms the optimal values of the force control parameter and the internal parameter according to the value of the objective function y. That is, when the value of the objective function y is the minimum value so far, the candidate value this time is updated to the new optimal value. On the other hand, when the value of the objective function y is not the minimum value so far, the previous optimal value is still maintained. The processing after step S270 is the same as that of the first embodiment, so the description is omitted.

[0109] Thus, in the second embodiment, the candidate values of the internal parameter W of the parameter determination function PF are searched according to the optimization algorithm using the objective function y. In addition, during the execution of the operation, the state of the robot 100 is observed, and the state observation value is input to the parameter determination function PF to obtain the candidate value of the force control parameter, and the force control action is executed using the force control parameter. As a result, the force control parameter can be adaptively determined according to the state of the robot 100, so that appropriate force control parameters can be used at each stage of the operation.

[0110] In addition, in the second embodiment, the objective function in the form of adding a penalty increased according to the excess amount of a specific force control characteristic value exceeding the allowable value to the measured value of the evaluation item is also used in the same way as the first embodiment. Therefore, it is possible to optimize the force control parameter while reducing the possibility of exceeding the limit value as a limiting condition. In addition, even if there is an output deviation in the force sensor, it is possible to reduce the possibility of the force control characteristic value exceeding the limit value as a limiting condition.

[0111] Other embodiments:

[0112] The present disclosure is not limited to the above embodiments and can be implemented in various ways without departing from its gist. For example, the present disclosure can also be implemented through the following aspects. To solve part or all of the technical problems of the present disclosure or to achieve part or all of the effects of the present disclosure, the technical features in the above embodiments corresponding to the technical features in each of the following aspects can be appropriately replaced and combined. In addition, if the technical feature is not described as an essential technical feature in this specification, it can be appropriately deleted.

[0113] (1) According to a first aspect of the present disclosure, a method for setting force control parameters in an operation of a robot is provided. The method includes: step (a) of setting a limit value and an objective function, the limit value defining a limit condition regarding a specific force control characteristic value detected in force control, the objective function regarding a specific evaluation item associated with the operation; step (b) of searching for an optimal value of the force control parameter using the objective function; and step (c) of determining a set value of the force control parameter based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to an actual measured value of the evaluation item.

[0114] According to this method, as an objective function for optimizing the force control parameter, an objective function having a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds the allowable value is added to the actual measured value of the evaluation item is used. Therefore, it is possible to optimize the force control parameter while reducing the possibility of exceeding the limit value that is a limit condition. In addition, even if there is an output deviation in the force sensor, it is possible to reduce the possibility of the force control characteristic value exceeding the limit value that is a limit condition.

[0115] (2) In the above method, it may also be that, when the force control characteristic value exceeds the limit value, the penalty of the objective function includes a first penalty component and a second penalty component. The first penalty component is obtained by multiplying the excess amount by which the force control characteristic value exceeds the allowable value by a first increase rate, and the second penalty component is obtained by multiplying the excess amount by which the force control characteristic value exceeds the limit value by a second increase rate.

[0116] According to this method, by giving a greater penalty to the excess amount exceeding the limit value, it is possible to reduce the possibility of exceeding the limit value during optimization.

[0117] (3) In the above method, it may also be that step (b) includes: step (b1) of maintaining the value of the force control parameter, trying the operation performed by the robot multiple times, and measuring the force control characteristic value at each try; step (b2) of updating the objective function based on the measured values of the force control characteristic value in the multiple tries; and step (b3) of performing the search for the optimal value using the updated objective function. Step (b2) includes the following steps: obtaining a standard deviation of the force control characteristic value in the multiple tries; and updating the allowable value such that a difference between the allowable value and the limit value is equal to a difference value that has a positive correlation with the standard deviation.

[0118] According to this method, it is possible to set an appropriate allowable value based on the deviation of the force control characteristic value.

[0119] (4) In the above method, it may also be that the process (b2) further includes the following processes: obtaining the average value of the force control characteristic values in the multiple trials; and increasing the first increase rate when the average value is greater than the allowable value by more than the first determination value, and decreasing the first increase rate when the average value is less than the allowable value by more than the second determination value.

[0120] According to this method, it is possible to adjust the first increase rate so that the force control characteristic values measured during the operation fall within a range close to the allowable value.

[0121] (5) In the above method, it may also be that the process (b2) further includes the following processes: obtaining the excess ratio, which is the ratio of the trials in which the force control characteristic values exceed the limit value in the multiple trials; and increasing the second increase rate when the excess ratio is equal to or greater than the first determination ratio, and decreasing the second increase rate when the excess ratio is less than or equal to the second determination ratio, which is smaller than the first determination ratio.

[0122] According to this method, it is possible to reduce the possibility that the force control characteristic values measured during the operation exceed the limit value.

[0123] (6) In the above method, it may also be that the process (b) includes: process (i), searching according to an optimization algorithm using the objective function to obtain candidate values of the internal parameters of a parameter determination function that takes a state observation value showing the state of the robot as input and the force control parameters as output; process (ii), setting the candidate values of the internal parameters obtained in process (i) in the parameter determination function; process (iii), obtaining the state observation value; process (iv), obtaining candidate values of the force control parameters by inputting the state observation value into the parameter determination function; process (v), performing the operation of the robot using the candidate values of the force control parameters to obtain measurement results including the force control characteristic values and the evaluation items; process (vi), calculating the value of the objective function according to the measurement results; and process (vii), repeating processes (i) to (vi) until the search end condition is satisfied.

[0124] According to this method, it is possible to adaptively determine the force control parameters according to the state of the robot.

[0125] (7) In the above method, it may also be that processes (iii) to (v) are repeatedly performed whenever a preset update condition is satisfied during one operation.

[0126] According to this method, different force control parameters can be used whenever the update condition is satisfied, so that appropriate force control parameters can be adopted in each stage of the operation.

[0127] (8) In the above method, it may also be that the evaluation item is the tact time of the operation, the force control characteristic value includes at least one of the maximum value of the resultant force and the maximum value of the resultant torque given by the robot to the workpiece as the object of the operation, and the force control parameter includes at least one of a virtual mass coefficient, a virtual viscosity coefficient, and a virtual elasticity coefficient.

[0128] According to this method, optimization regarding at least one of the virtual mass coefficient, the virtual viscosity coefficient, and the virtual elasticity coefficient can be performed using an objective function in a form in which a penalty corresponding to the maximum value of the resultant force and the resultant torque is added to the measured value of the tact time.

[0129] (9) According to the second aspect of the present disclosure, a robot system is provided. The robot system includes: a robot; a sensor that detects a specific force control characteristic value in an operation based on force control of the robot; and a parameter setting unit that performs a process of setting the force control parameter of the robot. The parameter setting unit performs: process (a) of setting a limit value and an objective function, the limit value defining a limitation condition regarding the force control characteristic value, and the objective function being related to a specific evaluation item associated with the operation; process (b) of searching for an optimal value of the force control parameter using the objective function; and process (c) of determining a set value of the force control parameter based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to the measured value of the evaluation item.

[0130] (10) According to the third aspect of the present disclosure, a computer program is provided that causes a processor to perform a process of setting a force control parameter in an operation of a robot. The computer program causes the processor to perform: process (a) of setting a limit value and an objective function, the limit value defining a limitation condition regarding a specific force control characteristic value detected in force control, and the objective function being related to a specific evaluation item associated with the operation; process (b) of searching for an optimal value of the force control parameter using the objective function; and process (c) of determining a set value of the force control parameter based on the result of the search. The objective function has a form in which a penalty that increases according to an excess amount by which the force control characteristic value exceeds an allowable value smaller than the limit value is added to the measured value of the evaluation item.

[0131] The present disclosure can also be implemented in various aspects other than those described above. For example, it can be implemented in aspects such as a robot system including a robot and a robot control device, a computer program for implementing the functions of the robot control device, and a non-transitory storage medium storing the computer program.

Claims

1. A method for setting force control parameters in the operation of a robot, characterized in that, include: Step (a), setting a limit value and an objective function, wherein the limit value specifies a restriction condition on a specific force control characteristic value detected in force control, and the objective function is related to a specific evaluation item associated with the operation; Step (b), searching for an optimal value of the force control parameter using the objective function; and Step (c), determining a setting value of the force control parameter according to the search result, The objective function has a form in which a penalty which increases according to the amount by which the force control characteristic value exceeds a permissible value smaller than the limit value is added to the actual measured value of the evaluation item, When the force control characteristic value exceeds the limit value, the penalty of the objective function includes a first penalty component and a second penalty component, wherein the first penalty component is obtained by multiplying the excess of the force control characteristic value over the allowable value by a first increase rate, and the second penalty component is obtained by multiplying the excess of the force control characteristic value over the limit value by a second increase rate.

2. The method according to claim 1, characterized in that, The step (b) comprises: Step (b1), maintaining the value of the force control parameter, trial-running the operation performed by the robot a plurality of times, and measuring the force control characteristic value during each trial; Step (b2), updating the objective function according to the measured values ​​of the force control characteristic value in a plurality of trials; and Step (b3), performing the search for the optimal value using the updated objective function, The step (b2) comprises the following steps: Determining a standard deviation of the force control characteristic values ​​in the plurality of trials; and The allowable value is updated so that the difference between the allowable value and the limit value is equal to a difference value that has a positive correlation with the standard deviation.

3. The method according to claim 2, characterized in that, The step (b2) further includes the following steps: Calculating an average value of the force control characteristic value in the plurality of trials; and When the average value is larger than the allowable value by a first determination value or more, the first increase rate is increased, and when the average value is smaller than the allowable value by a second determination value or more, the first increase rate is decreased.

4. The method according to claim 2, characterized in that, The step (b2) further includes the following steps: Calculating an excess ratio, wherein the excess ratio is a ratio of trials in which the force control characteristic value exceeds the limit value among the plurality of trials; as well as When the excess ratio is equal to or greater than a first determination ratio, the second increase rate is increased, and when the excess ratio is equal to or less than a second determination ratio which is smaller than the first determination ratio, the second increase rate is decreased.

5. The method according to claim 1, characterized in that, The step (b) comprises: Step (i), searching for candidate values ​​of internal parameters of a parameter determination function having a state observation value indicating a state of the robot as an input and having the force control parameter as an output according to an optimization algorithm using the objective function; step (ii), setting the candidate value of the internal parameter obtained in the step (i) in the parameter determination function; Step (iii), obtaining the state observation value; Step (iv) of obtaining a candidate value of the force control parameter by inputting the state observation value into the parameter determination function; Step (v), performing the operation of the robot by using a candidate value of the force control parameter, thereby obtaining a measurement result including the force control characteristic value and the evaluation item; Step (vi), calculating the value of the objective function based on the measurement result; and Step (vii), repeatedly performing Step (i) to Step (vi) until the search end condition is satisfied.

6. The method according to claim 5, characterized in that, Step (iii) to Step (v) are repeatedly performed whenever a preset update condition is satisfied during one operation.

7. The method according to claim 1, characterized in that, The evaluation item is the cycle time of the operation, The force control characteristic value includes at least one of the maximum value of the resultant force and the maximum value of the resultant torque given from the robot to a workpiece as an object of the operation, The force control parameter includes at least one of a virtual mass coefficient, a virtual viscous coefficient, and a virtual elastic coefficient.

8. A robot system, characterized in that, Comprising: A robot; A sensor that detects a specific force control characteristic value during an operation based on the force control of the robot; and A parameter setting unit that performs a process of setting the force control parameter of the robot, The parameter setting unit performs: Process (a), setting a limit value and an objective function, the limit value defining a limit condition regarding the force control characteristic value, and the objective function regarding a specific evaluation item associated with the operation; Process (b), searching for an optimal value of the force control parameter by using the objective function; And Process (c), determining a set value of the force control parameter based on the result of the search, The objective function has a form in which a penalty that increases according to an excess amount of the force control characteristic value exceeding an allowable value smaller than the limit value is added to the measured value of the evaluation item, When the force control characteristic value exceeds the limit value, the penalty of the objective function includes a first penalty component and a second penalty component. The first penalty component is obtained by multiplying the excess amount of the force control characteristic value exceeding the allowable value by a first increase rate, and the second penalty component is obtained by multiplying the excess amount of the force control characteristic value exceeding the limit value by a second increase rate.

9. A computer program, characterized in that, Causing a processor to perform a process of setting a force control parameter in an operation of a robot, The computer program causes the processor to perform: Process (a), setting a limit value and an objective function, the limit value defining a limit condition regarding a specific force control characteristic value detected in force control, and the objective function regarding a specific evaluation item associated with the operation; Process (b), searching for an optimal value of the force control parameter by using the objective function; And Process (c), determining a set value of the force control parameter based on the result of the search, The objective function has a form in which a penalty that increases according to an excess amount of the force control characteristic value exceeding an allowable value smaller than the limit value is added to the measured value of the evaluation item, When the force control characteristic value exceeds the limit value, the penalty of the objective function includes a first penalty component and a second penalty component. The first penalty component is obtained by multiplying the excess amount by which the force control characteristic value exceeds the allowable value by a first increase rate, and the second penalty component is obtained by multiplying the excess amount by which the force control characteristic value exceeds the limit value by a second increase rate.

Citation Information

Patent Citations

  • Control device and control system

    JP2020055095A

  • Method of adjusting force control parameters for force control motion of robot, robot system, and computer program

    JP2023042678A

  • Method Of Setting Force Control Parameter In Work Of Robot, Robot System, And Computer Program

    US20220410386A1