Machine learning device, servo control device, servo control system, and machine learning method

Through machine learning devices, the servo control unit is learned, and the position deviation and instructions are corrected using functions to solve the problem of interaxial interference and improve the follow-up ability of the command.

CN112445181BActive Publication Date: 2025-08-19FANUC LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010910848.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-05
Filing Date
2020-09-02
Publication Date
2025-08-19
Estimated Expiration
2040-09-02

AI Technical Summary

Technical Problem

When a plurality of servo control units drive motors of multiple shafts, driving of one shaft will cause interference to other shafts, resulting in a decrease in command follow-up.

Method used

The machine learning device is used to perform machine learning on the disturbed servo control unit, correct position deviation, speed command and torque command through functions, and use reinforcement learning to update the value function to correct interference, including state information acquisition, behavior information output and value function update.

Benefits of technology

Effectively correcting interaxial interference, improving the command follow-up of the servo control unit and simplifying the complicated adjustment process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112445181B_ABST
    Figure CN112445181B_ABST
Patent Text Reader

Abstract

The present invention provides a machine learning device, a servo control device, a servo control system, and a machine learning method. The machine learning device performs machine learning on multiple servo control units corresponding to multiple axes. The first servo control unit associated with the axis subject to interference includes a correction unit that corrects at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function containing at least one of a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference. The machine learning device obtains state information including first servo control information, second servo control information, and function coefficients, outputs behavior information including coefficient adjustment information to the correction unit, outputs a reward value from reinforcement learning using an evaluation function that is a function of the first servo control information, and updates a value function based on the reward value, state information, and behavior information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning device for performing machine learning on multiple servo control units that control multiple motors, a servo control device including the machine learning device, a servo control system, and a machine learning method, wherein the multiple motors drive a machine in which one of multiple axes is disturbed by the movement of at least one other axis. Background Art

[0002] For example, Patent Documents 1 and 2 describe devices including a plurality of servo control units that control a plurality of motors that drive a device having a plurality of axes.

[0003] Patent Document 1 describes a control device comprising: a first motor control unit that controls a first motor that drives a first axis associated with a machine tool, robot, or industrial machine; and a second motor control unit that controls a second motor that drives a second axis that rotates in a direction different from the first axis. Patent Document 1 also describes an evaluation program for evaluating the operational characteristics of the control device, which operates the first and second motor control units. The evaluation program operates the first and second motor control units so that the trajectory of a controlled object moving along the first and second axes driven by the first and second motors has at least the following shapes: a shape having a corner (corner) where the rotational directions of both the first and second motors do not reverse; and a shape depicting an arc where one of the first and second motors rotates in one direction and the other of the first and second motors rotates in a reversed direction.

[0004] Patent document 2 describes a position drive control system comprising: a position instruction control device; a position drive control device having a plurality of position drive control units arranged for each servo motor, which are given position instructions from the position instruction control device, and the position drive control system has a shared memory for storing control status data of each axis, and the position drive control unit has: an inter-axis correction speed and torque control unit, which obtains the control status data of other axes from the shared memory to calculate the inter-axis correction instruction value corresponding to the load change of other axes during synchronization and tuning control of multiple axes, and corrects the instruction value of the current axis by using the inter-axis correction instruction value calculated by the inter-axis correction speed and torque control unit.

[0005] Prior art literature

[0006] Patent Document 1: Japanese Patent Application Publication No. 2019-003404

[0007] Patent Document 2: Japanese Patent Application Laid-Open No. 2001-100819

[0008] When multiple motors driving multiple axes are controlled by multiple servo control units, when one servo control unit drives one axis, the drive of the one axis may interfere with the drive of other axes driven by other servo control units.

[0009] In order to improve the command following performance of the servo control unit on the side affected by the disturbance, it is desirable to correct the disturbance. Summary of the Invention

[0010] (1) A first aspect of the present disclosure provides a machine learning device that performs machine learning on a plurality of servo control units that control a plurality of motors that drive a machine having a plurality of axes, wherein one of the axes is disturbed by the motion of at least one other axis.

[0011] A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference;

[0012] The machine learning device has:

[0013] a state information acquisition unit configured to acquire state information including first servo control information of the first servo control unit, second servo control information of the second servo control unit, and coefficients of the function;

[0014] a behavior information output unit configured to output behavior information including adjustment information of the coefficient included in the state information to the correction unit;

[0015] a reward output unit that outputs a reward value in reinforcement learning using an evaluation function that is a function of the first servo control information; and

[0016] A value function updating unit updates a value function based on the reward value output by the reward output unit, the state information, and the behavior information.

[0017] (2) A second aspect of the present disclosure provides a servo control device comprising:

[0018] The machine learning device described in (1) above; and

[0019] a plurality of servo control units that control a plurality of motors that drive a machine having a plurality of axes, wherein one of the axes is disturbed by the motion of at least one other axis,

[0020] A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference;

[0021] The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

[0022] (3) A third aspect of the present disclosure provides a servo control system comprising:

[0023] The machine learning device described in (1) above; and

[0024] A servo control device comprising a plurality of servo control units for controlling a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, one of the plurality of axes being disturbed by the motion of at least one other axis,

[0025] A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference;

[0026] The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

[0027] (4) A fourth aspect of the present disclosure provides a machine learning method for a machine learning device, wherein the machine learning device performs machine learning on a plurality of servo control units that control a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis.

[0028] A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference;

[0029] In the machine learning method,

[0030] acquiring state information including first servo control information of the first servo control unit, second servo control information of the second servo control unit, and coefficients of the function;

[0031] outputting behavior information including adjustment information of the coefficient included in the state information to the correction unit,

[0032] outputting a reward value in reinforcement learning using an evaluation function, the evaluation function being a function of the first servo control information;

[0033] The value function is updated according to the reward value, the state information, and the behavior information.

[0034] Effects of the Invention

[0035] According to various aspects of the present disclosure, complicated adjustments can be avoided in the servo control unit related to the interfered axis, and the inter-axis interference can be corrected to improve the command followability. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 This is a block diagram showing a servo control device according to a first embodiment of the present disclosure.

[0037] Figure 2 This is a partial structural diagram of the spindle moving mechanism that moves the spindle of a four-axis machining center serving as a machine tool.

[0038] Figure 3 This is a partial structural diagram showing a table mechanism for mounting a workpiece of a five-axis machining center serving as a machine tool.

[0039] Figure 4 This is a block diagram showing a machine learning unit according to the first embodiment of the present disclosure.

[0040] Figure 5 Yes Figure 2 The characteristic diagram shown is a diagram showing the fluctuation of position feedback information before coefficient adjustment related to machine learning when driving a four-axis machining center.

[0041] Figure 6 Yes Figure 2 The characteristic diagram shown is a diagram showing changes in position feedback information after coefficient adjustment related to machine learning when driving a four-axis machining center.

[0042] Figure 7 Yes Figure 3 The diagram shows a characteristic diagram of changes in position feedback information of the rotation axis and the X axis before coefficient adjustment related to machine learning when driving a 5-axis machining center.

[0043] Figure 8 Yes Figure 3The diagram shows a characteristic diagram of changes in position feedback information of the rotation axis and the X axis after coefficient adjustment related to machine learning when driving a 5-axis machining center.

[0044] Figure 9 This is a flowchart illustrating the operation of the machine learning unit according to the first embodiment of the present disclosure.

[0045] Figure 10 This is a flowchart for explaining the operation of the optimization behavior information output unit of the machine learning unit according to the first embodiment of the present disclosure.

[0046] Figure 11 This is a block diagram showing a configuration example of a servo control system including a servo control device and a machine learning device.

[0047] Explanation of symbols

[0048] 10. 10-1 to 10-n Servo control device

[0049] 20, 20-1 to 20-n machine tools

[0050] 100, 200 servo control unit

[0051] 101, 201 subtractor

[0052] 102 Adder

[0053] 103, 202 Position control unit

[0054] 104, 203 Adder

[0055] 105, 204 Subtractor

[0056] 106, 205 speed control unit

[0057] 107 Adder

[0058] 108, 206 servo motors

[0059] 109, 207 rotary encoder

[0060] 110, 208 integrator

[0061] 111 Position deviation correction unit

[0062] 112 Speed command correction unit

[0063] 113 Torque command correction unit

[0064] 209 Position Feedforward Unit

[0065] 300 Machine Learning Department

[0066] 300-1~300-n Machine Learning Device

[0067] 400 Network DETAILED DESCRIPTION

[0068] Hereinafter, embodiments of the present disclosure will be described in detail using the accompanying drawings.

[0069] (First embodiment)

[0070] Figure 1 This is a block diagram showing a servo control device according to a first embodiment of the present disclosure.

[0071] like Figure 1 As shown, the servo control device 10 includes servo control units 100 and 200 and a machine learning unit 300. The machine learning unit 300 is a machine learning device. The machine learning unit 300 can be provided in the servo control unit 100 or the servo control unit 200. The machine tool 20 is driven by the servo control units 100 and 200.

[0072] While the machine tool 20 is used as an example of the control target of the servo control units 100 and 200, the controlled machine is not limited to a machine tool and may also be a robot, industrial machine, or the like. The servo control units 100 and 200 may also be provided as part of a machine tool, robot, industrial machine, or the like.

[0073] Servo control units 100 and 200 control two axes of machine tool 20. The machine tool may be, for example, a three-axis machining center, a four-axis machining center, or a five-axis machining center. The two axes may be, for example, two linear axes such as the Y-axis and the Z-axis, or a linear axis and a rotary axis such as the X-axis and the B-axis. The specific structure of machine tool 20 will be described later.

[0074] The servo control unit 100 includes a subtractor 101, an adder 102, a position control unit 103, an adder 104, a subtractor 105, a speed control unit 106, an adder 107, a servo motor 108, a rotary encoder 109, an integrator 110, a position deviation corrector 111, a speed command corrector 112, and a torque command corrector 113.

[0075] The servo control unit 200 includes a subtractor 201 , a position control unit 202 , an adder 203 , a subtractor 204 , a speed control unit 205 , a servo motor 206 , a rotary encoder 207 , an integrator 208 , and a position feedforward unit 209 .

[0076] The servo control unit 100 corresponds to a first servo control unit related to an axis receiving a disturbance, and the servo control unit 200 corresponds to a second servo control unit related to an axis providing a disturbance.

[0077] The difference between servo control unit 100 and servo control unit 200 is that servo control unit 100 includes a position deviation correction unit 111, a speed command correction unit 112, and a torque command correction unit 113. Position deviation correction unit 111, speed command correction unit 112, and torque command correction unit 113 are provided because when servo control unit 200 drives one axis of machine tool 20, the drive of that axis interferes with the drive of other axes driven by servo control unit 100. Servo control unit 100 corrects for the effects of the drive on that axis.

[0078] Figure 1 The position feedforward unit 209 is provided in the servo control unit 200 , but may not be provided. Furthermore, the position feedforward unit 209 may be provided in the servo control unit 100 or in both the servo control unit 100 and the servo control unit 200 .

[0079] The following further describes the various components of the servo control device 10 and the machine tool 20. First, the servo control unit 200 related to the axis that is subject to the disturbance will be described. The servo control unit 100 related to the axis that is subject to the disturbance will be described later.

[0080] <Servo Control Unit 200 >

[0081] A position command x is generated so that the pulse frequency can be changed according to a predetermined machining program by a higher-level control device or an external input device, thereby varying the speed of the servo motor 206. Position command x is a control command. Position command x is output to the subtractor 201, the position feedforward unit 209, the position deviation corrector 111, the speed command corrector 112, the torque command corrector 113, and the machine learning unit 300.

[0082] The subtractor 201 obtains the difference between the position command x and the detected position of the position feedback (position FB) (which becomes position feedback information x′), and outputs the difference as a position deviation to the position control unit 202 .

[0083] The position control unit 202 outputs a value obtained by multiplying the position gain Kp by the position deviation as a speed command to the adder 203 .

[0084] The adder 203 adds the speed command to the output value (position feedforward term) of the position feedforward unit 209 , and outputs the result to the subtracter 204 as a speed command for feedforward control.

[0085] The subtractor 204 obtains a difference between the output of the adder 203 and the speed detection value of the speed feedback, and outputs the difference as a speed deviation to the speed control unit 205 .

[0086] The speed control unit 205 adds a value obtained by multiplying the speed deviation by the integral gain K1v and an integrated value obtained by multiplying the speed deviation by the proportional gain K2v, and outputs the result as a torque command to the servo motor 206 .

[0087] The integrator 208 integrates the speed detection value output from the rotary encoder 207 and outputs a position detection value.

[0088] Rotary encoder 207 outputs the speed detection value as speed feedback information to subtractor 204. Integrator 208 calculates a position detection value from the speed detection value and outputs this position detection value as position feedback (position FB) information x' to subtractor 201. Position feedback (position FB) information x' is also output to machine learning unit 300, position deviation correction unit 111, speed command correction unit 112, and torque command correction unit 113.

[0089] The rotary encoder 207 and the integrator 208 are detectors, and the servo motor 206 may be a motor that performs rotational motion or a linear motor that performs linear motion.

[0090] The position feedforward unit 209 multiplies a value obtained by differentiating the position command value and multiplying the value by a constant by the position feedforward coefficient, and outputs the multiplied value as a position feedforward term to the adder 203 .

[0091] The servo control unit 200 is configured as described above.

[0092] <Servo Control Unit 100 >

[0093] Position command y is generated so that the pulse frequency can be changed according to a predetermined machining program by a host control device or an external input device to change the speed of servo motor 108. Position command y is a control command and is output to subtractor 101 and machine learning unit 300.

[0094] The subtractor 101 obtains the difference between the position command y and the detected position of the position feedback (which becomes position feedback information y′), and outputs the difference as a position deviation to the adder 102 .

[0095] The adder 102 calculates the difference between the position deviation and the position deviation correction value output from the position deviation correcting unit 111 , and outputs the difference as the corrected position deviation to the position control unit 103 .

[0096] The position control unit 103 outputs a value obtained by multiplying the position gain Kp by the corrected position deviation as a speed command to the adder 104 .

[0097] The adder 104 calculates the difference between the speed command and the speed command correction value output from the speed command correction unit 112 , and outputs the difference as a corrected speed command to the subtracter 105 .

[0098] The subtracter 105 obtains a difference between the output of the adder 104 and the speed detection value of the speed feedback, and outputs the difference as a speed deviation to the speed control unit 106 .

[0099] The speed control unit 106 adds a value obtained by multiplying the speed deviation by the integral gain K1v and integrating the value obtained by multiplying the speed deviation by the proportional gain K2v, and outputs the result as a torque command to the adder 107.

[0100] The adder 107 calculates the difference between the torque command and the torque command correction value output from the torque command correction unit 113 , and outputs the difference as a corrected torque command to the servo motor 108 .

[0101] The integrator 110 integrates the speed detection value output from the rotary encoder 109 to output a position detection value.

[0102] The rotary encoder 109 outputs the speed detection value as speed feedback information to the subtractor 105. The integrator 110 obtains a position detection value from the speed detection value and outputs the position detection value as position feedback information y' to the subtractor 101 and the machine learning unit 300.

[0103] The rotary encoder 109 and the integrator 110 are detectors, and the servo motor 108 may be a motor that performs rotational motion or a linear motor that performs linear motion.

[0104] The position deviation correction unit 111 receives the position feedback information x' output from the integrator 208 of the servo control unit 200, the position command x input to the servo control unit 200, and the change amount of the coefficients a1 to a6 of the function represented by the following mathematical formula 1 (hereinafter referred to as mathematical formula 1) output from the machine learning unit 300, and calculates the position deviation correction value Err using mathematical formula 1. comp And output to adder 102.

[0105]

Mathematical formula 1

[0106]

[0107] The speed command correction unit 112 receives the position feedback information x' output from the integrator 208 of the servo control unit 200, the position command x input to the servo control unit 200, and the change amount of the coefficients b1 to b6 of the function represented by the following mathematical formula 2 (hereinafter referred to as mathematical formula 2) output from the machine learning unit 300, and calculates the speed command correction value Vcmd using mathematical formula 2. compAnd output to adder 104.

[0108]

Mathematical formula 2

[0109]

[0110] The torque command correction unit 113 receives the position feedback information x' output from the integrator 208 of the servo control unit 200, the position command x input to the servo control unit 200, and the change amount of the coefficients c1 to c6 of the function represented by the following mathematical formula 3 (hereinafter referred to as mathematical formula 3) output from the machine learning unit 300, and calculates the torque command correction value Tcmd using mathematical formula 3. comp And output to adder 107.

[0111]

Mathematical formula 3

[0112]

[0113] The position deviation correction unit 111, the speed instruction correction unit 112 and the torque instruction correction unit 113 correspond to correction units, which use the position instruction x and the position feedback information x' of the servo control unit 200 to generate the correction value Err of the position deviation of the servo control unit 100. comp , speed command correction value Vcmd comp And the correction value of the torque command Tcmd comp The servo control unit 100 sets the correction value Err to the position deviation, speed command, and torque command regardless of the direction. comp , speed command correction value Vcmd comp And the correction value of the torque command Tcmd comp The scalar value is added. In this way, the interference associated with the axis driven by the servo control unit 200 can be eliminated from the position deviation, speed command, and torque command of the servo control unit 100. It is not necessary to provide all of the position deviation correction unit 111, speed command correction unit 112, and torque command correction unit 113; one or two of these units can be provided as needed.

[0114] In addition, mathematical formulas 1 to 3 are formulas that include position instruction x, the first-order differential of position instruction x, the second-order differential of position instruction x, position feedback information x', the first-order differential of position feedback information x', and the second-order differential of position feedback information x' as variables. However, mathematical formulas 1 to 3 do not have to include all of these variables, and one or more of them can be appropriately selected. For example, the correction value Err of the position deviation can be calculated from the second-order differential of position instruction x and the second-order differential of position feedback information x', that is, the acceleration of position instruction x and the acceleration of position feedback information x'. comp , speed command correction value Vcmd comp And the correction value of the torque command Tcmd comp .

[0115] The position command x, the first-order differential of the position command x, and the second-order differential of the position command x are variables related to the position command, and the position feedback information x', the first-order differential of the position feedback information x', and the second-order differential of the position feedback information x' are variables related to the position feedback information.

[0116] As described above, the servo control unit 100 is configured.

[0117] <Machine Tool 20>

[0118] The machine tool 20 is, for example, a three-axis machining center, a four-axis machining center, or a five-axis machining center.

[0119] Figure 2 This is a partial diagram of the main spindle moving mechanism that moves the main spindle of a 4-axis machining center. Figure 3 This diagram shows a partial configuration of the workpiece-carrying table mechanism of a 5-axis machining center.

[0120] Machine tool 20 is Figure 2 In the case of the four-axis machine tool 20A shown, for example, the servo control unit 200 controls the linear motion of the Y axis, and the servo control unit 100 controls the linear motion of the Z axis. In this case, the servo control unit 200 is the servo control unit for the axis that applies the disturbance, and the servo control unit 100 is the servo control unit for the axis that receives the disturbance.

[0121] like Figure 2 As shown, the X-axis movable table 22 is mounted on the stationary table 21 so as to be movable in the X-axis direction, and the Y-axis movable column 23 is mounted on the X-axis movable table 22 so as to be movable in the Y-axis direction. Furthermore, a spindle mounting table 24 is mounted on the side of the Y-axis movable column 23, and the spindle 25 is mounted on the spindle mounting table 24 so as to be rotatable relative to the B-axis and movable in the Z-axis direction. For example, when the Y-axis movable column 23 accelerates or decelerates in the Y-axis direction, the drive of the spindle 25 in the Z-axis direction is interfered with by the Y-axis.

[0122] Machine tool 20 is Figure 3In the case of the five-axis machining center 20B shown in FIG. 1 , for example, the servo control unit 200 controls the rotation of the rotary axis, and the servo control unit 100 controls the linear movement of the X-axis as the linear axis. Figure 3 As shown in FIG, when the rotating axis of the rotary indexing table 28 with centrifugal load is arranged on the linear axis, they affect each other and cause interference. In order to eliminate this interference, a correction unit is provided in at least one of the servo control unit 100 and the servo control unit 200. Here, the servo control unit 100 is provided with a position deviation correction unit 111, a speed instruction correction unit 112, and a torque instruction correction unit 113 as correction units. Figure 2 As in the four-axis machining center 20A shown, the servo control unit 200 is the servo control unit for the axis that applies the disturbance, and the servo control unit 100 is the servo control unit for the axis that receives the disturbance. The position command input to the servo control unit 200 is a command that specifies the rotation angle of the rotation axis.

[0123] like Figure 3 As shown, an X-axis movable table 27 is mounted on a stationary table 26 so as to be movable in the X-axis direction, and a rotary indexing table 28 is rotatably mounted on the X-axis movable table 27. Due to the influence of a workpiece or a workpiece holder mounted on the rotary indexing table 28, an eccentric load 29 may be generated at a position offset from the center of the rotation axis. When eccentric load 29 is generated, interference occurs between the X-axis movable table 27 and the rotary indexing table 28.

[0124] In addition, regarding the structures of the servo control unit 100 and the servo control unit 200, Figure 2 When the 4-axis machining center 20A shown is driven, Figure 3 The structure is the same when the 5-axis machining center 20B shown in the figure is driven. The values of the coefficients a1 to a6 of the mathematical formula 1 of the position deviation correction unit 111 of the servo control unit 100, the coefficients b1 to b6 of the mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of the mathematical formula 3 of the torque instruction correction unit 113 are the values of the interference given by the Y axis to the Z axis. Figure 2 The 4-axis machining center 20A shown in FIG. 1 has a rotation axis and an X-axis that interfere with each other. Figure 3 The five-axis machining centers 20B shown are different from each other.

[0125] <Machine Learning Department 300>

[0126] The machine learning unit 300 executes a pre-set machining program (hereinafter also referred to as the "learning machining program") and uses the position command y and position feedback (position FB) information y output from the servo control unit 100 to perform machine learning (hereinafter referred to as learning) on the coefficients a1 to a6 of the position deviation correction unit 111, the coefficients b1 to b6 of the speed command correction unit 112, and the coefficients c1 to c6 of the torque command correction unit 113. The machine learning unit 300 is a machine learning device. The learning performed by the machine learning unit 300 can be performed before or after shipment.

[0127] Hereinafter, a four-axis machining center 20A is used as the machine tool 20. The servo control unit 200 controls the servo motor 206 according to the machining program used during learning. The servo motor 206 drives the Y-axis of the four-axis machining center 20A. Furthermore, the servo control unit 100 controls the servo motor 108 according to the machining program used during learning. The servo motor 108 drives the Z-axis of the four-axis machining center 20A.

[0128] When learning the machining program for driving the four-axis machining center 20A, the Y-axis can be reciprocated by controlling the servo control unit 200 for the axis causing the interference, while the Z-axis can be reciprocated or not reciprocated by controlling the servo control unit 100 for the axis receiving the interference. The following description will describe the case where the Z-axis is not moved.

[0129] Based on the machining program during learning, the host control device or external input device outputs a position command to the servo control unit 200 for reciprocating Y-axis motion and a position command to the servo control unit 100 for stationary Z-axis motion. However, even if the position command to stationary Z-axis is input, the disturbance caused by the movement of the Y-axis in the servo control unit 100 affects the position deviation, speed command, and torque command of the servo control unit 100. Therefore, the machine learning unit 300 learns the coefficients a1 to a6 of the position deviation correction unit 111, the coefficients b1 to b6 of the speed command correction unit 112, and the coefficients c1 to c6 of the torque command correction unit 113, thereby setting the correction values of the position deviation, speed command, and torque command to the optimal values.

[0130] The machine learning unit 300 will be described in more detail below.

[0131] In the following description, the case where the machine learning unit 300 performs reinforcement learning is described. However, the learning performed by the machine learning unit 300 is not particularly limited to reinforcement learning. For example, the present invention can also be applied to the case where supervised learning is performed.

[0132] Before explaining the functional blocks within the machine learning unit 300, we'll first explain the basic structure of reinforcement learning. An agent (equivalent to the machine learning unit 300 in this embodiment) observes the state of its environment, selects a behavior, and the environment changes based on that behavior. As the environment changes, rewards are provided, and the agent learns to choose better behaviors (decisions).

[0133] Supervised learning involves a completely correct answer, while the rewards in reinforcement learning are mostly episodic values based on partial changes in the environment. Therefore, the agent learns to choose actions that maximize the total reward in the future.

[0134] In this way, reinforcement learning learns appropriate behaviors based on the interaction between behaviors and the environment, that is, it learns the method to be learned to maximize future rewards. This means that in this embodiment, it is possible to obtain behavioral information that influences future behaviors, such as selecting behavior information for correcting inter-axis interference in the servo control unit associated with the disturbed axis.

[0135] Here, any learning method can be used as reinforcement learning. In the following description, the case of using Q-learning (Q-learning) under a certain environmental state S is used as an example. The Q-learning is a method of learning the value Q(S, A) of selecting behavior A.

[0136] The purpose of Q-learning is to select the behavior A with the highest value Q(S, A) as the optimal behavior from among the possible behaviors A in a certain state S.

[0137] However, at the beginning of Q-learning, the correct value Q(S, A) for the combination of state S and action A is completely unknown. Therefore, the agent chooses various actions A in a certain state S, and selects the best action based on the reward given to the current action A, thereby continuing to learn the correct value Q(S, A).

[0138] Furthermore, we want to maximize the total amount of future rewards, so the goal is to eventually achieve Q(S, A) = E[Σ(γ t )r t Here, E[] represents the expected value, t represents the time, γ represents a parameter called the discount rate described later, and r t represents the reward at time t, and Σ is the sum at time t. The expected value in this equation is the expected value when the optimal behavior state changes. However, in the Q-learning process, since the optimal behavior is unknown, reinforcement learning is performed by performing various behaviors while searching. The updated formula for the value Q(S, A) can be expressed, for example, by the following mathematical formula 4 (hereinafter referred to as Formula 4).

[0139]

Mathematical formula 4

[0140]

[0141] In the above mathematical formula 4, S t represents the environmental state at time t, A t Represents the behavior at time t. Through behavior A t , the state changes to S t+1 . r t+1 Indicates the reward obtained by changing the state. In addition, the term with max is: t+1 In the following example, we multiply γ by the Q value when we select the action A with the highest known Q value. Here, γ is a parameter with a range of 0 < γ ≤ 1, and is called the discount rate. Furthermore, α is a learning coefficient, and the range of α is 0 < α ≤ 1.

[0142] The above mathematical formula 4 represents the following method: According to the experiment A t The result is the feedback r t+1 , update the state S t The following behavior A t The value of Q(S t 、A t ).

[0143] This update formula shows that if behavior A t The next state S t+1 The value of the best behavior under a Q(S t+1 , A) than state S t The following behavior A t The value of Q(S t 、A t ) is large, then Q(S t 、A t ), if it is small, reduce Q(S t 、A t ). That is, the value of a certain behavior in a certain state is close to the optimal behavior value in the next state caused by the behavior. Although the difference is due to the discount rate γ and the reward r t+1 The existence of the structure varies, but it is basically a structure in which the best behavior value in a certain state is propagated to the behavior value in its previous state.

[0144] Here, there is a Q-learning method that creates a table of Q(S, A) for all state-action pairs (S, A) and performs learning. However, in some cases, the number of states required to find the values of Q(S, A) for all state-action pairs is too large, making Q-learning convergence time-consuming.

[0145] Therefore, the well-known technology known as DQN (Deep Q-Network) can be utilized. Specifically, a suitable neural network is used to construct the value function Q. By adjusting the parameters of the neural network, the value function Q can be approximated by the appropriate neural network to calculate the value Q(S, A). Using DQN can shorten the time required for Q learning to converge. DQN is described in detail in, for example, the following non-patent literature.

[0146] <Non-patent literature>

[0147] "Human-level control through deep reinforcement learning," Volodymyr Mnih1, [online], [retrieved January 17, 2011], Internet (URL: http: / / files.davidqiu.com / research / nature14236.pdf)

[0148] The machine learning unit 300 performs the Q learning described above. Specifically, the machine learning unit 300 obtains a set of position instructions x and a set of position feedback information x' from the servo control unit 200 by executing the machining program during learning. Furthermore, the machine learning unit 300 obtains a set of position instructions y and a set of position feedback information y' from the servo control unit 100 by executing the machining program during learning. The position instructions y and the position feedback information y' constitute the first servo control information, while the position instructions x and the position feedback information x' constitute the second servo control information. The position instruction y instructs the Z axis to remain stationary. The set of position instructions x and the set of position feedback information x', and the set of position instructions y and the set of position feedback information y' constitute the state S. In addition, the machine learning unit 300 learns the following value Q: the adjustment of the values of the coefficients a1 to a6 of the mathematical formula 1 of the position deviation correction unit 111 of the servo control unit 100, the coefficients b1 to b6 of the mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of the mathematical formula 3 of the torque instruction correction unit 113 related to the state S is selected as behavior A.

[0149] The servo control unit 200 executes the machining program for learning and performs servo control on the servo motor 206 that drives the Y-axis. Furthermore, the servo control unit 100 executes the machining program for learning and performs servo control on the servo motor 108 to correct the position deviation, speed command, and torque command using the position deviation correction value, speed command correction value, and torque command correction value calculated using Mathematical Formula 1 with coefficients a1 to a6, Mathematical Formula 2 with coefficients b1 to b6, and Mathematical Formula 3 with coefficients c1 to c6, and to bring the Z-axis to a standstill based on the position command.

[0150] The machine learning unit 300 observes information about a state S, which includes a set of position instructions x and a set of position feedback information x' obtained by executing the machining program during learning, and a set of position instructions y and a set of position feedback information y' obtained by executing the machining program during learning, and determines an action A. The machine learning unit 300 provides a reward each time an action A is performed. The machine learning unit 300 searches for the optimal action A, for example, by trial and error, so as to maximize the total reward in the future. In this way, the machine learning unit 300 can select the optimal action A (i.e., coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6) for the state S, which includes a set of position instructions x and a set of position feedback information x' obtained by executing the machining program during learning, and a set of position instructions y and a set of position feedback information y' obtained by executing the machining program during learning based on the coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6.

[0151] That is, based on the value function Q learned by the machine learning unit 300, the behavior A with the largest Q value among the coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6 applied to a certain state S is selected, thereby selecting the behavior A (i.e., coefficients a1 to a6, coefficients b1 to b6, coefficients c1 to c6) that corrects the inter-axis interference generated by the machining program when executing the learning.

[0152] Figure 4 This is a block diagram showing the machine learning unit 300 according to the first embodiment of the present disclosure.

[0153] In order to carry out the above reinforcement learning, Figure 4 As shown, the machine learning unit 300 includes a state information acquisition unit 301, a learning unit 302, a behavior information output unit 303, a value function storage unit 304, and an optimized behavior information output unit 305. The learning unit 302 includes a reward output unit 3021, a value function update unit 3022, and a behavior information generation unit 3023.

[0154] The state information acquisition unit 301 acquires the state S, which includes the set of position instructions x and position feedback information x' (which becomes the second servo control information) of the servo control unit 200 obtained by executing the machining program during learning, and the set of position instructions y and position feedback information y' of the servo control unit 100 obtained by executing the machining program during learning based on the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque instruction correction unit 113. This state information S is equivalent to the environmental state S in Q learning. In addition, Figure 4In , coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6 are expressed as coefficients a, b, and c for simplicity.

[0155] The state information acquisition unit 301 outputs the acquired state information S to the learning unit 302 .

[0156] Furthermore, at the time when Q learning is first initiated, the coefficients a1 to a6 of mathematical formula 1 in the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 in the speed command correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 in the torque command correction unit 113 are previously generated by the user. In this embodiment, the initial setting values of the coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6 generated by the user are adjusted to the optimal values through reinforcement learning.

[0157] Furthermore, when the operator has adjusted the machine tool in advance, the adjusted values can be used as initial values to perform machine learning on the coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6.

[0158] The learning unit 302 is a part that learns the value Q(S, A) when a certain behavior A is selected under a certain environmental state S.

[0159] The reward output unit 3021 calculates a reward when action A is selected in a certain state S. State S' represents a state changed from state S by action A (correction of coefficients a1 to a6, coefficients b1 to b6, and coefficients c1 to c6).

[0160] The report output unit 3021 calculates the difference (y-y') between the position command y and the position feedback information y' in states S and S'. In the report output unit 3021, the position deviation calculated from the difference between the position command y and the position feedback information y' is referred to as the second position deviation. The set of differences (y-y') is referred to as the position deviation set. The position deviation set in state S is denoted by PD(S), and the position deviation set in state S' is denoted by PD(S').

[0161] When the position deviation (y-y') of the servo control unit 100 of the disturbed axis is represented by the position deviation e as the evaluation function f, for example, the following can be applied:

[0162] Function to calculate the integral of the absolute value of the position deviation

[0163] ∫|e|dt

[0164] Function for calculating the time-weighted integral of the absolute value of the position deviation

[0165] ∫t|e|dt

[0166] Function that calculates the integral value of the absolute value of the position deviation raised to the power of 2n (n is a natural number)

[0167] ∫e 2n dt (n is a natural number)

[0168] Function to calculate the maximum absolute value of position deviation

[0169] Max{|e|}.

[0170] The value of the evaluation function f obtained from the position difference set PD(S) is defined as the evaluation function value f(PD(S)), and the value of the evaluation function f obtained from the position difference set PD(S') is defined as the evaluation function value f(PD(S')).

[0171] The position command y input to the servo control unit 100 is not a command to stop the Z axis but a command to reciprocate the Z axis. The evaluation function can use the above-mentioned evaluation function f.

[0172] At this time, when the evaluation function value f(PD(S')) of the servo control unit 100 when operating based on the corrected position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113 related to the state information S' corrected by the behavior information A is greater than the evaluation function value f(PD(S)) of the servo control unit 100 when operating based on the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113 before the correction, the feedback output unit 3021 makes the feedback value negative.

[0173] On the other hand, when the evaluation function value f(PD(S′)) is smaller than the evaluation function value f(PD(S)), the reward output unit 3021 makes the reward value a positive value.

[0174] When the evaluation function f(PD(S′)) is equal to the evaluation function value f(PD(S)), the reward output unit 3021 sets the reward value to zero.

[0175] Furthermore, when the evaluation function value f(PD(S')) of the state S' after the execution of behavior A is greater than the evaluation function value f(PD(S)) in the previous state S, the absolute value of the negative value can be set larger according to the ratio. In other words, the absolute value of the negative value is increased according to the degree to which the value of f(PD(S')) increases. Conversely, when the evaluation function value f(PD(S')) of the state S' after the execution of behavior A is smaller than the evaluation function value f(PD(S)) in the previous state S, the positive value is set larger according to the ratio. In other words, the positive value is increased according to the degree to which the value of f(PD(S')) decreases.

[0176] The value function update unit 3022 performs Q learning based on the state S, behavior A, the state S′ when behavior A is applied to the state S, and the reward value calculated as described above, thereby updating the value function Q stored in the value function storage unit 304.

[0177] The update of the value function Q can be performed through online learning, batch learning, or mini-batch learning.

[0178] Online learning is a learning method that applies a certain action A to the current state S, immediately updating the value function Q each time the state S transitions to a new state S'. Batch learning, on the other hand, collects learning data by repeatedly applying a certain action A to the current state S, transitioning the state S to a new state S', and then updates the value function Q using all the collected learning data. Furthermore, mini-batch learning is a learning method intermediate between online and batch learning, updating the value function Q each time a certain amount of learning data has accumulated.

[0179] The behavior information generation unit 3023 selects behavior A during the Q learning process for the current state S. During the Q learning process, the behavior information generation unit 3023 generates behavior information A to correct the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed command correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque command correction unit 113 (equivalent to behavior A in Q learning), and outputs the generated behavior information A to the behavior information output unit 303. More specifically, the behavior information generating unit 3023, for example, adds or subtracts increments to the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque instruction correction unit 113 included in the state S, as well as to the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque instruction correction unit 113 included in the behavior A.

[0180] In addition, the following strategy can be adopted: when the behavior information generation unit 3023 transfers to state S' by increasing or decreasing the application coefficients a1~a6, coefficients b1~b6, and coefficients c1~c6 and gives a positive reward (positive reward), the next behavior A' is to add or subtract increments to the coefficients a1~a6, coefficients b1~b6, and coefficients c1~c6 in the same way as the previous action, and select a behavior A' with a smaller value of the evaluation function f.

[0181] In addition, the following strategy can be adopted conversely: when a negative reward (a reward with a negative value) is given, the behavior information generation unit 3023 selects the behavior A' whose evaluation function is smaller than the previous value as the next behavior A', for example, by subtracting or adding increments to the coefficients a1~a6, coefficients b1~b6, and coefficients c1~c6 in the opposite direction of the previous action.

[0182] In addition, the behavior information generation unit 3023 can also adopt the following strategy: select behavior A' through a greedy algorithm that selects behavior A' with the highest value Q(S, A) among the values of the currently estimated behavior A, randomly select behavior A' with a certain smaller probability ε, and select the behavior A' with the highest value Q(S, A) through the ε greedy algorithm, etc., which are well-known methods.

[0183] The behavior information output unit 303 transmits the behavior information A output from the learning unit 302 to the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113. As described above, the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113 make fine adjustments to the current state S, i.e., the currently set coefficients a1 to a6, b1 to b6, and c1 to c6, based on this behavior information, and transition to the next state S' (i.e., the corrected coefficients a1 to a6 of Formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of Formula 2 of the speed command correction unit 112, and the coefficients c1 to c6 of Formula 3 of the torque command correction unit 113).

[0184] The value function storage unit 304 is a storage device that stores the value function Q. The value function Q is stored, for example, as a table (hereinafter referred to as a behavior value table) corresponding to state S and behavior A. The value function Q stored in the value function storage unit 304 is updated by the value function update unit 3022. Furthermore, the value function Q stored in the value function storage unit 304 can be shared among other machine learning units 300. Sharing the value function Q among multiple machine learning units 300 allows reinforcement learning to be performed in a distributed manner by each machine learning unit 300, thereby improving the efficiency of reinforcement learning.

[0185] The optimized behavior information output unit 305 generates behavior information A (hereinafter referred to as "optimized behavior information") for causing the position deviation correction unit 111, the speed instruction correction unit 112, and the torque instruction correction unit 113 to perform an action with the maximum value Q (S, A) based on the updated value function Q performed by the value function update unit 3022 through Q learning.

[0186] More specifically, the optimized behavior information output unit 305 obtains the value function Q stored in the value function storage unit 304. As described above, this value function Q is updated by the value function update unit 3022 through Q learning. Furthermore, the optimized behavior information output unit 305 generates behavior information based on the value function Q and outputs this generated behavior information to the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113. This optimized behavior information, like the behavior information output by the behavior information output unit 303 during Q learning, includes information for correcting the coefficients a1 to a6 of Mathematical Formula 1 in the position deviation correction unit 111, the coefficients b1 to b6 of Mathematical Formula 2 in the speed command correction unit 112, and the coefficients c1 to c6 of Mathematical Formula 3 in the torque command correction unit 113.

[0187] The position deviation correction unit 111 , the speed command correction unit 112 , and the torque command correction unit 113 correct the coefficients a1 to a6 , the coefficients b1 to b6 , and the coefficients c1 to c6 based on the behavior information.

[0188] The machine learning unit 300 can perform the above actions to optimize the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque instruction correction unit 113, correct the inter-axis interference, and improve the instruction tracking performance.

[0189] Figure 5 Yes Figure 2 The characteristic diagram shown is a diagram showing a change in position feedback (position FB) information before coefficient (parameter) adjustment related to machine learning when the four-axis machining center 20A is driven. Figure 6 Yes Figure 2 The characteristic diagram shown is a diagram showing changes in position feedback (position FB) information after coefficient (parameter) adjustment related to machine learning when the four-axis machining center 20A is driven.

[0190] Figure 5 and Figure 6 The figure shows the change of the position feedback information of the servo control unit 100 when the servo control units 200 and 100 are driven in such a manner that the Y axis moves back and forth and the Z axis is stationary. Figure 6 As shown in the characteristic diagram, it can be seen that through the adjustment of the coefficients (parameters) involved in machine learning Figure 5 The position change of the characteristic diagram is improved and the command follow-up performance is enhanced.

[0191] Figure 7 Yes Figure 3The characteristic diagram shown is a diagram showing changes in position feedback (position FB) information of the rotation axis and the X axis before adjustment of coefficients (parameters) related to machine learning when the five-axis machining center 20B is driven. Figure 8 Yes Figure 3 The characteristic diagram shown is a diagram showing changes in position feedback (position FB) information of the rotation axis and the X axis after coefficients (parameters) related to machine learning are adjusted when the 5-axis machining center 20B is driven. Figure 7 and Figure 8 In the figure, the right vertical axis represents the value of the position feedback (position FB) information of the X-axis as the linear axis, and the left vertical axis represents the value of the position feedback (position FB) information of the rotary axis.

[0192] Figure 7 and Figure 8 100 and 100 are driven so that the rotation axis rotates and the X axis is stationary. Figure 8 As shown in the characteristic diagram, it can be seen that through the adjustment of the coefficients (parameters) involved in machine learning Figure 7 The position variation of the X-axis of the characteristic diagram is improved, and the command follow-up performance is enhanced.

[0193] As described above, by using the machine learning unit 300 according to the present embodiment, the coefficient adjustment of the position deviation correction unit 111 , the speed command correction unit 112 , and the torque command correction unit 113 can be simplified.

[0194] The functional blocks included in the servo control device 10 have been described above.

[0195] To implement these functional blocks, the servo control device 10 includes a CPU (Central Processing Unit) and other processing units. Furthermore, the servo control device 10 includes auxiliary storage devices such as HDDs (Hard Disk Drives) for storing various control programs, including application software and an OS (Operating System), and main storage devices such as RAM (Random Access Memory) for storing data temporarily needed by the processing unit after executing the programs.

[0196] Furthermore, the arithmetic processing unit in the servo control device 10 reads application software or an operating system from the auxiliary storage device, expands the read application software or operating system on the main storage device, and performs arithmetic processing based on the application software or operating system. Furthermore, the various hardware components of each device are controlled based on the results of these arithmetic operations. This realizes the functional blocks of this embodiment. In other words, this embodiment can be implemented through the collaboration of hardware and software.

[0197] The machine learning unit 300 utilizes a technology called GPGPU (General-Purpose Computing on Graphics Processing Units) (GPGPUs), which utilize GPUs in personal computers to perform high-speed processing of machine learning operations. Furthermore, to achieve even higher processing speeds, multiple computers equipped with such GPUs can be used to form a computer cluster, allowing parallel processing to be performed across the multiple computers within the cluster.

[0198] Next, refer to Figure 9 The operation of the machine learning unit 300 during Q learning in this embodiment will be described in the following process.

[0199] In step S11, the state information acquisition unit 301 acquires initial state information S0 from the servo control units 100 and 200. The acquired state information is output to the value function update unit 3022 or the behavior information generation unit 3023. As described above, this state information S is information corresponding to the state in Q learning.

[0200] At the time Q learning first begins, the set of position commands x and y in state S0 is obtained from a higher-level control device, an external input device, or the servo control unit 200 and the servo control unit 100. The set of position feedback information x' and y' in state S0 is obtained by operating the servo control unit 100 and the servo control unit 200 according to the machining program during learning. The set of position commands x input to the servo control unit 200 is a command for reciprocating the Y axis, while the set of position commands y input to the servo control unit 100 is a command for stationary Z axis. The position command x is input to the position feedforward unit 209, the subtractor 201, the position deviation corrector 111, the velocity command corrector 112, the torque command corrector 113, and the machine learning unit 300. The position command y is input to the subtractor 101 and the machine learning unit 300. The initial values for the coefficients a1 to a6 of Mathematical Formula 1 in the position deviation correction unit 111, the coefficients b1 to b6 of Mathematical Formula 2 in the speed command correction unit 112, and the coefficients c1 to c6 of Mathematical Formula 3 in the torque command correction unit 113 are pre-generated by the user. These initial values for the coefficients a1 to a6, b1 to b6, and c1 to c6 are then supplied to the machine learning unit 300. For example, the initial values may be: all coefficients a1 to a6 are set to 0; all coefficients b1 to b6 are set to 0; and all coefficients c1 to c6 are set to 0. Furthermore, the machine learning unit 300 may also extract the set of position commands x and position feedback information x', and the set of position commands y and position feedback information y' in the aforementioned state S0.

[0201] In step S12, the behavior information generation unit 3023 generates new behavior information A and outputs the generated new behavior information A to the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113 via the behavior information output unit 303. The behavior information generation unit 3023 outputs the new behavior information A according to the above-described strategy. Furthermore, upon receiving the behavior information A, the servo control unit 100 drives the machine tool including the servo motor 108 using a state S' in which the coefficients a1 to a6, b1 to b6, and c1 to c6 of the position deviation correction unit 111, the speed command correction unit 112, and the torque command correction unit 113 related to the current state S are corrected based on the received behavior information. As described above, this behavior information corresponds to behavior A in Q learning. The current state S is state S0 at the initial start of Q learning.

[0202] In step S13, the state information acquisition unit 301 acquires the set of position commands x and position feedback information x', the set of position commands y and position feedback information y', and the coefficients a1 to a6, b1 to b6, and c1 to c6 for the new state S'. Thus, the state information acquisition unit 301 acquires the set of position commands x and position feedback information x', and the set of position commands y and position feedback information y' for the state S' with the coefficients a1 to a6, b1 to b6, and c1 to c6. The acquired state information is output to the report output unit 3021.

[0203] In step S14, the reward output unit 3021 determines the magnitude relationship between the evaluation function value f(PD(S')) in state S' and the evaluation function value f(PD(S)) in state S. When f(PD(S'))>f(PD(S)), the reward is made negative in step S15. When f(PD(S'))<f(PD(S)), the reward is made positive in step S16. When f(PD(S'))=f(PD(S)), the reward is made zero in step S17. In addition, the negative and positive values of the reward can be weighted. In addition, the state S is the state S0 at the time of starting Q learning.

[0204] When any of steps S15, S16, and S17 is completed, in step S18, the value function updater 3022 updates the value function Q stored in the value function storage unit 304 based on the reward value calculated in that step. The process then returns to step S12, repeating the above process until the value function Q converges to an appropriate value. Alternatively, the process can be terminated upon repetition of the above process a predetermined number of times or for a predetermined period of time.

[0205] In addition, although step S18 exemplifies online updating, online updating may be replaced by batch updating or small batch updating.

[0206] Above, by reference Figure 9 The described action, in this embodiment, obtains the following effect by utilizing the machine learning unit 300: an appropriate value function for adjusting the coefficients a1 to a6 of mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of mathematical formula 3 of the torque instruction correction unit 113 can be obtained, which can simplify the optimization of the coefficients a1 to a6, the coefficients b1 to b6, and the coefficients c1 to c6.

[0207] Next, refer to Figure 10 The process of generating the optimization behavior information by the optimization behavior information output unit 305 is described.

[0208] First, in step S21, the optimization behavior information output unit 305 obtains the value function Q stored in the value function storage unit 304. As described above, the value function Q is a function updated by the value function update unit 3022 through Q learning.

[0209] In step S22 , the optimization behavior information output unit 305 generates optimization behavior information based on the cost function Q, and outputs the generated optimization behavior information to the servo control unit 100 .

[0210] In addition, by referring to Figure 10 The described action, in this embodiment, can generate optimization behavior information based on the value function Q obtained by learning by the machine learning unit 300. Based on the optimization behavior information, the adjustment of the coefficients a1 to a6 of the mathematical formula 1 of the currently set position deviation correction unit 111, the coefficients b1 to b6 of the mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of the mathematical formula 3 of the torque instruction correction unit 113 can be simplified, which can improve the quality of the workpiece processing surface.

[0211] The various components included in the servo control device and the machine learning unit described above can be implemented using hardware, software, or a combination thereof. Furthermore, the servo control method performed by the various components included in the servo control device can also be implemented using hardware, software, or a combination thereof. Here, "implemented using software" means that the computer executes the program by reading it.

[0212] Various types of non-transitory computer readable recording media (non-transitory computer readable medium) can be used to store the program and provide the program to the computer. Non-transitory computer readable recording media include various types of tangible recording media (tangible storage medium). Examples of non-transitory computer readable recording media include: magnetic recording media (e.g., hard disk drives), optical-magnetic recording media (e.g., optical magnetic disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (random access memory)).

[0213] The above-described embodiment is a preferred embodiment of the present invention. However, the scope of the present invention is not limited to the above-described embodiment, and the present invention can be implemented in various modified forms without departing from the spirit of the present invention.

[0214] For example, in the above embodiment, the machine learning unit 300 calculates the value of the evaluation function f by calculating the difference between the position command y and the position feedback information y' of the servo control unit 100 for the axis subject to interference. However, the value of the evaluation function f may also be calculated using the position deviation (y-y'), which is the output of the subtractor 101 of the servo control unit 100. The position deviation (y-y'), which is the output of the subtractor 101 of the servo control unit 100, is the first position deviation.

[0215] In addition, in the above embodiment, an example is described in which the machine learning unit 300 simultaneously learns the coefficients a1 to a6 of the mathematical formula 1 of the position deviation correction unit 111, the coefficients b1 to b6 of the mathematical formula 2 of the speed instruction correction unit 112, and the coefficients c1 to c6 of the mathematical formula 3 of the torque instruction correction unit 113. However, the machine learning unit 300 may first learn and optimize one of the coefficients a1 to a6, the coefficients b1 to b6, and the coefficients c1 to c6, and then learn and optimize the other coefficients in sequence.

[0216] Furthermore, in the above embodiment, the feedback output unit 3021 of the machine learning unit 300 uses the position deviation as the evaluation function, but the velocity deviation or the acceleration deviation may also be used.

[0217] The speed deviation can be obtained from the time differential of the position deviation, and the acceleration deviation can be obtained from the time differential of the speed deviation. The speed deviation can be obtained by using the output of the adder 104 , that is, the difference between the speed command and the speed feedback information, or the output of the subtractor 105 .

[0218] (Second embodiment)

[0219] In the first embodiment, an example was described in which the machine learning unit was provided as part of the servo control device. However, in this embodiment, an example is described in which the machine learning unit is provided externally to the servo control device, thereby forming a servo control system. Hereinafter, since the machine learning unit is provided independently of the servo control device, it will be referred to as a machine learning device.

[0220] Figure 11 This is a block diagram showing a configuration example of a servo control system including a servo control device and a machine learning device. Figure 11 The illustrated servo control system 30 includes n (n is a natural number greater than or equal to 2) servo control devices 10-1 to 10-n, n machine learning devices 300-1 to 300-n, and a network 400 connecting the servo control devices 10-1 to 10-n and the n machine learning devices 300-1 to 300-n. The n (n is a natural number greater than or equal to 2) servo control devices 10-1 to 10-n are connected to n machine tools 20-1 to 20-n.

[0221] Each of the servo control devices 10-1 to 10-n has the same functions as the servo control devices 10-1 to 10-n except that the servo control devices 10-1 to 10-n do not have a machine learning unit. Figure 1 The machine learning devices 300-1 to 300-n have the same structure as the servo control device 10. Figure 5 The machine learning unit 300 shown has the same structure.

[0222] Here, the servo control device 10-1 and the machine learning device 300-1 are connected in a one-to-one pair for communication. The servo control devices 10-2 to 10-n and the machine learning devices 300-2 to 300-n are also connected in the same manner as the servo control device 10-1 and the machine learning device 300-1. Figure 11 In the embodiment, n groups of servo control devices 10-1 to 10-n and machine learning devices 300-1 to 300-n are connected via a network 400. In each of the n groups of servo control devices 10-1 to 10-n and machine learning devices 300-1 to 300-n, the servo control devices and machine learning devices of each group can be directly connected via a connection interface. Multiple groups of these n groups of servo control devices 10-1 to 10-n and machine learning devices 300-1 to 300-n can be installed in the same factory or in different factories, for example.

[0223] The network 400 is, for example, a LAN (Local Area Network) built in a factory, the Internet, a public telephone network, or a combination thereof. The specific communication method in the network 400 is not particularly limited to wired or wireless connection.

[0224] <Degrees of freedom of system structure>

[0225] In the above-mentioned embodiment, the servo control devices 10-1 to 10-n and the machine learning devices 300-1 to 300-n are respectively connected in a one-to-one group in a communicative manner, but, for example, one machine learning device can also be connected to multiple motor control devices and multiple acceleration sensors via the network 400 in a communicative manner to implement machine learning of each motor control device and each machine tool.

[0226] In this case, a distributed processing system can be used to appropriately distribute the functions of a single machine learning device across multiple servers. Alternatively, the functions of a single machine learning device can be implemented using virtual server functions on the cloud.

[0227] Furthermore, if there are n machine learning devices 300-1 to 300-n corresponding to the servo control devices 10-1 to 10-n, each with the same model name, specifications, or series, the learning results of each machine learning device 300-1 to 300-n can be shared, thereby enabling the construction of a more ideal model.

[0228] The machine learning device, control system, and machine learning method involved in the present disclosure can be implemented in various ways including the above-mentioned implementations and having the following structures.

[0229] (1) One embodiment of the present disclosure provides a machine learning device (e.g., machine learning unit 300, machine learning devices 300-1 to 300-n) that performs machine learning on a plurality of servo control units (e.g., servo control units 100, 200) that control a plurality of motors driving a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis.

[0230] A first servo control unit (e.g., servo control unit 100) associated with the axis subjected to interference among the plurality of servo control units includes a correction unit (e.g., a position deviation correction unit 111, a speed command correction unit 112, a torque command correction unit 113) that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit (e.g., servo control unit 200) associated with the axis subject to interference.

[0231] The machine learning device has:

[0232] a state information acquisition unit (for example, the state information acquisition unit 301 ) that acquires state information including the first servo control information of the first servo control unit, the second servo control information of the second servo control unit, and coefficients of the function;

[0233] a behavior information output unit (e.g., behavior information output unit 303) configured to output behavior information including adjustment information of the coefficient included in the state information to the correction unit;

[0234] a reward output unit (for example, reward output unit 3021 ) that outputs a reward value in reinforcement learning using an evaluation function that is a function of the first servo control information; and

[0235] A value function updating unit (eg, value function updating unit 3022 ) updates the value function based on the reward value output by the reward output unit, the state information, and the behavior information.

[0236] According to this machine learning device, the correction unit coefficient of the servo control unit that corrects inter-axis interference can be optimized, complex adjustments in the servo control unit can be avoided, and the instruction follow-up performance of the servo control unit can be improved.

[0237] (2) In the machine learning device described in (1) above, the first servo control information includes a position instruction and position feedback information of the first servo control unit, or a first position deviation of the first servo control unit.

[0238] The evaluation function outputs the reward value based on the following values, including: the second position deviation or the first position deviation calculated from the position instruction and position feedback information of the first servo control unit, the absolute value of the first or second position deviation, or the square of the absolute value.

[0239] (3) In the machine learning device described in (1) or (2) above, the variable related to the position instruction of the second servo control unit is at least one of the position instruction of the second servo control unit, the first-order differential of the position instruction, and the second-order differential of the position instruction.

[0240] The variable related to the position feedback information of the second servo control unit is at least one of the position feedback information of the second servo control unit, a first-order differential of the position feedback information, and a second-order differential of the position feedback information.

[0241] (4) In the machine learning device described in any one of (1) to (3) above, the machining program for learning the control of the first servo control unit and the second servo control unit is such that, during machine learning, the axis that is subjected to the interference is moved and the axis that is subjected to the interference is stationary.

[0242] (5) In the machine learning device described in any one of (1) to (4) above, the machine learning device has: an optimization behavior information output unit that outputs adjustment information of the coefficient of the correction unit based on the value function updated by the value function updating unit.

[0243] (6) Another aspect of the present disclosure provides a servo control device (e.g., servo control device 10), comprising:

[0244] The machine learning device described in any one of (1) to (5) above (e.g., the machine learning unit 300); and

[0245] a plurality of servo controls (e.g., servo controls 100, 200) that control a plurality of motors that drive a machine having a plurality of axes where one axis is disturbed by motion of at least one other axis,

[0246] A first servo control unit (e.g., servo control unit 100) associated with the axis subjected to interference among the plurality of servo control units includes a correction unit (e.g., a position deviation correction unit 111, a speed command correction unit 112, a torque command correction unit 113) that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit (e.g., servo control unit 200) associated with the axis subject to interference.

[0247] The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

[0248] According to this servo control device, complicated adjustments in the servo control unit can be avoided, and inter-axis interference can be corrected to improve command followability.

[0249] (7) Another aspect of the present disclosure provides a servo control system (e.g., servo control system 30) comprising:

[0250] The machine learning device according to any one of (1) to (5) above; and

[0251] A servo control device (e.g., servo control devices 10-1 to 10-n) includes a plurality of servo control units (e.g., servo control units 100 and 200) for controlling a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis.

[0252] A first servo control unit (e.g., servo control unit 100) associated with the axis subjected to interference among the plurality of servo control units includes a correction unit (e.g., a position deviation correction unit 111, a speed command correction unit 112, a torque command correction unit 113) that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit (e.g., servo control unit 200) associated with the axis subject to interference.

[0253] The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

[0254] According to this servo control system, complicated adjustments in the servo control unit can be avoided, inter-axis interference can be corrected, and instruction followability can be improved.

[0255] (8) Another aspect of the present disclosure provides a machine learning method for a machine learning device (e.g., machine learning unit 300, machine learning devices 300-1 to 300-n), wherein the machine learning device performs machine learning on a plurality of servo control units (e.g., servo control units 100, 200) that control a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis,

[0256] A first servo control unit (e.g., servo control unit 100) associated with the axis subjected to interference among the plurality of servo control units includes a correction unit (e.g., a position deviation correction unit 111, a speed command correction unit 112, a torque command correction unit 113) that calculates a correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on a function, the function including at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit (e.g., servo control unit 200) associated with the axis subject to interference.

[0257] In the machine learning method,

[0258] acquiring state information including first servo control information of the first servo control unit, second servo control information of the second servo control unit, and coefficients of the function;

[0259] outputting behavior information including adjustment information of the coefficient included in the state information to the correction unit,

[0260] outputting a reward value in reinforcement learning using an evaluation function, the evaluation function being a function of the first servo control information;

[0261] The value function is updated according to the reward value, the state information, and the behavior information.

[0262] According to this machine learning method, the correction unit coefficient of the servo control unit that corrects inter-axis interference can be optimized, complex adjustments in the servo control unit can be avoided, and the instruction follow-up performance of the servo control unit can be improved.

[0263] (9) In the machine learning method described in (8) above, adjustment information of the coefficient of the correction unit is output as optimization behavior information based on the updated value function.

Claims

1. A machine learning device for performing machine learning on a plurality of servo control units that control a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis, wherein: The first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit including a position deviation correction unit, a speed command correction unit, and a torque command correction unit, wherein the position deviation correction unit, the speed command correction unit, and the torque command correction unit respectively calculate correction values for correcting the position deviation, the speed command, and the torque command of the first servo control unit based on one or more functions, wherein the function includes at least one variable selected from among a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference. The machine learning device has: a state information acquisition unit configured to acquire state information including first servo control information of the first servo control unit, second servo control information of the second servo control unit, and coefficients of the position deviation correction unit, the speed command correction unit, and the torque command correction unit based on the function; a behavior information output unit configured to output behavior information including adjustment information to the correction unit, the adjustment information being used to adjust the coefficients of the position deviation correction unit, the speed command correction unit, and the torque command correction unit included in the state information; a reward output unit that outputs a reward value for each execution of a behavior based on adjustment of the coefficient in reinforcement learning using an evaluation function that is a function of the first servo control information; as well as A value function updating unit updates a value function related to the coefficient adjustment information based on the reward value output by the reward output unit, the state information, and the behavior information.

2. The machine learning device according to claim 1, wherein The first servo control information includes a position instruction and position feedback information of the first servo control unit, or a first position deviation of the first servo control unit. The evaluation function outputs the reward value based on the following values, including: the second position deviation or the first position deviation calculated from the position instruction and position feedback information of the first servo control unit, the absolute value of the first position deviation or the second position deviation, or the square of the absolute value.

3. The machine learning device according to claim 1 or 2, wherein: The variable related to the position command of the second servo control unit is at least one of the position command of the second servo control unit, the first-order differential of the position command, and the second-order differential of the position command. The variable related to the position feedback information of the second servo control unit is at least one of the position feedback information of the second servo control unit, a first-order differential of the position feedback information, and a second-order differential of the position feedback information.

4. The machine learning device according to claim 1 or 2, wherein: The machining program for learning the control of the first servo control unit and the second servo control unit moves the axis that applies the interference and stops the axis that receives the interference during machine learning.

5. The machine learning device according to claim 1 or 2, wherein: The machine learning device includes an optimization behavior information output unit that outputs adjustment information for adjusting the coefficient included in the state information to the correction unit based on the value function updated by the value function update unit.

6. A servo control device, characterized in that: Include: The machine learning device according to any one of claims 1 to 5; and a plurality of servo control units that control a plurality of motors that drive a machine having a plurality of axes, wherein one of the axes is disturbed by the motion of at least one other axis, A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a corresponding correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on one or more functions, wherein the function includes at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference; The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

7. A servo control system, characterized in that: Include: The machine learning device according to any one of claims 1 to 5; and A servo control device comprising a plurality of servo control units for controlling a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, one of the plurality of axes being disturbed by the motion of at least one other axis, A first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit that calculates a corresponding correction value for correcting at least one of a position deviation, a speed command, and a torque command of the first servo control unit based on one or more functions, wherein the function includes at least one variable of a variable associated with a position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference; The machine learning device outputs behavior information including adjustment information of the coefficient to the correction unit.

8. A machine learning method for a machine learning device, wherein the machine learning device performs machine learning on a plurality of servo control units that control a plurality of motors, wherein the plurality of motors drive a machine having a plurality of axes, wherein one of the plurality of axes is disturbed by the motion of at least one other axis, wherein: The first servo control unit associated with the axis subjected to interference among the plurality of servo control units includes a correction unit including a position deviation correction unit, a speed command correction unit, and a torque command correction unit, wherein the position deviation correction unit, the speed command correction unit, and the torque command correction unit respectively calculate correction values for correcting the position deviation, the speed command, and the torque command of the first servo control unit based on one or more functions, wherein the function includes at least one variable selected from among a variable associated with the position command and a variable associated with position feedback information of a second servo control unit associated with the axis subject to interference. In the machine learning method, acquiring state information including first servo control information of the first servo control unit, second servo control information of the second servo control unit, and coefficients of the position deviation correction unit, the speed command correction unit, and the torque command correction unit based on the function; outputting behavior information including adjustment information to the correction unit, the adjustment information being used to adjust the coefficients of the position deviation correction unit, the speed instruction correction unit, and the torque instruction correction unit included in the state information, outputting a reward value for each execution of the behavior based on the adjustment of the coefficient in reinforcement learning using an evaluation function that is a function of the first servo control information, An updated value function associated with the coefficient adjustment information is performed according to the reward value, the state information, and the behavior information.

9. The machine learning method according to claim 8, wherein: Based on the updated cost function, adjustment information for adjusting the coefficient included in the state information is output to the correction unit.

Citation Information

Patent Citations

  • Position driving control system and synchronizing / tuning position driving control method

    JP2001100819A

  • Evaluation program, evaluation method, and control system

    JP2019003404A

  • Correction device, correction device controlling method, information processing program, and recording medium

    US20170153614A1

  • Machine learning apparatus and method for optimizing smoothness of feed of feed axis of machine and motor control apparatus including machine learning apparatus

    US20170154283A1

  • Computer readable information recording medium, evaluation method, and control device

    US20180364678A1