Machine learning device, control device and method for machine learning
The machine learning device uses reinforcement learning to optimize filter coefficients in servo control devices, addressing the challenge of oscillations caused by axis position or speed gain influences, and achieving stable filter settings.
Patent Information
- Application Number
- DE102020203758
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-15
- Filing Date
- 2020-03-24
- Publication Date
- 2025-05-28
- Estimated Expiration
- 2040-03-24
AI Technical Summary
In servo control devices, optimizing the filter characteristics to attenuate specific frequency components is challenging due to the influence of other axes' positions or speed gains, leading to potential oscillations even after optimizing for a specific position or speed gain.
A machine learning device performs reinforcement learning to optimize the coefficients of filters in servo control devices, using state information from a frequency characteristic calculation device to adjust filter settings and determine rewards based on evaluation values, thereby updating the action value function.
This approach allows for optimal filter setting even when machine characteristics change due to axis positions or interactions with other axes, effectively reducing oscillations and improving system stability.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTIONField of the invention
[0001] The present invention relates to a machine learning apparatus that performs reinforcement learning for optimizing a coefficient of at least one filter that attenuates at least one specific frequency component provided in a servo control apparatus for controlling a motor, to a control apparatus including such a machine learning apparatus, and to a machine learning method. State of the art
[0002] Devices that adjust the characteristics of a filter are described, for example, in Patent Documents 1 to 4. Patent Document 1 describes a vibration reduction device of a robot, comprising: a robot axis drive source provided in each axis of the robot and driving the robot axis according to an input control signal; and signal processing means that removes a frequency component corresponding to the natural frequency of the robot axis from the control signal and supplies the control signal, which has been subjected to signal processing in the signal processing means to reduce vibrations generated in the robot axis, to the robot axis drive source.In the robot vibration reduction device, a neural network is provided that inputs the current position of each axis of the robot to drive and output the natural frequency of each axis of the robot, and applies it to the signal processing device. The frequency component corresponding to the natural frequency of the robot axis output from the neural network is removed from the control signal. Patent Document 1 also describes that the signal processing means is a notch filter, and that a notch frequency is varied depending on the natural frequency of the robot axis output from the neutral network.
[0003] Patent Document 2 describes an XY stage control device in which movable guides that cross vertically and laterally are arranged on a stage and movable slides are arranged at their intersections. The XY stage control device includes: a variable notch filter that can variably adjust a notch frequency for absorbing the amplification of the resonant motion of the movable slides; and a switching device that inputs the position information of the movable slides on the stage and outputs a switching signal for switching the notch frequency of the notch filter.
[0004] Patent Document 3 describes a servo control device according to an embodiment, comprising: a command sampling that controls a servo amplifier to drive a moving element that performs a rotational motion or a reciprocal motion based on a torque command or a current command, and that samples the torque command or the current command for the servo amplifier when setting a speed control gain;and an operation processing unit that, when adjusting the speed control gain, converts the sample value of the torque command or the current command into the magnitude of the torque of the moving element at a frequency, and performs oscillation band determination to determine that a frequency band in which the magnitude of the torque of the moving element has a peak value is an oscillation band, and filter adjustment for setting a band-stop filter to attenuate the magnitude of the torque of the moving element in the oscillation band when adjusting the speed control gain;
[0005] Patent Document 4 describes a servo control device comprising: a speed command calculation unit; a torque command calculation unit; a speed detection unit; a speed control loop; a speed control loop gain adjustment unit; at least one filter that removes a specific band of torque command values; a sinusoidal disturbance input unit that performs sinusoidal sampling of the speed control loop; a frequency characteristic calculation unit that estimates the gain and phase of an input / output signal of the speed control loop; a resonance frequency detection unit; a filter adjustment unit that adjusts a filter according to a resonance frequency; a gain adjustment unit;a sequence control unit that automatically performs online detection of the resonance frequency, adjustment of a speed control loop gain, and filter adjustment; and a setting status display unit. The setting status display unit displays the setting stage and progress of the sequence control unit. Patent Document 1: Unexamined Japanese Patent Application, Publication No. JP H07-261853 A. Patent Document 2: Unexamined Japanese Patent Application, Publication No. JP S62-126402 A. Patent Document 3: Unexamined Japanese Patent Application, Publication No. JP 2013 - 126 266 A. Patent Document 3: Unexamined Japanese Patent Application, Publication No. JP 2017 - 022 855 A.
[0006] DE 10 2018 003 769 A1 discloses a machine learning device, servo control system, and machine learning method. A control gain is appropriately adjusted according to a phase of a motor. A machine learning device that performs reinforcement learning with respect to a servo control device that controls an operation of a control target device having a motor comprises: an action information output device for outputting action information including information for adjusting coefficients of a transfer function of a control gain to a control unit incorporated in the servo control device;A state information retrieval device for retrieving, from the servo control device, state information including a deviation between an actual operation of the control target device and a command input to the control unit, a phase of the motor, and the coefficients of the control gain transfer function when the control unit operates the control target device based on the action information; a reward output device for outputting a value of a reinforcement learning reward based on the deviation contained in the state information, and a value function update device for updating an action value function based on the reward value, the state information, and the action information. SUMMARY OF THE INVENTION
[0007] In a case where, when determining the characteristics of a filter such as a notch filter in a servo controller of one axis, a machine characteristic is affected by the position of another axis or the speed gain of a servo controller of the other axis, even if the filter characteristics are optimized with a specific position of the other axis or a specific speed gain, oscillation may occur in the other position or speed gain. Even in a case where the machine characteristic is not affected by the position of the other axis, oscillation may occur depending on the position of the current axis.Therefore, it is desirable to make the optimal setting of a filter characteristic even if a machine characteristic is changed by the position of the current axis or is influenced by another axis.
[0008] The object is achieved by a machine learning device having the features of patent claim 1 and by a control device having the features of patent claim 8. The object is further achieved by a machine learning method having the features of patent claim 9. (1) One aspect of the present invention is a machine learning device (400) that performs reinforcement learning in which a servo control device (100) for controlling a motor (150) is driven under a variety of conditions and that optimizes a coefficient of at least one filter (130) for attenuating at least one specific frequency component provided in the servo control device (100), and wherein the machine learning device (400) comprises: a state information acquiring unit (401) that acquires state information including the result of a calculation of a frequency characteristic calculating device (300) for calculating an input / output gain of the servo control device (100) and / or a phase delay of an input and an output, the coefficients of the filter (130), and the conditions;an action information output unit (403) that outputs action information including setting information of the coefficient included in the state information to the filter (130); a reward output unit (4021) that individually determines evaluation values under the conditions based on the result of the calculation, so as to output the value of a sum of the evaluation values as a reward; and a value function update unit (4022) that updates an action value function based on the value of the reward output by the reward output unit (4021), the state information, and the action information. (2) Another aspect of the present invention is a control device (10) comprising: the machine learning device (400) of (1) described above; the servo control device (100) comprising at least one filter (130) for attenuating at least one specific frequency component and controlling the motor (150); and the frequency characteristic calculation device (300) calculating the input / output gains of the servo control device (100) and / or the phase delay of the input and the output in the servo control device (100). (3) Another aspect of the present invention is a machine learning method of a machine learning device (400) that performs reinforcement learning in which a servo control device (100) for controlling a motor (150) is driven under a plurality of conditions and that optimizes a coefficient of at least one filter (130) for attenuating at least one specific frequency component provided in the servo control device (100).The machine learning method comprises: acquiring state information including the result of a calculation for calculating an input / output gain of the servo control device (100) and / or a phase delay of an input and an output, the coefficients of the filter (130), and the conditions; outputting action information including setting information of the coefficient included in the state information to the filter (130); individually determining evaluation values under the conditions based on the result of the calculation so as to determine the value of a sum of the evaluation values as a reward; and updating an action value function based on the value of the determined reward, the state information, and the action information.
[0009] In each of the aspects of the present invention, even when the machine characteristic of a machine tool, a robot, an industrial machine, or the like is changed depending on the conditions, for example, even when the machine characteristic is changed depending on the position of an axis or the machine characteristic is influenced by another axis, it is possible to perform the optimal setting of a filter characteristic. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram showing a control device including a machine learning device according to an embodiment of the present invention; Fig. 2 is a diagram showing a speed command serving as an input signal and a speed detection value serving as an output signal; Fig. 3 is a diagram showing an amplitude relationship between an input signal and an output signal and the frequency characteristic of a phase delay. Fig. 4 is a perspective view of a five-axis machine tool showing an example of the control target of the control device; Fig. 5 is a characteristic diagram showing an example of the frequency characteristic of an X-axis at the left end of the X-axis; Fig. 6 is a characteristic diagram showing an example of the frequency characteristic of the X-axis at the center of the X-axis; Fig. 7 is a characteristic diagram showing an example of the frequency characteristic of the X-axis at the right end of the X-axis; Fig. 8 is a schematic diagram showing how the servo strength of one axis changes the frequency characteristics of the input / output gain of the other axis; Fig. Figure 9 is a schematic diagram showing how the frequency characteristics of the input / output gain of the other axis are changed with the position of one axis; Fig. 10 is a block diagram showing a machine learning unit according to the embodiment of the present invention; Fig. 11 is a block diagram serving as a model for calculating the standard input / output gain model; Fig. 12 is a characteristic diagram showing the frequency characteristics of the input / output gain of a servo control unit in the standard model and a servo control unit before and after learning; Fig. 13 is a characteristic diagram showing a relationship between the bandwidth of a filter and a gain and a phase; Fig. 14 is a characteristic diagram showing a relationship between the attenuation coefficient of the filter and the gain and phase; Fig. 15 is a flowchart showing the operation of the machine learning unit at the time of Q-learning in the present embodiment; Fig. 16 is a flowchart illustrating the operation of an optimization action information output unit in the machine learning unit in the embodiment of the present invention; Fig. 17 is a block diagram showing an example in which a plurality of filters are directly connected to each other to form the filter; and Fig. 18 is a block diagram showing another configuration example of the control device. DETAILED DESCRIPTION OF THE INVENTION
[0010] An embodiment of the present invention will be described in detail below with reference to the drawings.
[0011] Fig. 1 is a block diagram showing a control device including a machine learning apparatus according to an embodiment of the present invention. Examples of the control target 500 of the control device 10 include a machine tool, a robot, and an industrial machine. The control device 10 can be provided as part of the control target such as a machine tool, a robot, or an industrial machine.
[0012] The control device 10 includes a servo control unit 100, a frequency generation unit 200, a frequency characteristic calculation unit 300, and a machine learning unit 400. The servo control unit 100 corresponds to a servo control device, the frequency characteristic calculation unit 300 corresponds to a frequency characteristic calculation device, and the machine learning unit 400 corresponds to a machine learning device. One or more of the frequency generation unit 200, the frequency characteristic calculation unit 300, and the machine learning unit 400 may be provided within the servo control unit 100. The frequency characteristic calculation unit 300 may be provided in the machine learning unit 400.
[0013] The servo control unit 100 includes a subtractor 110, a speed control unit 120, a filter 130, a current control unit 140, and a servo motor 150. The subtractor 110, the speed control unit 120, the filter 130, the current control unit 140, and the servo motor 150 form a speed feedback loop. A linear motor that performs a linear motion, a motor that includes a rotation axis, or the like can be used as the servo motor 150. The servo motor 150 can be provided as part of the control target 500.
[0014] The subtractor 110 determines a difference between a speed command value that is input and a feedback speed detection value to output the difference as a speed error to the speed control unit 120.
[0015] The speed control unit 120 adds a value obtained by multiplying the speed error by an integral gain K1v and integrating the result and a value obtained by multiplying the speed error by a proportional gain K2v to output the resultant value to the filter 130 as a torque command.
[0016] The filter 130 is a filter that attenuates a specific frequency component, and a notch filter or a low-pass filter, for example, is used. In a machine, e.g., a machine tool driven by a motor, a resonance point exists, and thus the resonance of the servo control unit 100 can be increased. A filter such as a notch filter is used, and thus it is possible to reduce the resonance. The output of the filter 130 is output as a torque command to the current control unit 140. Mathematical formula 1 (indicated below as Math. 1) specifies a transfer function F(s) of a notch filter serving as the filter 130. Parameters specify the coefficients ω c , τ and k. In the mathematical formula 1, the coefficient k is a damping coefficient, the coefficient ω can angular center frequency and the coefficient τ a subbandwidth. If the center frequency is assumed to be fc and a bandwidth is fw, the coefficient ω c by ω c = 2 π fc and the coefficient τ is represented by τ = fw / fc. F(s)=s2+2kτωcs+ωc2s2+2τωcs+ωc2
[0017] The current control unit 140 generates a current command for driving the servo motor 150 based on the torque command and outputs the current command to the servo motor 150. When the servo motor 150 is a linear motor, the position of a movable section is detected using a linear scale (not shown) provided in the servo motor 150, a position detection value is differentiated to determine a speed detection value, and the obtained speed detection value is input to the subtractor 110 as speed feedback. When the servo motor 150 is a motor with a rotary axis, a rotation angle position is detected using a rotary encoder (not shown) provided in the servo motor 150, and a speed detection value is input to the subtractor 110 as speed feedback.Although the servo control unit 100 is configured as described above, the control device 10 additionally includes the frequency generation unit 200, the frequency characteristic calculation unit 300, and the machine learning unit 400 to perform machine learning with optimal parameters for the filter.
[0018] The frequency generation unit 200 uses a sinusoidal signal as a speed command when changing the frequency to output it to the subtractor 110 of the servo control unit 100 and the frequency characteristic calculation unit 300.
[0019] The frequency characteristic calculation unit 300 uses the speed command (sine wave) generated in the frequency generation unit 200 and serving as an input signal, and the speed detection value (sine wave) output from the rotary encoder (not shown) and serving as an output signal or integration (sine wave) of a detection position serving as an output signal of the linear scale, and thereby determines an amplitude ratio (input / output gain) between the input signal and the output signal and a phase delay at each frequency specified by the speed command. Fig. 2 is a diagram showing the speed command as an input signal and the speed detection value as an output signal. Fig. Figure 3 is a graph showing the amplitude relationship between the input signal and the output signal and the frequency characteristic of the phase delay. According to the diagram in Fig. 2, the speed command output from the frequency generation unit 200 is varied in frequency and thus the input / output gain (amplitude ratio) and the frequency characteristic of the phase delay are obtained as shown in Fig. 3 is shown.
[0020] The machine learning unit 400 uses the input / output gain (amplitude ratio) output from the frequency characteristic calculation unit 300 and the phase delay to perform machine learning (hereinafter referred to as learning) with the coefficients ω c, τ, and k of the transfer function for the filter 130. Although learning is performed using the machine learning unit 400 before delivery, learning may be performed again after delivery. For example, the machine tool serving as the control target 500 is a five-axis machine tool having three linear axes consisting of an X-axis, a Y-axis, and a Z-axis, and two rotary axes consisting of a B-axis and a C-axis. Fig. 4 is a perspective view of a five-axis machine tool showing an example of the control target of the control device 10. Fig. 4 shows an example in which the servo motor 150 is provided in the machine tool serving as the control target 500. The servo motor 150 serving as the control target 500 and in Fig. The machine tool shown in Figure 4 includes linear motors 510, 520, and 530 that linearly move the tables 511, 521, and 531 in the X-axis direction, the Y-axis direction, and the Z-axis direction, respectively. The Y-axis linear motor 530 is mounted on the Z-axis linear motor 520. The machine tool also includes built-in motors 540 and 550 that rotate the tables 541 and 551 in the C-axis direction and the B-axis direction, respectively. For linear motors 510, 520, and 530, tables 511, 521, and 531 are movable sections. For built-in motors 540 and 550, tables 541 and 551 are movable sections. Therefore, linear motors 510, 520, and 530, and built-in motors 540 and 550, directly drive tables 511, 521, and 531, and tables 541 and 551, without the use of a gearbox or similar device. Linear motors 510, 520, and 530, and built-in motors 540 and 550, each correspond to servo motor 150.The tables 511, 521, and 531 can connect the rotation axis of the motor to a ball screw via a coupling, allowing them to be driven by a nut screwed onto the ball screw. The configuration and detailed operation of the machine learning unit 400 are described in more detail below. In the following description, the control target 500 is the position shown in . Fig. The machine tool shown in Figure 4 is used as an example. <Maschinenlerneinheit 400>
[0021] Although the following discussion describes a case where the machine learning unit 400 performs reinforcement learning, the learning performed by the machine learning unit 400 is not specifically limited to reinforcement learning, and the present invention can also be applied to a case where supervised learning is performed, for example.
[0022] Before describing the individual functional blocks provided in the machine learning unit 400, the basic mechanism of reinforcement learning will first be described. An agent (which in the present embodiment corresponds to the machine learning unit 400) observes the state of an environment and selects a specific action. Thus, the environment is changed due to the action described above. Depending on the change in the environment, a specific reward is given, and thus the agent learns to select (decision-making) a better action. While supervised learning provides a completely correct answer, the reward in reinforcement learning is often a fragmented value based on the change in part of the environment. Therefore, the agent learns to select an action in such a way that the overall reward for the future is maximized.
[0023] As described above, reinforcement learning involves learning an action and, therefore, a method for learning an appropriate action based on the interaction of the action with the environment, particularly learning to maximize the reward that can be achieved in the future. In the present embodiment, for example, this suggests that it is possible to acquire such an action that affects the future, i.e., an action of selecting action information to reduce vibrations at a machine end.
[0024] Although any learning method can be used as reinforcement learning here, the following discussion gives an example using Q-learning (which is a method for learning a value Q(S,A)) selected by an action A under a given environmental state S. A goal of Q-learning is to select an action A that has the highest value Q(S,A) as the optimal action among the actions A that can be performed in a given environmental state S.
[0025] However, at the time Q-learning is first started, the correct value of Q(S,A) for a combination of state S and action A is not found at all. Therefore, the agent learns the correct value Q(S,A) by choosing different actions A under a given state S and selecting a better action based on a reward given for action A at that time.
[0026] Since the goal is to maximize the total rewards achieved in the future, the ultimate goal is Q(S,A) = E[Σ(γ t )r t ] where E[] represents an expected value, t represents the time, γ represents a parameter called the discount rate, which will be described later, r t represents a reward at time t, and Σ represents the sum at time t. However, since it is unclear which action is the optimal action in the Q-learning process, reinforcement learning is performed while performing a search by performing various actions. A formula for updating the value Q(S,A) as described above can be represented, for example, by the mathematical formula 2 below (referred to as Math. 2 below). Q(St+1,At+1)←Q(St,At)+α(rt+1+γmaxQA(St+1,A)−Q(St,At))
[0027] In the mathematical formula 2 described above, S tfor the state of an environment at time t and A t for an action at time t. By the action A t the state is set to S t+1 changed. Here r t+1 represents a reward achieved by changing the state. A term with max is obtained by multiplying a Q value by γ when an action A, which has the highest Q value found at that time, is selected under state St+1. Here, γ is a parameter from 0<γ≤1 and is called the discount rate. Furthermore, α is a learning coefficient and is assumed to be in a range of 0<α≤1.
[0028] The mathematical formula 2 described above represents a method for updating a value Q(St,At) of an action A t in a state S t based on a reward r t+1 which, as a result of action A tThis update formula shows that if the value max a Q(S t+1 ,A) the best action in the subsequent state S t+1 , caused by action A t is higher than the value Q(St,At) of action A t in state S t , Q(St,At) is increased, while if the value max a Q(S t+1 ,A) is lower than the value Q(St,At), Q(St,At) is reduced. In other words, the value of a particular action in a particular state is adjusted to approach the value of the best action in the subsequent state caused by it. However, although this difference depends on the discount rate γ and the reward r t+1 is changed, the value of the best action in a given state is generally propagated to the value of an action in a state prior to the best action.
[0029] Here, in Q-learning, there is a method of constructing a table of Q(S,A) for all state-action pairs (S,A) to perform learning. However, since the number of states is too large to determine the values of Q(S,A) in all state-action pairs, it is likely that Q-learning will take a long time to converge.
[0030] Therefore, a well-known technology called DQN (Deep Q-Network) can be used. Specifically, the value of Q(S,A) can be calculated by forming a value function Q with a suitable neural network, adjusting the parameters of the neural network, and thereby approximating the value function Q with the appropriate neural network. By using DQN, it is possible to shorten the time required for Q-learning to converge. The details of DQN are described, for example, in a non-patent document below. <nichtpatentdokument>
[0031] "Human-level control through deep reinforcement learning", written by Volodymyr Mnihl, [online], [accessed on January 17, 2017], Internet<URL: http: / / files.davidqiu.com / research / nature14236.pdf>
[0032] The machine learning unit 400 performs the Q-learning described above. Specifically, the machine learning unit 400 learns the value Q in which the values of the coefficients ω c , τ and k of the transfer function for the filter 130, the input / output gain (amplitude ratio) output from the frequency characteristic calculation unit 300, and the phase delay are set to the state S.
[0033] The machine learning unit 400 controls the servo control unit 100 based on the coefficients ω c , τ and k of the transfer function for the filter 130 with the speed command described above, i.e., the sine wave whose frequency is changed, while observing the state information S obtained from the frequency characteristic calculation unit 300, which includes the input / output gain (amplitude ratio) and the phase delay at each frequency, to determine the action A. The machine learning unit 400 receives a reward each time the action A is executed. For example, the machine learning unit 400 searches for the optimal action A in a trial-and-error manner so that the total rewards for the future are maximized. In this way, the machine learning unit 400 controls, based on the coefficients ω c , τ and k of the transfer function for the filter 130, the servo control unit 100 with the speed command, which is the sine wave whose frequency is changed, and can thereby determine the optimal action A (ie, the optimal coefficients ω c , τ and k of the transfer function for the filter 130) for the state S obtained from the frequency characteristic calculation unit 300 and including the input / output gain (amplitude ratio) and the phase delay at each frequency.
[0034] In other words, based on the value function Q learned by the machine learning unit 400, among the actions A acting on the coefficients ω c , τ and k of the transfer function for the filter 130, which refer to a particular state S, such an action A is selected that the value of Q is maximized, and thus it is possible to determine such an action A (i.e. the coefficients ω c , τ and k of the transfer function for the filter 130) to minimize vibrations at the machine end caused by the execution of a program for generating a sinusoidal signal whose frequency is varied.
[0035] The state S includes the values of the coefficients ω c , τ, and k of the transfer function for the filter 130, the input / output gain (amplitude ratio), and the phase delay output from the frequency characteristic calculation unit 300 by driving the servo control unit under each of a plurality of conditions, and the plurality of conditions. The machine learning unit 400 determines an evaluation value under each of the conditions based on the input / output gain (amplitude ratio) and the phase delay included in the state S under each of the conditions, and adds the evaluation values under each condition to determine the reward. The details of a method for determining the reward will be described later. The action A is the modified information of the coefficients ω c , τ and k of the transfer function for the filter 130.
[0036] As a variety of conditions, three examples can be mentioned below.
[0037] (a) A plurality of positions of an axis (e.g., the X-axis) controlled by the servo control unit 100
[0038] The positions are a plurality of positions changed by the servo control unit 100, for example, a plurality of positions of the axis specified with a predetermined pitch such as 200 mm. The positions can be a plurality of positions determined as the left end, the center, and the right end of the axis. The positions can be four or more points. In the case of the machine tool, for example, the position of the axis corresponds to the position of the table. If the servo motor 150 is a linear motor, the position of the X-axis controlled by the servo control unit 100 is determined by the detection position of the moving part (table) of the linear motor, which is detected with the linear scale. The detection position of the moving part is input from the linear scale to the machine learning unit 400. If the servo motor 150 is a motor that uses, for example,has a rotation axis, the rotation axis of the motor is connected to a ball screw via a coupling, so that a nut screwed onto the ball screw drives the table. Therefore, the position of the axis controlled by the servo control unit 100 is determined by detecting the movement of the table with the linear scale attached to the table and using the detection position of the linear scale. The detection position (the position of the axis) of the table is input to the machine learning unit 400 as state S. Fig. Figure 1 shows how the detection position (the position of the axis) of the table, which serves as part of the control target 500 and is attached to the table, is input to the machine learning unit 400. The state S includes the values of the coefficients ω c , τ and k of the transfer function for the filter 130, the input / output gain (amplitude ratio) and the phase delay for each condition output from the frequency characteristic calculation unit 300 by driving the servo control unit under each of a plurality of conditions (a plurality of X-axis positions), and the detection position (the position of the axis) of the table corresponding to each of the conditions.
[0039] The Fig. Figures 5 to 7 are characteristic diagrams showing an example of the frequency characteristics (frequency characteristics of input / output gain and phase delay) of the X-axis at the left end, middle, and right end of the X-axis. As shown in a range of the frequency characteristics of the input / output gain in Fig. 5 and Fig. 7, which is surrounded by a dashed line, the resonance is increased at the left end and the right end of the X-axis, and as in a range of the frequency characteristic of the input / output gain in Fig. 6, which is surrounded by a dashed line, the resonance is reduced at the center of the X-axis. The machine learning unit 400 determines the evaluation value under each of the conditions based on the input / output gain (amplitude ratio) and phase delay at a plurality of positions (e.g., left end, center, and right end of the X-axis) of the X-axis included in state S and corresponding to each of the conditions, and sums the evaluation values to determine the reward.
[0040] (b) A plurality of speed gains of the servo control unit that controls an axis (e.g., the Z-axis) different from an axis (e.g., the Y-axis) controlled by the servo control unit 100
[0041] Fig. Figure 8 is a schematic characteristic diagram showing how the servo strength of one axis changes the frequency characteristics of the input / output gain of the other axis. Here, the servo strength indicates the strength of noise and Fig. Figure 8 shows that when the servo strength decreases, the change in the frequency characteristic of the input / output gain of the other axis increases with the servo strength of one axis. When the speed gain of the servo control unit controlling the Z-axis is decreased, the servo strength of the Y-axis is decreased, while when the speed gain of the servo control unit controlling the Z-axis is increased, the servo strength of the Y-axis is increased. Therefore, a variety of speed gains are calculated taking into account the Fig. 8 is set. Although the frequency characteristic of the Y-axis is described here, when the speed gains of the servo control unit that controls the Z-axis are different, three or more speed gains of the servo control unit that controls the Z-axis can be set.
[0042] The velocity gains of the servo control unit controlling the Z-axis are input as state S to the machine learning unit 400, which optimizes the coefficients for the filter of the Y-axis servo control unit 100. State S includes the values of the coefficients ω c , τ, and k of the transfer function for the filter 130, the input / output gain (amplitude ratio), and the phase lag output from the frequency characteristic calculation unit 300 by driving the servo control unit under each of a plurality of conditions (a plurality of speed gains), which are based on each of the conditions and the speed gain of the servo control unit that corresponds to each of the conditions and controls the Z-axis. The machine learning unit 400 determines an evaluation value under each of the conditions based on the input / output gain (amplitude ratio) and the phase lag of the Y-axis in the speed gain of the servo control unit that controls the Z-axis, which are included in the state S and correspond to each of the conditions, and sums the evaluation values to determine the reward.
[0043] (c) A plurality of positions of an axis (e.g., the Y-axis) that are different from an axis controlled by the servo control unit 100 (e.g., the Z-axis)
[0044] The frequency characteristic of an axis controlled by the servo control unit 100 can be changed by the position of another axis. As an example, there is a case where, as shown in Fig. 4, the Y-axis is arranged on the Z-axis, and the frequency characteristic of the Z-axis is changed depending on a plurality of positions of the Y-axis. The positions are changed by the servo control unit (not shown) of the Y-axis and represent, for example, a plurality of positions on the axis specified at a predetermined pitch, such as 200 mm. The positions may be a plurality of positions, such as the upper end and the lower end of the Y-axis, which are specified. The positions may be three or more points. When the servo motor 150 is a linear motor, the position of the Y-axis controlled by the servo control unit is determined by the detection position of the movable part of the linear motor, which is detected with the linear scale.The detection position of the movable section is input from the linear scale to the machine learning unit 400, which optimizes the coefficients for the filter of the Z-axis servo control unit 100.
[0045] If the servo motor 150 is a motor with a rotation axis, for example, the rotation axis of the motor is connected to a ball screw via a coupling, so that a nut screwed onto the ball screw drives the control target stage. Therefore, the position of the Y-axis controlled by the servo control unit is determined by detecting the movement of the stage with the linear scale attached to the stage and using the detection position of the linear scale. The detection position of the stage is input as state S to the machine learning unit 400, which optimizes the coefficients for the Z-axis servo control unit's filter. Fig. Figure 9 is a schematic characteristic diagram showing how the position of one axis changes the frequency characteristic of the input / output gain of the other axis. Fig. 9 shows how the position (the axis position A and the axis position B of Fig. 9) of one axis, the position and magnitude of an increase in the input / output gain of the other axis is changed.
[0046] The state S includes the values of the coefficients ω c , τ, and k of the transfer function for the filter 130, the input / output gain (amplitude ratio), and the phase delay for each condition output from the frequency characteristic calculation unit 300 by driving the servo control unit under each of multiple conditions (multiple Y-axis positions), and the detection position (axis position) of the Y-axis table corresponding to each of the conditions. The machine learning unit 400 determines, based on the input / output gain (amplitude ratio) and the Z-axis phase delay at the Y-axis positions (e.g., the upper end and lower end of the Y-axis) included in the state S, an evaluation value corresponding to each of the conditions at each Y-axis position, and sets the sum of the evaluation values as the reward.The machine learning unit 400 determines, based on the input / output gain (amplitude ratio) and the phase delay under each of a plurality of conditions (e.g., upper end and lower end of the Y-axis) of the Z-axis frequency characteristic in the positions (e.g., upper end and lower end of the Y-axis) of the Y-axis included in the state S, the evaluation value under each of the conditions and sums the evaluation values to determine the reward.
[0047] Although (b) describes the case where the frequency characteristic of the Y-axis is changed by the speed gain of the servo control unit that controls the Z-axis, the frequency characteristic of the Z-axis below the Y-axis can be changed by the speed gain of the servo control unit that controls the Y-axis. Although (c) describes the case where the frequency characteristic of the Z-axis controlled by the servo control unit 100 is changed by the position of the Y-axis, the frequency characteristic of the Y-axis controlled by the servo control unit 100 can be changed by the position of the Y-axis.
[0048] By using the reward which is the sum of the individual evaluation values under a plurality of conditions in any one of (a) to (c) described above, the machine learning unit 400 performs the learning, and thus, even in a machine in which the frequency characteristic (the frequency characteristic of the input / output gain and the phase delay) is changed by a plurality of conditions, it is possible to obtain the optimal coefficients ω c , τ and k of the transfer function for the filter 130.
[0049] If the calculated input / output gain is less than or equal to the input / output gain of a standard model, the evaluation value is a positive value given when the phase delay is decreased, a negative value given when the phase delay is increased, or a zero value given when the phase delay is not changed. The standard model is a model of the servo control unit that has an ideal characteristic without any oscillation. The input / output gain of the standard model will be described later. Even if the reward is determined by the sum of the individual evaluation values under a variety of conditions, thus changing the frequency characteristics of the input / output gain or the phase delay under each of the conditions, it is possible to perform efficient learning in which the adjustment of the filter is performed stably.
[0050] A weight can be assigned to the evaluation value corresponding to each of a plurality of conditions. As described above, a weight is assigned to the evaluation value so that even if the influences of each condition on a machine property differ from each other, the reward corresponding to the influence can be determined. For example, in (a), as described above, the evaluation values determined at the left end, middle, and right end positions of the X-axis are assumed to be Es(L), Es(C), and Es(R), and the reward is assumed to be Re. The weighting coefficients of the evaluation values Es(L), Es(C), and Es(R) are assumed to be coefficients a, b, and c, and the reward Re is determined by Re = a × Es(L) + b × Es(C) + c × Es(R). The coefficients a, b, and c can be determined as needed, and e.g.In the case of a machine tool where oscillation is unlikely to occur in the center of the X-axis, coefficient b can be set lower than coefficients a and c.
[0051] If the reward is determined by the sum of the individual evaluation values corresponding to each condition, even if one evaluation value is a negative value, the other evaluation values can be positive values, so the reward is a positive value. Therefore, the reward can only be determined by the sum of the individual evaluation values corresponding to each condition when all evaluation values are zero or positive values. Then, if there is even one negative value among all evaluation values, the reward is set to a negative value. Preferably, this negative value is set to a large value (e.g., -∞), thus preventing a case from being selected where even one negative value exists among all evaluation values. In this way, it is possible to perform efficient learning in which the adjustment of the filter is performed stably in any position.
[0052] Fig. 10 is a block diagram showing the machine learning unit 400 according to the embodiment of the present invention. To implement the previously described reinforcement learning as shown in Fig. 10, the machine learning unit 400 includes a state information acquisition unit 401, a learning unit 402, an action information output unit 403, a value function storage unit 404, and an optimization action information output unit 405. The learning unit 402 includes a reward output unit 4021, a value function update unit 4022, and an action information generation unit 4023.
[0053] The state information acquisition unit 401 acquires from the frequency characteristic calculation unit 300 based on the coefficients ω c , τ, and k of the transfer function for the filter 130, the state S representing the input / output gain (amplitude ratio) and the phase delay under each of the conditions obtained by driving the servo control unit 100 with the speed command (sine wave). This state information S corresponds to an environmental state S in Q-learning. The state information acquisition unit 401 outputs the acquired state information S to the learning unit 402.
[0054] The coefficients ω c , τ and k of the transfer function for the filter 130 at the time of the first start of Q-learning are previously generated by a user. In the present embodiment, the user-generated initial setting values of the coefficients ω c , τ and k of the transfer function for the filter 130 are optimally adjusted by reinforcement learning. If an operator adjusts the machine tool beforehand, the adjusted values of the coefficients ω c , τ and k are set to the initial values and machine learning is performed.
[0055] The learning unit 402 is a unit that learns the value Q(S,A) under a certain environmental condition S when a certain action A is selected.
[0056] The reward output unit 4021 is a unit that calculates the reward under the certain state S when the action A is selected. If the coefficients ω c , τ, and k of the transfer function for the filter 130 are modified, the reward output unit 4021 compares an input / output gain Gs calculated under each of the conditions with an input / output gain Gb at each frequency of the preset standard model. If the calculated input / output gain Gs is greater than the input / output gain Gb of the standard model, the reward output unit 4021 provides a first negative evaluation value. On the other hand, if the calculated input / output gain Gs is less than or equal to the input / output gain Gb of the standard model, the reward output unit 4021 provides a positive evaluation value when the phase delay is decreased, provides a second negative evaluation value when the phase delay is increased, or provides a zero evaluation value when the phase delay is not changed.Preferably, the absolute value of the second negative value is set lower than the absolute value of the first negative value, thus preventing a case from being selected where the calculated input / output gain Gs is larger than the input / output gain Gb of the standard model.
[0057] An operation in which the negative evaluation value is provided with the reward output unit 4021 when the calculated input / output gain Gs is larger than the input / output gain Gb of the standard model will first be described with reference to Fig. 11 and Fig. 12. The reward output unit 4021 stores the standard model of the input / output gain. The standard model is a model of the servo control unit that has an ideal characteristic without any oscillation. The standard model can be calculated, for example, from the moment of inertia Ja, a torque constant Kt, a proportional gain K p , an integral gain K I and a derivative gain KD according to Fig. 11. The moment of inertia Ja is a value resulting from the addition of motor inertia and machine inertia. Fig. Figure 12 is a map showing the frequency characteristics of the input / output gains of the servo control unit in the standard version and the servo control unit 100 before and after learning. As can be seen from the map in Fig. 12, the Standard Model has a region A, which is a frequency range in which an ideal input / output gain equal to or greater than a constant input / output gain, e.g., equal to or greater than -20db, is provided, and a region B, which is a frequency range in which an input / output gain smaller than the constant input / output gain is provided. In the region A of Fig. 12, the ideal input / output gain in the standard model is given by a curve MC 1 (thick line). In area B of Fig. 12, the ideal virtual input / output gain in the standard model is given by a curve MC 11 (a thick dashed line), and the input / output gain in the standard model is set to a constant value and is represented by a straight line MC 12 (a thick line). In areas A and B of Fig. 12 are the input / output gain curves of the servo control unit before and after learning by the curves RC 1 and RC 2 marked.
[0058] In area A, when the curve RC 1 before learning the calculated input / output gain, the curve MC 1 the ideal input / output gain in the standard model, the reward output unit 4021 provides the first negative evaluation value. In the region B where the frequency is exceeded at which the input / output gain is sufficiently reduced, even if the curve RC 1 the input / output gain before learning the curve MC 11 the ideal virtual input / output gain in the Standard Model, the influence on stability is reduced. Therefore, in region B, as described above, the input / output gain in the Standard Model is not the curve MC11 of the ideal gain characteristic, but the straight line MC 12 of constant input / output gain (e.g. - 20dB). However, since instability may be caused if the curve RC1 of the calculated input / output gain before learning crosses the straight line MC 12 the input / output gain exceeds the constant value, the first negative value is specified as the evaluation value.
[0059] The following describes a process in which, when the calculated input / output gain Gs is less than or equal to the input / output gain Gb in the standard model, the reward output unit 4021 determines the evaluation value based on the information about the phase delay calculated under each of the conditions to determine the reward from the sum of the evaluation values. In the following description, a phase delay, which is a state variable related to the state information S, is represented by D(S), and a phase delay, which is a state variable related to a state S' changed from the state S by the action information A (the modification of the coefficients ω c , τ and k of the transfer function for the filter 130) is represented by D(S').
[0060] The reward output unit 4021 determines the evaluation value under each of the conditions and determines the sum of the evaluation values under each condition to set it to the reward. As a method for determining the evaluation value based on the information of the phase delay with the reward output unit 4021, for example, a method for determining the evaluation value depending on whether the frequency at which the phase delay reaches 180 degrees is increased, decreased, or not changed when the state S is changed to the state S' can be adopted. Although the case where the phase delay is 180 degrees is described here, there is no particular limitation to 180 degrees, and another value may be adopted. For example, when the phase delay is determined by the phase diagram in Fig. 3 is displayed when the state S is changed to the state S' and the curve is modified so that the frequency at which the phase delay reaches 180 degrees is reduced (towards X2 in Fig. 3), the phase delay is increased. On the other hand, if the state S is changed to the state S' and the curve is changed so that the frequency at which the phase delay reaches 180 degrees is increased (towards X1 in Fig. 3), the phase delay is reduced.
[0061] Therefore, when the state S is changed to the state S' and the frequency at which the phase delay reaches 180 degrees is decreased, it is defined as phase delay D(S) < phase delay D(S'), and the reward output unit 4021 sets the evaluation value to the second negative value. The absolute value of the second negative value is set lower than the first negative value. On the other hand, when the state S is changed to the state S' and the frequency at which the phase delay reaches 180 degrees is increased, it is defined as phase delay D(S) > phase delay D(S'), and the reward output unit 4021 sets the evaluation value to a positive value.When the state S is changed to the state S' and the frequency at which the phase delay reaches 180 degrees is not changed, it is defined as phase delay D(S) = phase delay D(S'), and the reward output unit 4021 sets the evaluation value to zero. The method for determining the evaluation value based on the phase delay information is not limited to the method described above, and another method may be adopted.
[0062] If it is defined that the phase delay D(S') in state S' after performing action A is greater than the phase delay D(S) in the previous state S, with respect to a negative value, the negative value may be increased according to a ratio. For example, in the first method described above, the negative value is preferably increased according to the degree of reduction in frequency. In contrast, with respect to a positive value, if it is defined that the phase delay D(S') in state S' after performing action A is less than the phase delay D(S) in the previous state S, the positive value may be increased according to a ratio. For example, in the first method described above, the positive value is preferably increased according to the degree of increase in frequency.
[0063] The reward output unit 4021 determines the evaluation value under each of the conditions. Then, the reward output unit 4021 adds the evaluation values under each condition to determine the reward. This reward is the sum of the evaluation values under each condition of the machine tool. As described above, the reward output unit 4021 provides the first negative evaluation value when the curve RC 1 before learning the calculated input / output gain, the curve MC 1 of the ideal input / output gain in the standard model. Since the curve RC 1 before learning the calculated input / output gain, the curve MC 1 of the ideal input / output gain in the standard model, the reward output unit 4021 does not determine the evaluation value based on the phase delay. The evaluation value is the first negative evaluation value when the curve RC 1 before learning the calculated input / output gain, the curve MC 1 of the ideal input / output gain in the Standard Model.
[0064] The value function updating unit 4022 performs Q-learning based on the state S, the action A, the state S' when the action A is applied to the state S, and the reward calculated as described above to update the value function Q stored in the value function storage unit 404. The updating of the value function Q can be performed by online learning, batch learning, or mini-batch learning. Online learning is a learning method in which the value function Q is immediately updated each time a certain action A is applied to the current state S, so that the state S is changed to the new state S'.Batch learning is a learning technique in which the application of a specific action A to the current state S is repeated in such a way that the state S is changed to the new state S', in which learning data is thus collected and in which all collected learning data is used to update the value function Q. Furthermore, mini-batch learning is a learning technique that is halfway between online learning and batch learning, in which the value function Q is updated every time a certain amount of learning data is stored.
[0065] The action information generation unit 4023 selects the action A in the process of Q-learning for the current state S. In order to perform an operation (corresponding to the action A in Q-learning) for modifying the coefficients ω c , τ and k of the transfer function for the filter 130 in the process of Q-learning, the action information generation unit 4023 generates the action information A and outputs the generated action information A to the action information output unit 403. More specifically, the action information generation unit 4023 can, for example, generate the coefficients ω c , τ and k of the transfer function for the filter 130 contained in the action A incrementally to or from the coefficients ω c , τ and k of the transfer function for the filter 130 contained in state S.
[0066] Although all coefficients ω c , τ and k can be changed, some of the coefficients can be changed. The center frequency fc at which the oscillation occurs is easy to find, and thus the center frequency fc is easy to identify. In order to perform an operation to temporarily set the center frequency fc, change the bandwidth fw and the damping coefficient k, that is, to set the coefficient ω c (= 2π fc) and to change the coefficient τ (= fw / fc) and the damping coefficient k, the action information generation unit 4023 can therefore generate the action information A and output the generated action information A to the action information output unit 403. In the characteristic of the filter 130, as shown in Fig. 13, the gain and phase are changed by the bandwidth fw of the filter 130. In Fig. 13, a dashed line indicates a case where the bandwidth fw is large, and a solid line indicates a case where the bandwidth fw is small. In the characteristic of the filter 130, as shown in Fig. 14, the gain and phase are changed by the damping coefficient k of the filter 130. In Fig. 14, a dashed line indicates a case where the damping coefficient k is low, and a solid line indicates a case where the damping coefficient k is high.
[0067] The action information generation unit 4023 may take measures to select an action A' by a known method, such as a greedy method to select the action A' having the highest value Q(S,A) among the values of the currently estimated actions A, or an ε-greedy method to randomly select the action A' with a small probability ε or otherwise to select the action A' having the highest value Q(S,A).
[0068] The action information output unit 403 is a unit that transmits the action information A output by the learning unit 402 to the filter 130. As described above, based on this action information, the filter 130 modifies the current state S, that is, the coefficients ω c , τ and k, which are currently set to change to the subsequent state S' (i.e., the modified coefficients of the filter 130).
[0069] The value function storage unit 404 is a storage unit that stores the value function Q. The value function Q can be stored as a table (hereinafter referred to as an action value table), for example, for each state S or each action A. The value function Q stored in the value function storage unit 404 can be shared with another machine learning unit 400. The value function Q is shared by multiple machine learning units 400, so that reinforcement learning can be performed in the machine learning unit 400 by distributing it, with the result that the efficiency of reinforcement learning can be improved.
[0070] The optimization action information output unit 405 generates, based on the value function Q updated by performing Q-learning with the value function updating unit 4022, the action information A (hereinafter referred to as "optimization action information") that causes the filter 130 to perform such an operation as to maximize the value Q(S,A). More specifically, the optimization action information output unit 405 acquires the value function Q stored in the value function storage unit 404. As described above, this value function Q has been updated by performing Q-learning with the value function updating unit 4022. Then, the optimization action information output unit 405 generates the action information based on the value function Q and outputs the generated action information to the filter 130.As with the action information output by the action information output unit 403 in the process of Q-learning, the optimization action information includes information on changing the coefficients ω. c , τ and k of the transfer function for the filter 130.
[0071] In the filter 130, based on this action information, the coefficients ω c , τ and k of the transfer function are modified. Through the operation described above, the machine learning unit 400 optimizes the coefficients ω c , τ, and k of the transfer function for the filter 130, so that the machine learning unit 400 can be operated to reduce vibrations on the machine side. Then, the machine learning unit 400 can perform optimal adjustment of the filter characteristics even when the machine characteristics change depending on conditions, for example, even when the machine characteristics change depending on the position of an axis or when the machine characteristics are influenced by another axis. As described above, the machine learning unit 400 of the present invention is used, and thus it is possible to simplify the setting of the parameters of the filter 130.
[0072] The functional blocks included in the control device 10 have been described above. To implement these functional blocks, the control device 10 includes an operation processing device such as a CPU (Central Processing Unit). The control device 10 also includes an auxiliary storage device such as an HDD (Hard Disk Drive) for storing various control programs such as application software and an operating system, and a main storage device such as a RAM (Random Access Memory) for storing data temporarily required when the operation processing device executes programs.
[0073] In the control device 10, the operation processing device reads the application software and the operating system from the auxiliary storage device and performs operation processing based on the application software and the operating system, while developing the application software and the operating system that are read in the main storage device. The control device 10 also controls various types of hardware provided in individual devices based on the result of the operation. In this way, the functional blocks of the present embodiment are realized. In other words, the present embodiment can be realized through the cooperation of hardware and software.
[0074] Since the machine learning unit 400 has a large number of operations related to machine learning, GPUs (Graphics Processing Units), for example, are incorporated into a personal computer and preferably used for machine learning-related operation processing through a technology called GPGPUs (General-Purpose Computing on Graphics Processing Units), so that high-speed processing can be performed. To perform higher-speed processing, a computer cluster may be constructed including a plurality of computers equipped with such GPUs, and the computers included in the computer cluster may perform parallel processing.
[0075] The operation of the machine learning unit 400 at the time of Q-learning in the present embodiment will then be explained using the flowchart in Fig. 15 described.
[0076] In step S11, the state information acquisition unit 401 acquires the initial state information S from the servo control unit 100 and the frequency generation unit 200. The acquired state information is output to the value function update unit 4022 and the action information generation unit 4023. As described above, the state information S is information corresponding to a state in Q-learning.
[0077] An input / output gain (amplitude ratio) Gs(S 0 ) and a phase delay D(S 0 ) under each of the conditions in a state S 0 at the time of the first start of Q-learning are obtained by the frequency characteristic calculation unit 300 by driving the servo control unit 100 with the speed command, which is the sine wave whose frequency is changed. The speed command value and the speed detection value are input to the frequency characteristic calculation unit 300, and the input / output gain (amplitude ratio) Gs(S 0 ) and the phase delay D(S 0 ) under each of the conditions output from the frequency characteristic calculation unit 300 are sequentially input to the state information acquisition unit 401 as initial state information. The initial values of the coefficients ω c , τ and k of the transfer function for the filter 130 are previously generated by the user, and the initial values of the coefficients ω c , τ and k are supplied to the state information acquisition unit 401 as initial state information.
[0078] In step S12, the action information generation unit 4023 generates new action information A and outputs the generated new action information A to the filter 130 via the action information output unit 403. The action information generation unit 4023 outputs the new action information A based on the above-described measures. The servo control unit 100, which has received the action information A, drives the servo motor 150 with the speed command, which is the sine wave whose frequency is changed, based on the received action information in the state S' in which the coefficients ω c , τ and k of the transfer function for the filter 130 are changed with respect to the current state S. As described above, this action information corresponds to action A in Q-learning.
[0079] In step S13, the state information acquisition unit 401 acquires, as new state information, the input / output gain (amplitude ratio) Gs(S'), the phase delay D(S'), and the coefficients ω c , τ and k of the transfer function from the filter 130 in the new state S'. The obtained information about the new state is output to the reward output unit 4021.
[0080] In step S14, the reward output unit 4021 determines whether the input / output gain G(S') at each frequency in state S' is less than or equal to the input / output gain Gb at each frequency in the standard model. If the input / output gain G(S') at each frequency is greater than the input / output gain Gb at each frequency in the standard model (No in step S14), the reward output unit 4021 sets the evaluation value to the first negative value in step S15 and returns to step S12.
[0081] If the input / output gain G(S') at each frequency in state S' is less than or equal to the input / output gain Gb at each frequency in the standard model (Yes in step S14), the reward output unit 4021 provides a positive evaluation value when the phase delay D(S') is smaller than the phase delay D(S), provides a negative evaluation value when the phase delay D(S') is greater than the phase delay D(S), or provides an evaluation value of zero when the phase delay D(S') does not change compared to the phase delay D(S). Although the method described above is mentioned as a method for determining the evaluation value such that, for example, the phase delay is reduced, there is no particular limitation to this method, and another method may be used.
[0082] In step S16, specifically e.g. when the state S in the phase diagram of Fig. 3 is changed to the state S' and the frequency at which the phase delay is 180 degrees is decreased, it is defined as phase delay D(S) < phase delay D(S'), and the reward output unit 4021 sets the evaluation value to the second negative value in step S17. The absolute value of the second negative value is set lower than the first negative value. When the state S is changed to the state S' and the frequency at which the phase delay is 180 degrees is increased, it is defined as phase delay D(S) > phase delay D(S'), and the reward output unit 4021 sets the evaluation value to a positive value in step S18.When the state S is changed to the state S' and the frequency at which the phase delay is 180 degrees is not changed, it is defined as phase delay D(S) = phase delay D(S'), and the reward output unit 4021 sets the evaluation value to a zero value in step S19.
[0083] When any one of steps S17, S18, and S19 is completed, it is determined in step S20 whether or not the evaluation values are determined under a plurality of conditions. If the evaluation values are not determined under the conditions, that is, if one of the conditions exists under which the evaluation value is not determined, the process returns to step S13, the condition is changed to the condition under which the evaluation value is not determined, and thus the state information is acquired. If the evaluation values are determined under the conditions in step S21, the evaluation values (evaluation values calculated in any one of steps S17, S18, and S19) determined under each condition are added together, and the sum of the evaluation values is set to the reward.Then, in step S22, based on the reward value calculated by the value function in step S21, the value function Q stored in the value function storage unit 404 is updated by the value function updating unit 4022. Then, the process returns to step S12 again, the above-described processing is repeated, and thus the value function Q converges to an appropriate value. The above-described processing can be completed under the condition that the processing is repeated a predetermined number of times or for a predetermined time. Although the online update is illustrated in step S21, the online update may be replaced by a batch update or mini-batch update.
[0084] As described above, in the present embodiment, in the manner described with reference to Fig. 15, the machine learning unit 400 is used, and thus it is possible to determine the corresponding value function for setting the coefficients ω c , τ and k of the transfer function for the filter 130. An operation in which optimization action information is generated with the optimization action information output unit 405 will be described below with reference to the flowchart in Fig. 16. In step S23, the optimization measure information output unit 405 first acquires the value function Q stored in the value function storage unit 404. As described above, the value function Q has been updated by performing Q-learning with the value function updating unit 4022.
[0085] In step S24, the optimization action output unit 405 generates the optimization action information based on the value function Q and outputs the generated optimization action information to the filter 130.
[0086] In the present embodiment, it is possible to use the method described with reference to Fig. 16, it is possible to generate the information about the optimization action on the basis of the value function Q determined by learning with the machine learning unit 400, on the basis of the information about the optimization action, the setting of the coefficients ω c , τ and k of the transfer function for the filter 130, which are currently set, to reduce the vibrations at the machine end and to improve the quality of the machined surface of a workpiece.
[0087] In the above-discussed embodiment, a description is given by the example of learning by changing a plurality of the above-discussed conditions (a), (b), and (c), as well as the frequency characteristics of the input / output gain and the phase delay. However, the above-described conditions (a), (b), and (c) can be combined as needed so that they can be learned by the machine learning unit 400. For example, although the frequency characteristics of the Y-axis can be affected by the position of the Y-axis itself, the position of the Z-axis, and the speed gain of the Z-axis servo control unit, they can be combined to set a variety of conditions.Specifically, the Y-axis machine learning unit 400 may combine a plurality of conditions as needed among first multiple conditions of the positions of the left end, center, and right end of the Y-axis itself, second multiple conditions of the positions of the left end, center, and right end of the Z-axis, and third multiple conditions of the speed gain of the Z-axis servo control unit to perform the learning.
[0088] The individual components of the control device described above can be implemented using hardware, software, or a combination thereof. A servo control method performed through the interaction of the individual components contained in the control device described above can also be implemented using hardware, software, or a combination thereof. Software implementation here refers to implementation achieved by reading and executing programs with a computer.
[0089] The programs can be stored and delivered to the computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media includes various types of tangible storage media. Examples of non-transitory computer-readable media are a magnetic recording medium (e.g., a hard disk drive), a magneto-optical recording medium (e.g., a magneto-optical disk), a CD-ROM (read-only memory), a CD-R, a CD-R / W, and semiconductor memories (e.g., a mask ROM, a PROM (programmable ROM), an EPROM (erasable PROM), a flash ROM, and a RAM (random access memory)). The programs can be delivered to the computer using various types of transient computer-readable media.
[0090] Although the above-described embodiment is a preferred embodiment of the present invention, the scope of the present invention is not limited to the above-described embodiment only, and various modifications may be made without departing from the gist of the present invention.
[0091] Although the above-discussed embodiment describes the case where the machine driven by the servomotor 150 has a single resonance point, the machine may have multiple resonance points. When the machine has multiple resonance points, multiple filters corresponding to the resonance points are provided and connected in series, making it possible to dampen the entire resonance. Fig. Figure 17 is a block diagram showing an example where multiple filters are directly connected to form the filter. Fig. Since 17 m (m is a natural number of two or more) resonance points are present, the filter 130 is formed by connecting m filters 130-1 to 130-m in series. The optimal values for the attenuation of the resonance points are sequentially determined by machine learning with respect to the coefficients ω c , τ and k of the m filters 130-1 to 130-m.
[0092] The control device may also be a device other than that in Fig. 1 shown configuration. <Variante, bei der die Maschinenlerneinheit außerhalb der Servosteuereinheit vorgesehen ist>
[0093] Fig. 18 is a block diagram showing another configuration example of the control device. Fig. The control device 10A shown in Figure 18 differs from that shown in Fig. 1 in that n (n is a natural number of two or more) servo control units 100A-1 to 100A-n are connected to n machine learning units 400A-1 to 400A-n via a network 600, and each of them includes the frequency generation unit 200 and the frequency characteristic calculation unit 300. The machine learning units 400A-1 to 400A-n have the same configuration as those shown in Fig. 10. The servo control units 100A-1 to 100A-n each correspond to the servo control device, and the machine learning units 400A-1 to 400A-n each correspond to the machine learning device. Of course, one or both of the frequency generation unit 200 and the frequency characteristic calculation unit 300 may be provided outside the servo control units 100A-1 to 100A-n.
[0094] Here, the servo control unit 100A-1 and the machine learning unit 400A-1 are sequentially paired and connected to each other so that they can communicate with each other. The servo control units 100A-2 to 100A-n and the machine learning units 400A-2 to 400A-n are connected in the same way as the servo control unit 100A-1 and the machine learning unit 400A-1. Although in Fig. 18 n pairs of servo control units 100A-1 to 100A-n and machine learning units 400A-1 to 400A-n are connected via the network 600, the n pairs of servo control units 100A-1 to 100A-n and machine learning units 400A-1 to 400A-n can be connected so that the servo control unit and the machine learning unit of each pair are directly connected via a connection interface. Regarding the n pairs of servo control units 100A-1 to 100A-n and machine learning units 400A-1 to 400A-n, for example, multiple pairs can be deployed in the same factory or in different factories.
[0095] The network 600 is, for example, a LAN (Local Area Network) established within a factory, on the Internet, in a public switched telephone network, or a combination thereof. In the network 600, a specific communication method, such as whether the network 600 uses a wired connection or a wireless connection, and the like, is not particularly limited. <Flexibilität der Systemkonfiguration>
[0096] Although in the above-described embodiment, the servo control units 100A-1 to 100A-n and the machine learning units 400A-1 to 400A-n are sequentially paired and connected so as to be able to communicate with each other, a configuration may be adopted, for example, in which one machine learning unit is connected to a plurality of servo control units via the network 600 so as to be able to communicate with them, and thus machine learning is performed on the servo control units. In this case, a distributed processing system may be adopted in which the functions of the one machine learning unit are distributed among a plurality of servers as needed. The functions of the one machine learning unit can be realized by utilizing a virtual server function or the like in a cloud.
[0097] When there are n machine learning units 400A-1 to 400A-n, each corresponding to n servo control units 100A-1 to 100A-n of the same model name, specifications, or series, the machine learning units 400A-1 to 400A-n can be configured to share the learning results among the machine learning units 400A-1 to 400A-n. This allows for the construction of a more optimal model.
[0098] The machine learning apparatus, the control apparatus, and the machine learning apparatus according to the present invention may adopt not only the above-described embodiment but also various kinds of embodiments having configurations as described below. (1) A machine learning device (machine learning unit 400) that performs reinforcement learning in which a servo control device (servo control unit 100) for controlling a motor (servo motor 150) is driven under a variety of conditions and that optimizes a coefficient of at least one filter (filter 130) for attenuating at least one specific frequency component provided in the servo control device, the machine learning device comprising: a state information acquisition unit (state information acquisition unit 401) that acquires state information including the result of calculation of a frequency characteristic calculation device (frequency characteristic calculation unit 300) for calculating an input / output gain of the servo control device and / or a phase delay of an input and an output, the coefficients of the filter, and the conditions;an action information output unit (action information output unit 403) that outputs action information including coefficient setting information included in the state information to the filter; a reward output unit (reward output unit 4021) that individually determines evaluation values under the conditions based on the result of the calculation, thereby outputting the value of a sum of the evaluation values as a reward;and a value function updating unit (value function updating unit 4022) that updates an action value function based on the value of the reward output by the reward output unit, the state information, and the action information. In the machine learning device described above, it is possible to perform optimal adjustment of the filter characteristic even when the machine characteristic changes depending on conditions, for example, even when the machine characteristic changes depending on the position of an axis or even when the machine characteristic is affected by another axis. (2) The machine learning device according to (1) described above, in which the motor drives an axis in a machine tool, a robot, or an industrial machine, and in which the conditions are a plurality of positions of the axis. In the machine learning device described above, it is possible to perform the optimal setting of the filter characteristic even when the machine characteristic is changed depending on multiple positions of an axis in a machine tool, a robot, or an industrial machine. (3) The machine learning device according to (1) described above, wherein the motor drives one axis in a machine tool, a robot, or an industrial machine, and the conditions are a plurality of positions of another axis located on the axis or below the axis. In the machine learning device described above, it is possible to perform optimal adjustment of the filter characteristic even when the machine characteristic is changed depending on a plurality of positions of another axis located on one axis or below the one axis in a machine tool, a robot, or an industrial machine. (4) The machine learning device according to (1) described above, wherein the motor drives one axis in a machine tool, a robot, or an industrial machine, and the conditions represent a plurality of speed gains of the servo control device that drives another axis arranged on the axis or below the axis. In the machine learning device described above, it is possible to perform the optimal setting of the filter characteristic even when the machine characteristic is changed depending on a plurality of speed gains of the servo control device that drives another axis arranged on one axis or below the one axis in a machine tool, a robot, or an industrial machine. (5) The machine learning device according to any one of (1) to (4) described above, wherein the frequency characteristic calculation device uses a sinusoidal input signal whose frequency is changed and speed feedback information of the servo control device to calculate at least one of the input / output gain and the phase delay of the input and the output. (6) The machine learning device according to any one of (1) to (5) described above, wherein a weight is set for each of the evaluation values according to each of the conditions. In the machine learning device described above, the weight for each of the evaluation values can be set according to the degree of influence, even if the influences of the individual conditions exerted on the machine learning device are different from each other. (7) The machine learning device according to any one of (1) to (6) described above, comprising: an optimization action information output unit (optimization action information output unit 405) that outputs the adjustment information of the coefficient based on the value function updated by the value function updating unit. (8) A control device comprising: the machine learning device (machine learning unit 400) according to any one of (1) to (7) described above; the servo control device (servo control unit 100) having the at least one filter for attenuating the at least one specific frequency component and controlling the motor; and the frequency characteristic calculation device (frequency characteristic calculation unit 300) calculating the input / output gain of the servo control device and / or the phase delay of the input and output in the servo control device. In the control device described above, it is possible to perform the optimal adjustment of the filter characteristic even when the machine characteristic is changed depending on the conditions, e.g.even if the machine characteristics are changed depending on the position of an axis or even if the machine characteristics are influenced by another axis. (9) A machine learning method of a machine learning device (machine learning unit 400) that performs reinforcement learning in which a servo control device (servo control unit 100) for controlling a motor (servo motor 150) is driven under a variety of conditions and that optimizes a coefficient of at least one filter (filter 130) for attenuating at least one specific frequency component provided in the servo control device, the machine learning method comprising: acquiring state information including the result of a calculation for calculating an input / output gain of the servo control device and / or a phase delay of an input and an output, the coefficients of the filter, and the conditions; outputting action information including setting information of the coefficient included in the state information to the filter;Determining evaluation values individually under the conditions based on the calculation result, thereby determining the value of a sum of the evaluation values as a reward; and updating an action value function based on the value of the determined reward, the state information, and the action information. With the machine learning method described above, it is possible to perform optimal adjustment of the filter characteristic even when the machine characteristic changes depending on conditions, for example, even when the machine characteristic changes depending on the position of an axis or even when the machine characteristic is affected by another axis. EXPLANATION OF REFERENCE SYMBOLS 10, 10A control device 100, 100-1 to 100-n servo control unit 110 subtractors 120 Speed control unit 130 filters 140 Power control unit 150 servo motor 200 Frequency generation unit 300 Frequency characteristic calculation unit 400 machine learning units 400A-1 to 400A-n machine learning unit 401 Status information acquisition unit 402 Learning Unit 403 Action information output unit 404 Value function storage unit 405 Optimization action information output unit 500 control target 600 network< / nichtpatentdokument>
Claims
[1] A machine learning unit (400) that performs reinforcement learning in which a servo control learning unit (100) for controlling a motor (150) is driven under a plurality of conditions, and that optimizes a coefficient of at least one filter (130) for attenuating at least one specific frequency component provided in the servo control learning unit (100), the machine learning unit (400) comprising: a state information acquisition unit (401) that acquires state information including a result of calculation by a frequency characteristic calculation unit (300) for calculating an input / output gain of the servo control unit (100) and / or a phase delay of an input and an output, the coefficients of the filter (130), and the conditions; an action information output unit (403) that outputs to the filter (130) action information comprising setting information of the coefficient contained in the state information; a reward output unit (4021) that individually determines evaluation values under the conditions based on the result of the calculation so as to output a value of a sum of the evaluation values as a reward; and a value function updating unit (4022) that updates an action value function based on a value of the reward output by the reward output unit (4021), the state information, and the action information. [2] The machine learning unit (400) according to claim 1, wherein the motor (150) drives an axis in a machine tool, a robot, or an industrial machine, and the conditions are a plurality of positions of the axis. [3] The machine learning unit (400) according to claim 1, wherein the motor (150) drives an axis in a machine tool, a robot, or an industrial machine, and the conditions are a plurality of positions of another axis that is on the axis or located below the axis. [4] The machine learning unit (400) according to claim 1, wherein the motor (150) drives an axis in a machine tool, a robot, or an industrial machine, and the conditions are a plurality of speed gains of the servo control unit (100) that drives another axis arranged on the axis or located below the axis. [5] The machine learning unit (400) according to any one of claims 1 to 4, wherein the frequency characteristic calculation unit (300) uses a sinusoidal input signal whose frequency is changed and speed feedback information of the servo controller (100) to calculate the input / output gain and / or the phase delay of the input and the output. [6] The machine learning unit (400) according to any one of claims 1 to 5, wherein a weight is set for each of the evaluation values corresponding to each of the conditions. [7] The machine learning unit (400) according to any one of claims 1 to 6, comprising: an optimization action information output unit (405) that outputs the adjustment information of the coefficient based on the value function updated by the value function updating unit (4022). [8] Control device (10) comprising: the machine learning unit (400) according to any one of claims 1 to 7; the servo control unit (100) comprising the at least one filter (130) for attenuating the at least one specific frequency component and controlling the motor (150); and the frequency characteristic calculation unit (300) that calculates the input / output gain of the servo control unit (100) and / or the phase delay of the input and the output in the servo control unit (100). [9] A machine learning method of a machine learning unit (400) that performs reinforcement learning in which a servo control unit (100) for controlling a motor (150) is driven under a plurality of conditions and that optimizes a coefficient of at least one filter (130) for attenuating at least one specific frequency component provided in the servo control unit (100), the machine learning method comprising: acquiring state information including a result of the calculation for calculating an input / output gain of the servo control unit (100) and / or a phase delay of an input and an output, the coefficients of the filter (130), and the conditions; outputting to the filter (130) action information including setting information of the coefficient contained in the state information; individually determining evaluation values under conditions based on the result of the calculation, so as to determine a value of a sum of the evaluation values as a reward; and updating an action value function based on a value of the determined reward, the state information, and the action information.
Citation Information
Patent Citations
machine learning apparatus, servo control system and machine learning method
DE102018003769A1
XY stage control device
JP1987126402A
Vibration reduction device for robot
JP1995261853A
Servo controller and adjustment method of the same
JP2013126266A
Servo control device having function for displaying online automatic adjustment situation of control system
JP2017022855A