Machine learning device, control device, and machine learning method

Through machine learning, optimized filter parameters of servo control devices, the problem of difficulty in parameter adjustment in the prior art is solved, and the optimal setting of the filter and the improvement of servo control performance are achieved.

CN111082729BActive Publication Date: 2025-07-11FANUC LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN201910926693.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-10-02
Filing Date
2019-09-27
Publication Date
2025-07-11
Estimated Expiration
2039-09-27

AI Technical Summary

Technical Problem

In the prior art, when determining the parameters of the notch filter, it is difficult to adjust to the optimal value, resulting in insufficient resonance suppression or deterioration of servo control performance.

Method used

The machine learning device is used to optimize the filter parameters of the servo control device, obtain input and output gain and phase delay information through the measurement device, and adjust the coefficients of the filter using reinforcement learning to optimize the characteristics of the filter.

Benefits of technology

The optimal setting of filter parameters is achieved, effectively suppressing mechanical resonance and improving servo control performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111082729B_ABST
    Figure CN111082729B_ABST
Patent Text Reader

Abstract

The present invention provides a machine learning device, a control device, and a machine learning method, which can easily set parameters for determining filter characteristics. The machine learning device performs machine learning for optimizing the coefficients of at least one filter provided in a servo control device that controls the rotation of a motor. The filter is a filter that attenuates specific frequency components, and the coefficients of the filter are optimized based on measurement information from a measurement device that measures at least one of the input-output gain and the input-output phase delay of the servo control device according to an input signal and an output signal that vary in frequency in the servo control device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a machine learning device that performs machine learning for optimizing coefficients of a filter in a servo control device, a control device including the machine learning device, and a machine learning method, wherein the servo control device controls the rotation of a motor of a machine tool, a robot, an industrial machine, or the like. Background Art

[0002] For example, Patent Documents 1 and 2 disclose devices for automatically adjusting filter characteristics.

[0003] Patent Document 1 discloses a servo actuator that has a speed feedback loop for controlling the speed of a motor, inserts a notch filter unit into the speed feedback loop to remove mechanical resonance, and the servo actuator includes: a data collection unit that acquires data representing the frequency response characteristics of the speed feedback loop; a moving average unit that performs a moving average process on the data acquired by the data collection unit; a comparison unit that compares the data acquired by the moving average unit with the data acquired by the data collection unit and extracts the resonance characteristics of the speed feedback loop; a notch filter setting unit that sets the frequency and Q value of the notch filter unit based on the resonance characteristics extracted by the comparison unit.

[0004] Patent Document 2 discloses a servo actuator that, in a tuning mode, superimposes an AC signal obtained by scanning a frequency on a signal of a speed command value, detects the amplitude of a torque command value signal obtained from a speed control unit as a result of the superimposition, and sets the frequency of the torque command value signal when the change rate of the amplitude changes from positive to negative as the center frequency of the notch filter.

[0005] Patent Document 3 discloses a control device for a motor that includes: a notch filter capable of changing parameters of the notch filter including a notch frequency and a notch width; a vibration frequency estimation unit that estimates a vibration frequency. It also includes: a notch filter parameter setting unit that sets the frequency between the notch frequency of the notch filter and the estimated vibration frequency as the new notch frequency of the notch filter and changes the notch width so that the original notch frequency component and the estimated frequency component are attenuated.

[0006] Prior Art Documents

[0007] Patent Document 1: Japanese Patent Application Laid-Open No. 2009-104439

[0008] Patent Document 2: Japanese Patent Application Laid-Open No. 5-19858

[0009] Patent Document 3: Japanese Patent Application Laid-Open No. 2008-312339

[0010] The servo actuator in Patent Document 1 adjusts the characteristics of the notch filter according to the frequency response characteristics of the speed feedback loop. The servo actuator in Patent Document 2 uses the torque command value signal to adjust the characteristics of the notch filter. In Patent Document 3, the frequency between the notch frequency of the notch filter and the estimated vibration frequency is set as the new notch frequency of the notch filter, and the notch width is changed so that the original notch frequency component and the estimated frequency component are attenuated, thereby adjusting the characteristics of the notch filter.

[0011] However, when determining the characteristics of the notch filter, it is necessary to determine multiple parameters such as the attenuation coefficient, the center frequency of the frequency band to be removed, and the bandwidth, and it is difficult to adjust these parameters to obtain the optimal value. And when the settings of these parameters are inappropriate, sometimes resonance cannot be sufficiently suppressed, or the phase delay of the servo control unit increases, resulting in deterioration of the servo control performance. Summary of the Invention

[0012] An object of the present invention is to provide a machine learning device, a control device including the machine learning device, and a machine learning method that can set the optimal parameters of the filter of the servo control device.

[0013] (1) A machine learning device according to the present invention (for example, the machine learning unit 400 described later) performs machine learning for optimizing the coefficients of at least one filter provided in a servo control device (for example, the servo control unit 100 described later), and the servo control device controls the rotation of a motor (for example, the servo motor 150 described later).

[0014] The filter is a filter that attenuates specific frequency components (for example, the filter 130 described later).

[0015] The machine learning device optimizes the coefficients of the filter according to the measurement information of a measurement device (for example, the measurement unit 300 described later), wherein the measurement device measures at least one of the input-output gain and the input-output phase delay of the servo control device based on the input signal and the output signal whose frequencies change in the servo control device.

[0016] (2) In the machine learning device of the above (1), it may be that

[0017] The input signal whose frequency changes is a sine wave whose frequency changes, and this sine wave is generated by a frequency generation device, and the frequency generation device is provided inside or outside the servo control device.

[0018] (3) In the machine learning device of the above (1) or (2), it may be that

[0019] The machine learning device has:

[0020] A state information acquisition unit (for example, the state information acquisition unit 401 described later) that acquires state information including the measurement information and the coefficients of the filter;

[0021] An action information output unit (for example, the action information output unit 403 described later) that outputs action information to the filter, where the action information includes adjustment information of the coefficients included in the state information;

[0022] A reward output unit (for example, the reward output unit 4021 described later) that outputs a reward value in reinforcement learning based on the measurement information; and

[0023] A value function update unit (for example, the value function update unit 4022 described later) that updates the action value function according to the reward value output by the reward output unit, the state information, and the action information.

[0024] (4) In the machine learning device in (3) above, it may be that

[0025] The measurement information includes the input-output gain and the phase delay of the input-output;

[0026] When the input-output gain of the servo control device included in the measurement information is less than or equal to the input-output gain of the standard model of the input-output gain calculated according to the characteristics of the servo control device, the reward output unit calculates a reward based on the phase delay of the input-output.

[0027] (5) In the machine learning device in (4) above, it may be that

[0028] The input-output gain of the standard model is a fixed value above a specified frequency.

[0029] (6) In the machine learning device in (4) or (5) above, it may be that

[0030] The reward output unit calculates a reward to reduce the phase delay of the input-output.

[0031] (7) In the machine learning device according to any one of (3) to (6) above, it may be that

[0032] The machine learning device has an optimized action information output unit (for example, the optimized action information output unit 405 described later) that outputs adjustment information of the coefficients according to the value function updated by the value function update unit.

[0033] (8) A control device according to the present invention (for example, the control device 10 described later) has:

[0034] The machine learning device according to any one of the above (1) to (7) (for example, the machine learning unit 400 described later);

[0035] A servo control device (for example, the servo control unit 100 described later), which has at least one filter that attenuates a specific frequency component, and the servo control device controls the rotation of a motor; and

[0036] A measuring device (for example, the measuring unit 300 described later), which measures at least one of the input-output gain and the phase delay of the input-output of the servo control device based on the input signal and the output signal whose frequencies change in the servo control device.

[0037] (9) A machine learning method for a machine learning device (for example, the machine learning unit 400 described later) according to the present invention, where the machine learning device performs machine learning to optimize the coefficients of at least one filter provided in a servo control device (for example, the servo control unit 100 described later), the servo control device controls the rotation of a motor (for example, the servo motor 150 described later), and the filter attenuates a specific frequency component,

[0038] Optimize the coefficients of the filter according to the measurement information of a measuring device (for example, the measuring unit 300 described later), where the measuring device measures at least one of the input-output gain and the phase delay of the input-output in the servo control device based on the input signal and the output signal whose frequencies change in the servo control device.

[0039] Advantages of the Invention

[0040] According to the present invention, there is provided a machine learning device, a control device including the machine learning device, and a machine learning method, which can set the optimal parameters of the filter of the servo control device. Description of the Drawings

[0041] Figure 1 It is a block diagram showing a control device including a machine learning device according to an embodiment of the present invention.

[0042] Figure 2 It is a diagram showing the speed command as the input signal and the detected speed as the output signal.

[0043] Figure 3 It is a diagram showing the frequency characteristics of the amplitude ratio and the phase delay between the input signal and the output signal.

[0044] Figure 4 It is a block diagram showing the machine learning unit according to an embodiment of the present invention.

[0045] Figure 5It is a block diagram of a model that becomes a model for calculating the input-output gain standard model.

[0046] Figure 6 It is a characteristic diagram showing the frequency characteristics of the input-output gain of the servo control unit representing the standard model and the servo control units before and after learning.

[0047] Figure 7 It is a characteristic diagram showing the relationship between the bandwidth, gain, and phase of the filter.

[0048] Figure 8 It is a characteristic diagram showing the relationship between the attenuation coefficient, gain, and phase of the filter.

[0049] Figure 9 It is a flowchart showing the operation of the machine learning unit during Q learning in this embodiment.

[0050] Figure 10 It is a flowchart explaining the operation of the optimization behavior information output unit of the machine learning unit according to an embodiment of the present invention.

[0051] Figure 11 It is a block diagram showing an example of constructing a filter by directly connecting multiple filters.

[0052] Figure 12 It is a block diagram showing another structural example of the control device.

[0053] Symbol Explanation

[0054] 10, 10A Control Device

[0055] 100, 100-1 to 100-n Servo Control Unit

[0056] 110 Subtractor

[0057] 120 Speed Control Unit

[0058] 130 Filter

[0059] 140 Current Control Unit

[0060] 150 Servo Motor

[0061] 200 Frequency Generation Unit

[0062] 300 Measurement Unit

[0063] 400 Machine Learning Unit

[0064] 400A-1 to 400A-n Machine Learning Unit

[0065] 401 State Information Acquisition Unit

[0066] 402 Learning Department

[0067] 403 Behavioral Information Output Department

[0068] 404 Value Function Storage Department

[0069] 405 Optimized Behavioral Information Output Department

[0070] 500 Network Specific Embodiment

[0071] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0072] Figure 1 is a block diagram of a control device of a machine learning device including an embodiment of the present invention. The control object of the control device 10 is, for example, a machine tool, a robot, or industrial machinery. The control device 10 can also be provided as a part of a control object such as a machine tool, a robot, or industrial machinery.

[0073] The control device 10 includes: a servo control unit 100, a frequency generation unit 200, a measurement unit 300, and a machine learning unit 400. The servo control unit 100 corresponds to a servo control device, the measurement unit 300 corresponds to a measurement device, and the machine learning unit 400 corresponds to a machine learning device.

[0074] In addition, one or more of the frequency generation unit 200, the measurement unit 300, and the machine learning unit 400 can be provided inside the servo control unit 100.

[0075] The servo control unit 100 includes: a subtractor 110, a speed control unit 120, a filter 130, a current control unit 140, and a servo motor 150. The subtractor 110, the speed control unit 120, the filter 130, the current control unit 140, and the servo motor 150 form a speed feedback loop.

[0076] The subtractor 110 calculates the difference between the input speed command and the detected speed of the speed feedback, and outputs this difference as a position deviation to the speed control unit 120.

[0077] The speed control unit 120 adds the value obtained by integrating the integral gain K1v multiplied by the speed deviation and the value obtained by multiplying the proportional gain K2v by the speed deviation, and outputs the result as a torque command to the filter 130.

[0078] The filter 130 is a filter that attenuates specific frequency components. For example, a notch filter is used. In a machine such as a machine tool driven by a motor, there are resonance points, and sometimes the resonance is increased due to the servo control unit 100. The resonance can be reduced by using a notch filter. The output of the filter 130 is output as a torque command to the current control unit 140.

[0079] Mathematical formula 1 (hereinafter referred to as Mathematical formula 1) represents the transfer function F(s) of the notch filter of the filter 130. The parameters represent the coefficients ω c , τ, k.

[0080] The coefficient k of Mathematical formula 1 is the attenuation coefficient, and the coefficient ω c is the center angular frequency, and the coefficient τ is a specific frequency band. Setting the center frequency to fc and the bandwidth to fw, then the coefficient ω c is represented by ω = 2πfc c , and the coefficient τ is represented by τ = fw / fc.

[0081]

Mathematical formula 1

[0082]

[0083] The current control unit 140 generates a current command for obtaining the current of the servo motor 150 according to the torque command, and outputs the current command to the motor 150.

[0084] The rotational angle position of the servo motor 150 is detected by a rotary encoder (not shown) provided in the servo motor 150, and the speed detection value is input to the subtractor 110 as speed feedback.

[0085] The servo control unit 100 is configured as described above. However, in order to perform machine learning on the optimal parameters of the filter, the control device 10 further includes: a frequency generation unit 200, a measurement unit 300, and a machine learning unit 400.

[0086] The frequency generation unit 200 outputs a sine wave signal as a speed command to the subtractor 110 and the measurement unit 300 of the servo control unit 100 while changing the frequency.

[0087] The measurement unit 300 uses the speed command (sine wave) that becomes the input signal generated by the frequency generation unit 200 and the detected speed (sine wave) that becomes the output signal output from the rotary encoder (not shown), and obtains the amplitude ratio (input-output gain) and phase delay of the input signal and the output signal for each frequency specified by the speed command. Figure 2 is a diagram showing the speed command that becomes the input signal and the detected speed that becomes the output signal. Figure 3 is a diagram showing the frequency characteristics of the amplitude ratio and phase delay of the input signal and the output signal.

[0088] As Figure 2 shown, the speed command output from the frequency generation unit 200 changes the frequency, and the frequency characteristics of the input-output gain (amplitude ratio) and phase delay as Figure 3 shown are obtained.

[0089] The machine learning unit 400 uses the input-output gain (amplitude ratio) and phase delay output from the measurement unit 300 to perform machine learning (hereinafter referred to as learning) on the coefficients ω, τ, and k of the transfer function of the filter 130. Learning of the machine learning unit 400 is performed before shipment, but relearning can also be performed after shipment. c Before the description of each functional block included in the machine learning unit 400, first, the basic structure of reinforcement learning will be described. An agent (corresponding to the machine learning unit 400 in the present embodiment) observes the environmental state, selects a certain action, and the environment changes according to this action. As the environment changes, a certain reward is provided, and the agent learns better action selection (decision-making).

[0090] <Machine learning unit 400>

[0091] In the following description, the case where reinforcement learning is performed on the machine learning unit 400 will be described, but the learning performed by the machine learning unit 400 is not particularly limited to reinforcement learning. For example, the present invention can also be applied to the case of performing supervised learning.

[0092] Supervised learning represents a complete correct answer, while the rewards in reinforcement learning are mostly fragmentary values based on partial changes in the environment. Therefore, the agent learns to select actions to maximize the sum of future rewards.

[0093] In this way, in reinforcement learning, by learning actions, appropriate actions are learned based on the interaction between the actions and the environment, that is, the learning method for maximizing the future rewards obtained is learned. This means that in the present embodiment, it is possible to obtain, for example, action information such as selecting an action for suppressing vibration at the mechanical end, which affects future actions.

[0094] Here, any learning method can be used as reinforcement learning. In the following description, the case of using Q-learning in a certain environmental state S will be described as an example. Q-learning is a method of learning the value Q(S, A) of selecting an action A.

[0095] Q-learning aims to select the action A with the highest value Q(S, A) from the actions A that can be taken in a certain state S as the optimal action.

[0096] However, at the time when Q-learning starts for the first time, the correct value of the value Q(S, A) is completely unknown for the combination of the state S and the action A. Therefore, the agent selects various actions A in a certain state S, and for the action A at that time, selects a better action according to the given reward, and thus continues to learn the correct value Q(S, A).

[0097] But at the initial time point of starting Q-learning, the correct value of the value Q(S, A) for the combination of the state S and the action A is completely unknown. Therefore, the agent selects various actions A in a certain state S, and for the action A at that time, selects a better action according to the given reward, and thus continues to learn the correct value Q(S, A).

[0098] In addition, in order to maximize the sum of future rewards, the goal is ultimately to make Q(S, A) = E[Σ(γ t )r t . Here, E[] represents the expected value, t represents the time, γ represents a parameter called the discount rate described later, and r t represents the reward at time t, and Σ is the sum at time t. The expected value in this mathematical formula is the expected value when the optimal behavior state changes. However, during the Q-learning process, since the optimal behavior is unknown, various behaviors are performed, and reinforcement learning is carried out while searching. The update formula for such a value Q(S, A) can be expressed, for example, by the following Mathematical Formula 2 (hereinafter referred to as Mathematical Formula 2).

[0099]

Mathematical Formula 2

[0100]

[0101] In the above Mathematical Formula 2, S t represents the environmental state at time t, and A t represents the behavior at time t. Through the behavior A t , the state changes to S t+1 . r t+1 represents the reward obtained through this state change. In addition, the term with max is: when in the state S t+1 , γ is multiplied by the Q value when selecting the behavior A with the highest known Q value at that time. Here, γ is a parameter where 0 < γ ≤ 1 and is called the discount rate. In addition, α is the learning coefficient, and the range of α is set to 0 < α ≤ 1.

[0102] The above Mathematical Formula 2 represents the following method: according to the reward r t fed back as a result of trying A t+1 , the value Q(S t , A t ) of the behavior A t in the state S t is updated.

[0103] This update formula indicates that: if the value of the optimal behavior max t Q(S t+1 , A) in the next state S a caused by the behavior A t+1 is greater than the value Q(S t , A t ) of the behavior A t in the state S t , then Q(S t , A t ) is increased; conversely, if it is small, then Q(S t , A t)。That is, to make the value of a certain behavior in a certain state close to the optimal behavior value in the next state caused by that behavior. Among them, although this difference varies due to the existence form of the discount rate γ and the return r t+1 , it is basically a structure in which the optimal behavior value in a certain state spreads to the behavior value in its previous state.

[0104] Here, the Q-learning has the following method: A table of Q(S, A) for all state-action pairs (S, A) is created for learning. However, sometimes the number of states is too large to find the values of Q(S, A) for all state-action pairs, making it take a long time for Q-learning to converge.

[0105] Therefore, a well-known technique called DQN (Deep Q-Network) can be used. Specifically, an appropriate neural network can be used to construct the value function Q, and the parameters of the neural network are adjusted. Thus, the value Q(S, A) is calculated by approximating the value function Q through an appropriate neural network. By using DQN, the time required for Q-learning to converge can be shortened. In addition, DQN is described in detail in the following non-patent literature, for example.

[0106] <Non-patent literature>

[0107] “Human-level control through deep reinforcement learning”, by Volodymyr Mnih1 [online], [searched on January 17, 2017], Internet <URL: http: / / files.davidqiu.com / research / nature14236.pdf>

[0108] The machine learning unit 400 performs the Q-learning described above. Specifically, the machine learning unit 400 learns the following value Q: regarding the values of the coefficients ω c , τ, k of the transfer function of the filter 130, the input-output gain (amplification ratio) and phase delay output from the measurement unit 300 as the state S, and the adjustment selection of the values of the coefficients ω c , τ, k of the transfer function of the filter 130 related to the state S is selected as the action A.

[0109] The machine learning unit 400 determines according to the coefficients ω of the transfer function of the filter 130 c, τ, k, determines the behavior A based on the observation state information S, where the state information S includes the input-output gain (amplification ratio) and phase delay for each frequency obtained from the measurement unit 300 by driving the servo control circuit 100 using a sine wave whose frequency changes, i.e., the speed command. The machine learning unit 400 returns a reward whenever the behavior A is performed. The machine learning unit 400, for example, tries to search for the optimal behavior A by trial and error to maximize the sum of future rewards. By doing so, the machine learning unit 400 can select the optimal behavior A for the state S (i.e., the optimal coefficients ω c , τ, k) of the transfer function of the filter 130, where the state S includes, according to the coefficients ω c , τ, k of the transfer function of the filter 130, the input-output gain (amplification ratio) and phase delay for each frequency obtained from the measurement unit 300 by driving the servo control unit 100 using a sine wave whose frequency changes, i.e., the speed command.

[0110] That is, according to the value function Q learned by the machine learning unit 400, the behavior A that maximizes the Q value is selected among the behaviors A that apply the coefficients ω c , τ, k of the transfer function of the filter 130 related to a certain state S. Thus, the behavior A that minimizes the vibration of the mechanical end generated by executing the learning program can be selected (i.e., the coefficients ω c , τ, k of the transfer function of the filter 130).

[0111] Figure 4 is a block diagram of the machine learning unit 400 showing an embodiment of the present invention.

[0112] To perform the above-mentioned reinforcement learning, as Figure 4 shown, the machine learning unit 400 has: a state information acquisition unit 401, a learning unit 402, a behavior information output unit 403, a value function storage unit 404, and an optimal behavior information output unit 405. The learning unit 402 has: a reward output unit 4021, a value function update unit 4022, and a behavior information generation unit 4023.

[0113] The state information acquisition unit 401 acquires the state S from the measurement unit 300, where the state S includes the input-output gain (amplitude ratio) and phase delay obtained by driving the servo motor 150 using a speed command (sine wave) according to the coefficients ω c , τ, k of the transfer function of the filter 130. This state information S corresponds to the environmental state S in Q-learning.

[0114] The state information acquisition unit 401 outputs the acquired state information S to the learning unit 402.

[0115] In addition, the coefficients ω c of the transfer function of the filter 130 at the time point when Q-learning starts initially, τ, and k are generated in advance by the user. In the present embodiment, the initial setting values of the coefficients ω c of the transfer function of the filter 130 created by the user are adjusted to be optimal through reinforcement learning.

[0116] In addition, regarding the coefficients ω c of the transfer function of the filter 130, τ, and k, when the operator has adjusted the machine tool in advance, the adjusted values can be used as initial values for machine learning.

[0117] The learning unit 402 is a part that learns the value Q(S, A) when selecting a certain action A in a certain environmental state S.

[0118] The reward output unit 4021 is a part that calculates the reward when an action A is selected in a certain state S.

[0119] The reward output unit 4021 compares the input-output gain Gs measured when the coefficients ω c of the transfer function of the filter 130 are corrected, τ, and k with the input-output gain Gb of each frequency of the preset standard model. When the measured input-output gain Gs is larger than the input-output gain Gb of the standard model, the reward output unit 4021 gives a negative reward. On the other hand, when the measured input-output gain Gs is less than or equal to the input-output gain Gb of the standard model, the reward output unit 4021 gives a positive reward when the phase delay becomes smaller, gives a negative reward when the phase delay becomes larger, and gives a zero reward when the phase delay remains unchanged.

[0120] First, the actions of the reward output unit 4021 giving a negative reward when the measured input-output gain Gs is larger than the input-output gain Gb of the standard model will be described using Figure 5 and Figure 6 .

[0121] The reward output unit 4021 stores the standard model of the input-output gain. The standard model is a model of the servo control unit with ideal characteristics of no resonance. The standard model can be calculated, for example, based on the inertia Ja, torque constant K Figure 5 shown in the model, the proportional gain K t , the integral gain K p , the differential gain K I , and the differential gain K D . The inertia Ja is the sum of the motor inertia and the mechanical inertia.

[0122] Figure 6It is a characteristic diagram showing the frequency characteristics of the input-output gains of the servo control unit representing the standard model and the servo control units 100 before and after learning. As Figure 6 shown in the characteristic diagram, the standard model has: Region A and Region B. Among them, Region A is a frequency region where the input-output gain becomes a certain value or more, for example, an ideal input-output gain of -20 dB or more, and Region B is a frequency region where the input-output gain is less than a certain value. In Figure 6 Region A, the ideal input-output gain of the standard model is represented by the curve MC1 (thick line). In Figure 6 Region B, the ideal virtual input-output gain of the standard model is represented by the curve MC 11 (thick dashed line), and the input-output gain of the standard model is represented as a fixed value by the straight line MC 12 (thick line). In Figure 6 Regions A and B, the curves of the input-output gains of the servo control units before and after learning are represented by the curves RC1 and RC2 respectively.

[0123] The reward output unit 4021 gives a first negative reward in Region A when the curve RC1 of the input-output gain before learning exceeds the curve MC1 of the ideal input-output gain of the specified model.

[0124] In Region B where the frequency at which the input-output gain becomes sufficiently small is exceeded, even if the curve RC1 of the input-output gain before learning exceeds the curve MC of the ideal virtual input-output gain of the standard model 11 , the influence on stability becomes small. Therefore, in Region B, as described above, the curve MC of the input-output gain of the standard model is not the ideal gain characteristic 11 , but the straight line MC of the input-output gain with a fixed value (for example, -20 dB) 12 is used. However, when the curve RC1 of the measured input-output gain before learning exceeds the straight line MC of the input-output gain with a fixed value 12 , it may be unstable, so a first negative value is given as a reward.

[0125] Next, an explanation will be given of the operation of the reward output unit 4021 to determine the reward based on the phase delay information when the measured input-output gain Gs is less than or equal to the input-output gain Gb of the standard model.

[0126] In the following explanation, D(S) represents the state variable related to the state information S, that is, the phase delay, and D(S’) represents the state variable related to the state S’ that has changed from the state S through the behavior information A (the correction of the coefficients ω c , τ, k of the transfer function of the filter 130), that is, the phase delay.

[0127] The reward output unit 4021 determines the method of reward based on the information of phase delay. For example, there are the following three methods.

[0128] The first method is: when changing from state S to state S', the method of determining the reward by the frequency of the phase delay being 180 degrees becoming larger, smaller, or the same. Here, the case where the phase delay is 180 degrees is listed, but it is not particularly limited to 180 degrees and can also be other values.

[0129] For example, when Figure 3 the phase delay is represented by the shown phase line graph, when changing from state S to state S', if the curve changes such that the frequency of the phase delay being 180 degrees becomes smaller (in the Figure 3 X2 direction of Figure 3 ), then the phase delay becomes larger. On the other hand, when changing from state S to state S', if the curve changes such that the frequency of the phase delay being 180 degrees becomes larger (in the Figure 3 X1 direction of

[0130] ), then the phase delay becomes smaller.

[0130] Therefore, when changing from state S to state S', when the frequency of the phase delay being 180 degrees becomes smaller, it is defined that the phase delay D(S) < phase delay D(S'), and the reward output unit 4021 sets the reward value to the second negative value. In addition, the absolute value of the second negative value is set to be smaller than the first negative value.

[0131] On the other hand, when changing from state S to state S', when the frequency of the phase delay being 180 degrees becomes larger, it is defined that the phase delay D(S) > phase delay D(S'), and the reward output unit 4021 sets the reward value to a positive value.

[0132] In addition, when changing from state S to state S', when the frequency of the phase delay being 180 degrees remains unchanged, it is defined that the phase delay D(S) = phase delay D(S'), and the reward output unit 4021 sets the reward value to zero.

[0133] The second method is: when changing from state S to state S', the method of determining the reward by the absolute value of the phase delay when the input-output gain crossover is 0 dB becoming larger, smaller, or the same.

[0134] For example, when the input gain is represented by the shown gain line graph in state S, the phase delay corresponding to the point where the crossover is 0 dB (hereinafter, referred to as the "zero crossover point") in the Figure 3 shown phase line graph is -90 degrees. Figure 3 shown in

[0135] When changing from state S to state S', when the absolute value of the phase delay at the zero crossing point increases, it is defined that phase delay D(S) < phase delay D(S'), and the reward output unit 4021 sets the reward value to the second negative value.

[0136] On the other hand, when changing from state S to state S', when the absolute value of the phase delay at the zero crossing point decreases, it is defined that phase delay D(S) > phase delay D(S'), and the reward output unit 4021 sets the reward value to a positive value.

[0137] In addition, when changing from state S to state S', when the absolute value of the phase delay at the zero crossing point remains unchanged, it is defined that phase delay D(S) = phase delay D(S'), and the reward output unit 4021 sets the reward value to zero.

[0138] The third method is: when changing from state S to state S', a method of determining the reward by the phase margin increasing, decreasing, or remaining the same. The so-called phase margin refers to: when the gain is 0 dB, the amount indicating how many degrees the phase is from -180 degrees. For example, in Figure 3 when the gain is 0 dB, the phase is -90 degrees, so the phase margin is 90 degrees.

[0139] When changing from state S to state S', when the phase margin decreases, it is defined that phase delay D(S) < phase delay D(S'), and the reward output unit 4021 sets the reward value to the second negative value.

[0140] On the other hand, when changing from state S to state S', when the phase margin increases, it is defined that phase delay D(S) > phase delay D(S'), and the reward output unit 4021 sets the reward value to a positive value.

[0141] In addition, when changing from state S to state S', when the phase margin remains unchanged, it is defined that phase delay D(S) = phase delay D(S'), and the reward output unit 4021 sets the reward value to zero.

[0142] In addition, as the negative value when the phase delay D(S') after executing the action A is defined to be larger than the phase delay D(S) in the previous state S, the negative value can be set larger according to a ratio. For example, in the above first method, the negative value can be set larger according to the degree of frequency decrease. Conversely, as the positive value when the phase delay D(S') after executing the action A is defined to be smaller than the phase delay D(S) in the previous state S, the positive value can be set larger according to a ratio. For example, in the above first method, the positive value can be set larger according to the degree of frequency increase.

[0143] The value function update unit 4022 performs Q-learning based on the state S, the action A, the state S' when the action A is applied to the state S, and the reward value calculated as described above, thereby updating the value function Q stored in the value function storage unit 404.

[0144] The update of the value function Q can be performed by online learning, batch learning, or mini-batch learning.

[0145] Online learning is a learning method as follows: by applying a certain action A to the current state S, whenever the state S transitions to a new state S', the value function Q is updated immediately. In addition, batch learning is a learning method as follows: by repeatedly applying a certain action A to the current state S and the state S transitions to a new state S', learning data is collected, and all the collected learning data is used to update the value function Q. Furthermore, mini-batch learning is a learning method intermediate between online learning and batch learning, and is a learning method in which the value function Q is updated whenever a certain amount of learning data is accumulated.

[0146] The action information generation unit 4023 selects the action A during Q-learning for the current state S. During Q-learning, the action information generation unit 4023 generates action information A for the actions (equivalent to the action A in Q-learning) of the coefficients ω c , τ, k of the transfer function of the correction filter 130, and outputs the generated action information A to the action information output unit 403.

[0147] More specifically, the action information generation unit 4023, for example, adds or subtracts an increment to the coefficients ω c , τ, k of the transfer function of the filter 130 included in the state S, for the coefficients ω c , τ, k of the transfer function of the filter 130 included in the action A.

[0148] And the following strategy can be adopted: when the action information generation unit 4023 applies an increase or decrease in the coefficients ω c , τ, k of the transfer function of the filter 130 and transitions to the state S' and returns a positive reward (a positive-valued reward), as the next action A', it selects an action A' such that the coefficients ω c , τ, k of the transfer function of the filter 130 increase or decrease by the same increment as the previous action, etc., so that the measured phase delay is smaller than the previous phase delay.

[0149] In addition, conversely, the following strategy can be adopted: when a negative reward (a negative-valued reward) is returned, the action information generation unit 4023, as the next action A', for example, selects an action such that the coefficients ω c、τ, k decrease or increase by an increment, etc., in the opposite direction to the previous action, such that when the measured input-output gain is greater than the input-output gain of the standard model, the difference in the input gain compared to the previous time is smaller, or such that the measured phase delay is smaller than the previous phase delay, for behavior A'.

[0150] In addition, each coefficient ω c 、τ, k can all be corrected, or a part of the coefficients can be corrected. The center frequency fc at which resonance occurs is easy to find and the center frequency fc is easy to determine. Therefore, in order to temporarily fix the center frequency fc, correct the bandwidth fw, and the attenuation coefficient k, that is, in order to fix the coefficient ω c (= 2πfc), correct the coefficient τ(= fw / fc), and the action of the attenuation coefficient k, the behavior information generation unit 4023 can generate the behavior information A and output the generated behavior information A to the behavior information output unit 403.

[0151] In addition, the characteristics of the filter 130 are as Figure 7 shown, and the gain and phase change due to the bandwidth fw of the filter 130. In Figure 7 , the dashed line represents the case where the bandwidth fw is large, and the solid line represents the case where the bandwidth fw is small. In addition, the characteristics of the filter 130 are as Figure 8 shown, and the gain and phase change due to the attenuation coefficient k of the filter 130. In Figure 8 , the dashed line represents the case where the attenuation coefficient k is small, and the solid line represents the case where the attenuation coefficient k is large.

[0152] In addition, the behavior information generation unit 4023 can also adopt the following strategy: by using the well-known method such as the greedy algorithm that selects the behavior A' with the highest value Q(S, A) among the values of the currently estimated behavior A, or randomly selecting the behavior A' with a certain small probability ε, and otherwise selecting the behavior A' with the highest value Q(S, A) of the ε-greedy algorithm, to select the behavior A'.

[0153] The behavior information output unit 403 is the part that sends the behavior information A output from the learning unit 402 to the filter 130. As described above, based on this behavior information, the filter 130 makes a fine correction to the current state S, that is, the currently set coefficients ω c 、τ, k and transfers to the next state S' (that is, the coefficients of the corrected filter 130).

[0154] The value function storage unit 404 is a storage device that stores the value function Q. The value function Q can be stored as a table (hereinafter referred to as the action value table) according to the state S and the action A, for example. The value function Q stored in the value function storage unit 404 is updated by the value function update unit 4022. In addition, the value function Q stored in the value function storage unit 404 can also be shared among other machine learning units 400. If the value function Q is shared among multiple machine learning units 400, reinforcement learning can be performed dispersedly by each machine learning unit 400, and thus the efficiency of reinforcement learning can be improved.

[0155] The optimized action information output unit 405 generates action information A (hereinafter referred to as "optimized action information") for causing the filter 130 to perform the action that maximizes the value Q(S, A) according to the value function Q updated by the value function update unit 4022 through Q learning.

[0156] More specifically, the optimized action information output unit 405 acquires the value function Q stored in the value function storage unit 404. As described above, the value function Q is a function updated by the value function update unit 4022 through Q learning. And the optimized action information output unit 405 generates action information according to the value function Q and outputs the generated action information to the filter 130. This optimized action information is the same as the action information output by the action information output unit 403 during the process of Q learning and includes information on the coefficients ω c , τ, k for correcting the transfer function of the filter 130.

[0157] In the filter 130, the coefficients ω c , τ, k of the transfer function are corrected according to this action information.

[0158] The machine learning unit 400 can perform actions to optimize the coefficients ω c , τ, k of the transfer function of the filter 130 through the above actions and suppress the vibration at the mechanical end.

[0159] As described above, by using the machine learning unit 400 according to the present invention, the parameter adjustment of the filter 130 can be simplified.

[0160] The above has described the functional blocks included in the control device 10.

[0161] To implement these functional blocks, the control device 10 has an arithmetic processing device such as a CPU (Central Processing Unit). In addition, the control device 10 also has an auxiliary storage device such as an HDD (Hard Disk Drive) that stores various control programs such as application software or an OS (Operating System), and a main storage device such as a RAM (Random Access Memory) that stores data temporarily required after the arithmetic processing device executes the program.

[0162] Moreover, in the control device 10, the arithmetic processing device reads the application software or the OS from the auxiliary storage device, expands the read application software or the OS on the main storage device, and performs arithmetic processing according to these application software or the OS. In addition, according to the arithmetic result, various hardware devices are controlled. Thus, the functional blocks of the present embodiment are implemented. That is to say, the present embodiment can be implemented by the cooperation of hardware and software.

[0163] Regarding the machine learning unit 400, since the amount of computation associated with machine learning increases, for example, a technology called GPGPU (General-Purpose computing on Graphics Processing Units), which utilizes a GPU (Graphics Processing Units) installed in a personal computer, can perform high-speed processing when the GPU is used for the arithmetic processing associated with machine learning. And, to perform even faster processing, a computer cluster can be constructed using multiple computers equipped with such GPUs, and parallel processing can be performed by the multiple computers included in the computer cluster.

[0164] Next, with reference to Figure 9 the following process, the operation of the machine learning unit 400 during Q-learning in the present embodiment will be described.

[0165] In step S11, the state information acquisition unit 401 acquires the initial state information S from the servo control unit 100 and the control device of the frequency generation unit 200. The acquired state information is output to the value function update unit 4022 or the behavior information generation unit 4023. As described above, this state information S is information corresponding to the state in Q-learning.

[0166] The servo control circuit 100 is driven by a sine wave with a changing frequency, i.e., a speed command, to obtain the input-output gain (amplitude ratio) Gs(S0) and the phase delay D(S0) at the state S0 at the time point when Q-learning starts initially from the measurement unit 300. The speed command and the detected speed are input to the measurement unit 300, and the input-output gain (amplitude ratio) Gs(S0) and the phase delay D(S0) output from the measurement unit 300 are input to the state information acquisition unit 401 as the initial state information. The initial values of the coefficients ω c , τ, and k of the transfer function of the filter 130 are generated by the user in advance, and the initial values of the coefficients ω c , τ, and k are sent to the state information acquisition unit 401 as the initial state information.

[0167] In step S12, the behavior information generation unit 4023 generates new behavior information A, and outputs the generated new behavior information A to the filter 130 via the behavior information output unit 403. The behavior information generation unit 4023 outputs new behavior information A according to the above strategy. In addition, the servo control unit 100 that receives the behavior information A corrects the state S' of the coefficients ω c , τ, and k of the transfer function of the filter 130 related to the current state S, and drives the servo motor 150 using a sine wave with a changing frequency, i.e., a speed command. As described above, this behavior information corresponds to the behavior A in Q-learning.

[0168] In step S13, the state information acquisition unit 401 acquires the input-output gain (amplitude ratio) Gs(S'), the phase delay D(S'), and the coefficients ω c , τ, and k of the transfer function of the filter 130 as new state information in the new state S'. The acquired new state information is output to the reward output unit 4021.

[0169] In step S14, the reward output unit 4021 determines whether the input-output gain Gs(S') at each frequency in the state S' is less than or equal to the input-output gain Gb at each frequency of the standard model. If the input-output gain Gs(S') at each frequency is greater than the input-output gain Gb at each frequency of the standard model (step S14: no), in step S15, the reward output unit 4021 sets the reward to the first negative value and returns to step S12.

[0170] If the input-output gains Gs(S’) at each frequency in state S’ are less than the input-output gains Gb at each frequency of the standard model (step S14 is YES), when the phase delay D(S’) is less than the phase delay D(S), the reward output unit 4021 gives a positive reward; when the phase delay D(S’) is greater than the phase delay D(S), the reward output unit 4021 gives a negative reward; when the phase delay D(S’) has no change compared to the phase delay D(S), the reward output unit 4021 gives a zero reward. As described above, there are three methods for determining the reward to reduce the phase delay, for example. However, in the following example, the first method is described.

[0171] In step S16, specifically, for example, in Figure 3 the phase line graph, when changing from state S to state S’, when the frequency at a phase delay of 180 degrees becomes smaller, it is defined as phase delay D(S) < phase delay D(S’), and in step S17, the reward output unit 4021 sets the reward value to the second negative value. In addition, the absolute value of the second negative value is set to be smaller than the first negative value. When changing from state S to state S’, when the frequency at a phase delay of 180 degrees becomes larger, it is defined as phase delay D(S) > phase delay D(S’), and in step S18, the reward output unit 4021 sets the reward value to a positive value. Furthermore, when changing from state S to state S’, when the frequency at a phase delay of 180 degrees remains unchanged, it is defined as phase delay D(S) = phase delay D(S’), and in step S19, the reward output unit 4021 sets the reward value to zero.

[0172] When one of step S17, step S18, and step S19 ends, in step S20, according to the reward value calculated in that step, the value function update unit 4022 updates the value function Q stored in the value function storage unit 404. Then, it returns to step S11 again, and the above process is repeated, and the value function Q converges to an appropriate value. In addition, the process can be ended on the condition of repeating the above process a specified number of times or repeating the above process for a specified time.

[0173] In addition, step S20 illustrates online update, and it can be replaced with batch update or mini-batch update instead of online update.

[0174] As described above, through the actions described by referring to Figure 9 in this embodiment, by using the machine learning unit 400, an appropriate value function for obtaining the adjustment of each coefficient ω c , τ, k of the transfer function of the filter 130 can be obtained, and the optimization of each coefficient ω c , τ, k of the transfer function of the filter 130 can be simplified.

[0175] Next, with reference to Figure 10 the process of Figure 10 , the operation of the optimization behavior information output unit 405 when generating the optimization behavior information will be described.

[0176] First, in step S21, the optimization behavior information output unit 405 acquires the value function Q stored in the value function storage unit 404. As described above, the value function Q is a function updated by Q-learning by the value function update unit 4022.

[0177] In step S22, the optimization behavior information output unit 405 generates optimization behavior information based on the value function Q, and outputs the generated optimization behavior information to the filter 130.

[0178] In addition, by referring to Figure 10 the operations described in Figure 10 , in the present embodiment, it is possible to generate optimization behavior information based on the value function Q obtained by learning by the machine learning unit 400, and based on this optimization behavior information, simplify the coefficients ω c , τ, k of the transfer function of the currently set filter 130, and it is possible to suppress the vibration at the mechanical end and improve the quality of the machined surface of the workpiece.

[0179] Each structural part included in the above control device can be implemented by hardware, software, or a combination thereof. In addition, the servo control method performed by the cooperation of each structural part included in the above control device can also be implemented by hardware, software, or a combination thereof. Here, implementing by software means that the computer reads and executes a program to achieve it.

[0180] Various types of non-transitory computer-readable recording media can be used to store a program and provide the program to a computer. The non-transitory computer-readable recording media include various types of tangible storage media. Examples of non-transitory computer-readable recording media include: magnetic recording media (e.g., hard disk drive), magneto-optical recording media (e.g., magneto-optical disk), CD-ROM (Read Only Memory), CD-R, CD-R / W, semiconductor memories (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, RAM (random access memory). In addition, the program can also be supplied to the computer through various types of transitory computer-readable media.

[0181] The above-described embodiments are preferred embodiments of the present invention. However, the scope of the present invention is not limited only to the above-described embodiments, and various modifications can be made without departing from the spirit of the present invention.

[0182] In addition, in the above-described embodiment, the case where there is one resonance point in the machine driven by the servo motor 150 has been described. However, there are sometimes multiple resonance points in the machine. When there are multiple resonance points in the machine, a plurality of filters are provided corresponding to each resonance point and connected in series, whereby all resonances can be attenuated. Figure 11 It is a block diagram showing an example of constructing a filter by directly connecting a plurality of filters. In Figure 11 when there are m (m is a natural number of 2 or more) resonance points, the filter 130 is constituted by connecting m filters 130-1 to 130-m in series. For the respective coefficients ω c τ, k of the m filters 130-1 to 130-m, the optimum values for attenuating the resonance points are obtained in sequence by machine learning.

[0183] In addition, the structure of the control device has the following structure in addition to the Figure 1 structure.

[0184] <Modification example where the machine learning unit is provided outside the servo control unit>

[0185] Figure 12 It is a block diagram showing another structural example of the control device. Figure 12 The control device 10A shown is different from the Figure 1 control device 10 shown in that n (n is a natural number of 2 or more) servo control units 100A-1 to 100A-n are connected to n machine learning units 400A-1 to 400A-n via the network 500, and each has a frequency generation unit 200 and a measurement unit 300. The machine learning units 400A-1 to 400A-n have the same structure as the Figure 4 machine learning unit 400 shown. The servo control units 100A-1 to 100A-n respectively correspond to servo control devices. In addition, the machine learning units 400A-1 to 400A-n respectively correspond to machine learning devices. Of course, one or both of the frequency generation unit 200 and the measurement unit 300 can be provided outside the servo control units 100A-1 to 100A-n.

[0186] Here, the servo control unit 100A-1 and the machine learning unit 400A-1 are a one-to-one group and can be communicatively connected. The servo control units 100A-2 to 100A-n and the machine learning units 400A-2 to 400A-n are also connected in the same manner as the servo control unit 100A-1 and the machine learning unit 400A-1. InFigure 11 Among them, n groups of the servo control units 100A-1 to 100A-n and the machine learning units 400A-1 to 400A-n are connected via the network 500. For the n groups of the servo control units 100A-1 to 100A-n and the machine learning units 400A-1 to 400A-n, the servo control unit and the machine learning unit of each group can be directly connected via a connection interface. For these n groups of the servo control units 100A-1 to 100A-n and the machine learning units 400A-1 to 400A-n, for example, multiple groups can be set in the same factory, or they can be set in different factories respectively.

[0187] In addition, the network 500 is, for example, a LAN (Local Area Network) built within a factory, the Internet, a public telephone network, or a combination thereof. There is no particular limitation on whether the specific communication method in the network 500 is a wired connection or a wireless connection, etc.

[0188] <Degree of freedom of system structure>

[0189] In the above-described embodiment, the servo control units 100A-1 to 100A-n and the machine learning units 400A-1 to 400A-n are respectively connected in a one-to-one group in a communicable manner. However, for example, one machine learning unit can also be communicably connected to multiple servo control units via the network 500 to perform machine learning on each servo control unit.

[0190] At this time, each function of one machine learning unit can be used as a distributed processing system appropriately dispersed among multiple servers. In addition, each function of one machine learning unit can also be implemented by using a virtual server function, etc. on the cloud.

[0191] In addition, when there are n machine learning units 400A-1 to 400A-n respectively corresponding to n servo control units 100A-1 to 100A-n of the same model name, the same specification, or the same series, the learning results in each of the machine learning units 400A-1 to 400A-n can be shared. In this way, a more ideal model can be constructed.

Claims

1. A machine learning device that performs machine learning to optimize the coefficients of at least one filter provided in a servo control device, where the servo control device controls the rotation of a motor, characterized in that the filter is a filter that attenuates specific frequency components, the machine learning device includes: a state information acquisition unit that acquires state information including measurement information of a measurement device and the coefficients of the filter, where the measurement device measures at least one of the input-output gain and the input-output phase delay of the servo control device based on an input signal and an output signal with frequency variations in the servo control device; a behavior information output unit that outputs behavior information to the filter, and the behavior information includes adjustment information of the coefficients included in the state information; a reward output unit that outputs a reward value in reinforcement learning based on the measurement information; and a value function update unit that updates the value function according to the reward value, the state information, and the behavior information output by the reward output unit, the measurement information includes the input-output gain and the input-output phase delay, when the input-output gain of the servo control device included in the measurement information is less than or equal to the input-output gain of the standard model of the input-output gain calculated according to the characteristics of the servo control device, the reward output unit calculates the reward based on the input-output phase delay.

2. The machine learning device according to claim 1, characterized in that the input signal with frequency variations is a sine wave with frequency variations, and the sine wave is generated by a frequency generation device provided inside or outside the servo control device.

3. The machine learning device according to claim 1 or 2, characterized in that the input-output gain of the standard model is a fixed value above a specified frequency.

4. The machine learning device according to claim 1 or 2, characterized in that the reward output unit calculates the reward to make the input-output phase delay smaller.

5. The machine learning device according to claim 1 or 2, characterized in that the machine learning device has an optimized behavior information output unit that outputs the adjustment information of the coefficients according to the value function updated by the value function update unit.

6. A control device, characterized in that, Having: the machine learning device according to any one of claims 1 to 5; a servo control device that has at least one filter that attenuates specific frequency components, and the servo control device controls the rotation of a motor; and a measurement device that measures at least one of the input-output gain and the input-output phase delay of the servo control device based on an input signal and an output signal with frequency variations in the servo control device.

7. A machine learning method for a machine learning device that performs machine learning to optimize the coefficients of at least one filter provided in a servo control device, where the servo control device controls the rotation of a motor and the filter attenuates specific frequency components, characterized in that Obtain status information including the measurement information of the measuring device and the coefficients of the filter, where the measuring device measures at least one of the input-output gain and the input-output phase delay of the servo control device based on the input signal and the output signal with frequency variation in the servo control device. Output behavior information to the filter, where the behavior information includes adjustment information of the coefficients included in the status information. Output the reward value in the reinforcement learning based on the measurement information. Update the value function according to the output reward value, the status information, and the behavior information. The measurement information includes the input-output gain and the input-output phase delay. When the input-output gain of the servo control device included in the measurement information is less than or equal to the input-output gain of the standard model of the input-output gain calculated according to the characteristics of the servo control device, calculate the reward based on the input-output phase delay.

Citation Information

Patent Citations

  • Servo actuator

    JP1993019858A

  • Controller for motors

    JP2008312339A

  • Servo actuator

    JP2009104439A

  • Control device and control method

    WO2018151215A1