Stability margin setting support device, control system, and setting support method
The setting assistance device simplifies the adjustment of stability margins in servo control devices by allowing users to interactively modify gain and phase margins on a complex plane, enhancing usability and flexibility.
Patent Information
- Application Number
- JP2023554199
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-10-22
- Publication Date
- 2025-09-17
- Estimated Expiration
- 2041-10-22
AI Technical Summary
Setting stability margins in servo control devices requires specialized knowledge and is not easily adjustable.
A setting assistance device that includes a closed curve drawing unit, change unit, and scaling unit to facilitate user interaction in adjusting gain and phase margins on a complex plane, along with an adjustment unit for filter coefficients and feedback gains.
Enables easy and flexible setting and changing of stability margins in servo control devices.
Smart Images

Figure 0007741190000003 
Figure 0007741190000004 
Figure 0007741190000005
Abstract
Description
[Technical Field]
[0001] The present invention relates to a setting support device that supports a user in setting a stability margin of a servo control device, a control system including the setting support device, and a setting support method. [Background technology]
[0002] Patent Document 1 describes a setting support device that supports the setting of a plurality of control parameters used in control processing in a motor control device. Patent Document 1 describes that the setting support device includes a test operation instruction unit that changes the value of at least one of the control parameters to control the servo driver to actually perform a test operation or controls the test operation by a simulation using a virtual model, and a performance index calculation unit that calculates a performance index of control by the servo driver according to the result of the test operation. Patent Document 1 also describes that a speed proportional gain and a position proportional gain can be used as the control parameters, and a phase margin can be used as the performance index.
[0003] Patent Document 2 describes a design support device that supports the design of a DC power supply system in which power is supplied from a DC power supply via a DC bus to a plurality of servo devices each including an inverter circuit and an electric motor. Patent Document 2 describes a design support device that generates time-series current data indicating a time-varying pattern of the total current value supplied to multiple servo devices via a DC bus based on system information of a DC power supply system and operation pattern information contained therein that indicates the operation patterns of each of multiple servo devices, outputs information indicating the stability of the DC power supply system when the current flowing through the DC bus is at its maximum value of the generated time-series current data based on the output impedance Zo(s) of the power supply side of the DC power supply system and the input impedance Zin(s) of the load side of the DC power supply system, Zin(s) being a function of the current value flowing through the DC bus, etc., and changes the operation pattern information based on the information. Patent Document 2 also describes that the design support device displays a Nyquist diagram corresponding to the actual operation of the DC power supply system. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2017-167607 [Patent Document 2] Japanese Patent Publication No. 2020-120573 Summary of the Invention [Problem to be solved by the invention]
[0005] The stability margin during gain adjustment of a servo control device is often evaluated using multiple indices such as the gain margin, phase margin, and maximum closed-loop gain. Even in the gain of a servo control device that includes a filter and the automatic adjustment function of the filter, these indices must be determined in some way, but specialized knowledge is required to set them appropriately. Therefore, when a user sets the stability margin of a servo control device, it is desired that the stability margin be easily set and changed. [Means for solving the problem]
[0006] (1) A first aspect of the present disclosure is a setting assistance device that assists a user in setting a stability margin of a servo control device, comprising: a closed curve drawing unit that draws a closed curve on a complex plane that includes (-1, 0) on the complex plane and passes through a gain margin and a phase margin; a change unit that changes the gain margin and the phase margin based on a user operation; a closed curve scaling unit that scales or shrinks the closed curve in conjunction with the amount of change in the gain margin and the phase margin made by the change unit.
[0007] (2) A second aspect of the present disclosure is a setting support device according to (1) above, The control system includes an adjustment unit provided in the servo control device for adjusting the coefficient and feedback gain of at least one filter.
[0008] (3) A third aspect of the present disclosure is a computer that: A process of setting a first gain margin and a first phase margin, which are reference stability margins; A process of drawing a closed curve on the complex plane that includes (-1, 0) on the complex plane and passes through the first gain margin and the first phase margin; a process of changing the first gain margin and the first phase margin to a second gain margin and a second phase margin based on a user operation; expanding or contracting the closed curve in conjunction with an amount of change from the first gain margin and the first phase margin to the second gain margin and the second phase margin; This is a setting support method for performing the above. [Effects of the Invention]
[0009] According to each aspect of the present disclosure, when a user sets a stability margin of a servo control device, the stability margin can be easily set and changed. [Brief explanation of the drawings]
[0010] [Figure 1]FIG. 1 is a block diagram illustrating a control system according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of a setting support unit included in the control system. [Figure 3] FIG. 10 is a diagram showing a display screen of a display unit that displays a complex plane and a slide bar. [Figure 4] FIG. 2 is an enlarged view of a portion of a complex plane showing a unit circle and a circle that is a closed curve on the complex plane. [Figure 5] FIG. 10 is a diagram showing the display screen of the display unit that displays a slide bar and a circle on a complex plane when a slider is moved. [Figure 6] FIG. 2 is an enlarged view of a portion of a complex plane showing a unit circle on the complex plane and two circles that form closed curves. [Figure 7] FIG. 1 is a diagram illustrating a Nyquist locus, a unit circle, and a circle passing through a gain margin and a phase margin, drawn on a complex plane. [Figure 8] 10 is a flowchart showing the operation of a setting support unit. [Figure 9] FIG. 10 is a block diagram showing a modified example of the control system in which the adjustment unit 400 is replaced with a machine learning unit. [Figure 10] FIG. 2 is a block diagram showing the configuration of a machine learning unit. [Figure 11] A closed-loop Bode diagram. [Figure 12] FIG. 1 is a block diagram showing a reference model of a closed loop. [Figure 13] 10 is a characteristic diagram showing frequency characteristics of input / output gains of a servo control unit of a reference model and servo control units before and after learning. FIG. [Figure 14] FIG. 10 is a block diagram showing an example of a filter configured by directly connecting a plurality of filters. [Figure 15] FIG. 10 is a block diagram showing another example of the configuration of the control system. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings.
[0012] FIG. 1 is a block diagram illustrating a control system according to an embodiment of the present disclosure. The control system 10 includes a servo control unit 100, a frequency generation unit 200, a frequency characteristic calculation unit 300, an adjustment unit 400, and a setting support unit 500. The servo control unit 100 corresponds to a servo control device, and the setting support unit 500 corresponds to a setting support device. One or more of the frequency generating unit 200, the frequency characteristic calculating unit 300, the adjusting unit 400, and the setting assisting unit 500 may be provided within the servo control unit 100. The frequency characteristic calculating unit 300 may be provided within the adjusting unit 400. Furthermore, the setting assisting unit 500 may be provided within the adjusting unit 400.
[0013] The servo control unit 100 includes a subtractor 110, a speed control unit 120, a filter 130, a current control unit 140, and a motor 150. The subtractor 110, the speed control unit 120, the filter 130, the current control unit 140, and the motor 150 form a servo system with a closed speed feedback loop. The motor 150 may be a linear motor that moves in a straight line, or a motor with a rotating shaft. The object driven by the motor 150 may be, for example, a mechanical part of a machine tool, a robot, or an industrial machine. The motor 150 may be provided as part of the machine tool, the robot, the industrial machine, or the like. The control system 10 may be provided as part of the machine tool, the robot, the industrial machine, or the like. The configuration of the servo control unit 100 will be described in detail below.
[0014] The frequency generating unit 200 outputs a sinusoidal signal as a speed command to the subtractor 110 and the frequency characteristic calculating unit 300 of the servo control unit 100 while changing the frequency.
[0015] The frequency characteristic calculation unit 300 uses the speed command (sine wave) generated by the frequency generation unit 200 as an input signal, and the detected speed (sine wave) as an output signal output from a rotary encoder (not shown) provided on the motor 150 or the integral of the detected position (sine wave) as an output signal output from the linear scale, to determine the amplitude ratio (input / output gain) and phase delay between the input signal and the output signal for each frequency specified by the speed command, and outputs them to the adjustment unit 400.
[0016] The setting support unit 500 supports the user in setting the gain margin and phase margin (which become the stability margin) of the open-loop circuit of the servo control unit 100, and outputs image data including a closed curve passing through the gain margin and phase margin on the complex plane. The open-loop circuit is composed of the speed control unit 120, filter 130, current control unit 140, and motor 150 shown in FIG. 1.
[0017] The adjustment unit 400 adjusts one or both of the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficient ω of the transfer function of the filter 130 so that the gain margin and phase margin of the servo control unit 100 are equal to or greater than the values set by the user, using a closed curve passing through the gain margin and phase margin included in the image data output from the setting support unit 500 and a Nyquist locus obtained from the input / output gain (amplitude ratio) and phase delay output from the frequency characteristic calculation unit 300. c , τ, and / or δ is adjusted.
[0018] The servo control unit 100, the adjustment unit 400, and the setting support unit 500 will be described in further detail below. <Servo control unit 100> As already explained, the servo control unit 100 includes the subtractor 110, the speed control unit 120, the filter 130, the current control unit 140, and the motor 150.
[0019] The subtractor 110 calculates the difference between the input speed command and the detected speed that has been fed back, and outputs this difference to the speed control unit 120 as the speed deviation.
[0020] The speed control unit 120 adds together the value obtained by multiplying the speed deviation by integral gain K1v and integrating the result, and the value obtained by multiplying the speed deviation by proportional gain K2v, and outputs the result as a torque command to the filter 130. The speed control unit 120 is a control unit that sets the feedback gain.
[0021] The filter 130 is a filter that attenuates specific frequency components, and may be, for example, a notch filter, a low-pass filter, or a band-stop filter. A resonance point exists in a machine such as a machine tool having a mechanical part driven by a motor 150, and resonance may increase in the servo control unit 100. Resonance can be reduced by using a filter such as a notch filter. The output of the filter 130 is output to the current control unit 140 as a torque command. Equation 1 (hereinafter referred to as Equation 1) shows the transfer function F(s) of the notch filter as the filter 130. The parameters are coefficients ω c , τ, and δ are shown. The coefficient δ in Equation 1 is the damping coefficient, and the coefficient ω c is the central angular frequency, and the coefficient τ is the fractional bandwidth. If the central frequency is fc and the bandwidth is fw, then the coefficient ω c ω c =2πfc, and the coefficient τ is expressed as τ=fw / fc.
number
[0022] The current control unit 140 generates a current command for driving the motor 150 based on the torque command, and outputs the current command to the motor 150 . When the motor 150 is a linear motor, the position of the movable part is detected by a linear scale (not shown) provided on the motor 150, and the position detection value is differentiated to obtain a speed detection value, which is input to the subtractor 110 as speed feedback. If the motor 150 is a motor having a rotating shaft, the rotation angle position is detected by a rotary encoder (not shown) provided on the motor 150, and the detected speed value is input to the subtractor 110 as speed feedback.
[0023] <Setting support unit 500> Fig. 2 is a block diagram showing an example of the configuration of the setting support unit 500. As shown in Fig. 2, the setting support unit 500 includes a reference stability margin setting unit 501, a closed curve drawing unit 502, a display unit 503, a stability margin changing unit 504 serving as a changing unit, and a closed curve scaling unit 505.
[0024] The reference stability margin setting unit 501 sets a reference stability margin (gain margin and phase margin) (hereinafter referred to as reference stability margin). The reference stability margin may be set to a predetermined value in advance, or may be arbitrarily set by the user.
[0025] The closed curve drawing unit 502 draws a unit circle whose circumference passes through (-1,0) on a complex plane 5031 in FIGS. 3 and 4 (described later), and draws a closed curve that intersects with this unit circle and includes (-1,0) on the complex plane based on the reference stability margin output from the reference stability margin setting unit 501. In FIG. 3, the center of the circle is on the real axis, but it does not have to be on the real axis. As shown in FIG. 4, the point where the circle intersects with the real axis is the gain margin, and the point where the circle intersects with the unit circle is the phase margin. The closed curve may be a closed curve other than a circle, such as a rhombus, a rectangle, or an ellipse. In the following description, the closed curve is assumed to be a circle.
[0026] The display unit 503 displays a complex plane 5031, a slide bar 5032, and an end button 5033 on a display screen 503A shown in Fig. 3. Fig. 3 is a diagram showing the display screen 503A of the display unit 503. The complex plane 5031 is output from the closed curve drawing unit 502 as image data. The display unit 503 is configured with a liquid crystal display device, and the slide bar 5032 serves as an operation unit for the stability margin changing unit 504, which will be described later. FIG. 4 is an enlarged view of a part of a complex plane 5031, and shows a unit circle and a circle that is a closed curve on the complex plane.
[0027] The stability margin and slider 5032A of slide bar 5032 shown in Fig. 3 are linked, and the user can change the stability margin (gain margin and phase margin) by moving slider 5032A of slide bar 5032 shown in Fig. 3 left and right. In Fig. 3, slider 5032A moves left and right, but the direction in which the slider moves is not particularly limited, and it may also move up and down.
[0028] The stability margin changing unit 504 is a changing unit that changes the gain margin and the phase margin based on a user operation, and is incorporated as part of the display unit 503. The stability margin changing unit 504 is, for example, a slide bar, and the slide bar is an input element that displays a small mark (acting as a slider) indicating the current position on a bar-shaped area on the display screen, and is operated by sliding the mark within the bar-shaped area. 3, as described above, the stability margin changing unit 504 creates a slide bar 5032 as an operation unit and displays it on the display screen 503A. When the user moves a slider 5032A of the slide bar 5032 shown in FIG. 3 left or right, the stability margin changing unit 504 detects the amount of movement of the slider 5032A and outputs the amount of movement to the closed curve scaling unit 505. The amount of movement of the slider 5032A corresponds to the amount of change in the gain margin and the phase margin, and the gain margin and the phase margin can be changed by moving the slider 5032A. The stability margin changing unit 504 may be provided separately from the display unit 503. In this case, the slide bar 5032 is displayed on the display screen of a display device separate from the display unit 503.
[0029] The stability margin changing unit 504 may be an operating tool that is operated by sliding a knob-like slider, such as that used on a control panel of a machine. When the stability margin changing unit 504 is an operating tool, the stability margin changing unit 504 is provided separately from the display unit 503.
[0030] The closed curve scaling unit 505 scales or reduces the diameter of a circle that passes through the stability margins (gain margin and phase margin) with the center of the circle as the reference point, in accordance with the amount of movement of the slider 5032 A. Note that the reference point for the scale-up or scale-down may not be the center of the circle, but may be any other point on the real axis. If the user places importance on the stability of the servo system, the user moves slider 5032A of slide bar 5032 from the center to the left in Fig. 5, and if the user places importance on the responsiveness of the servo system, the user moves slider 5032A of slide bar 5032 from the center to the right in Fig. 5. In Fig. 3, for example, the center of slider 5032A of slide bar 5032 is taken as the standard, and the gain margin is set to 6 dB and the phase margin to 30 degrees. Also in Fig. 3, the limit values for emphasizing stability are set to 10 dB and 45 degrees, respectively, and the limit values for emphasizing responsiveness are set to 4.5 dB and 25 degrees, respectively.
[0031] When the user places importance on the stability of the servo system and moves the slider 5032A of the slide bar 5032 from the center to the left side in Figure 5, the closed curve scaling unit 505 expands the diameter of the circle to increase the stability margin (gain margin and phase margin), and outputs information about the circle with the expanded diameter (coordinates, radius, etc.) to the closed curve drawing unit 502. The closed curve drawing unit 502 draws a circle with an enlarged diameter on a complex plane, and the display unit 503 displays on the display screen the complex plane on which the circle with an enlarged diameter is drawn and a slide bar with the slider moved to the left side of Figure 5. The user looks at the complex plane on which the circle with an enlarged diameter is drawn, displayed on the display screen of the display unit 503, and if the user wants to change the stability margin again, moves the slider 5032A of the slide bar 5032 left or right. If the user does not want to change the stability margin, the user clicks the end button 5033 shown in Fig. 5. When the user clicks the end button 5033, the closed curve drawing unit 502 outputs image data relating to the complex plane including the circle that becomes the closed curve and the unit circle to the adjustment unit 400.
[0032] Based on the two reference stability margins set by the reference stability margin setting unit 501, the closed curve drawing unit 502 draws a unit circle whose circumference passes through (-1, 0) on the complex plane 5031 as shown in FIG. 6 , and can also draw two circles, a first circle and a second circle, that intersect with this unit circle based on the two reference stability margins (gain margin and phase margin). For example, the reference stability margin setting unit 501 reduces the stability margin for cutting feed of the machine tool and increases the stability margin for rapid feed. The closed curve drawing unit 502 draws the circle for cutting feed as the first circle and the circle for rapid feed as the second circle. Then, when the user moves the slider 5032A of the slide bar 5032, the closed curve scaling unit 505 changes the radii of the first and second circles. In FIG. 6, the first circle and the second circle have the same center, but the center of the first circle and the center of the second circle may be different. The number of circles that form the closed curve in the closed curve drawing unit 502 is not limited to two, and may be three or more as necessary. By using multiple closed curves in this way, it is possible to use closed curves that are appropriate for multiple modes of the servo control device, for example, cutting feed and rapid feed, for each mode.
[0033] <Adjustment section 400> The adjustment unit 400 acquires image data relating to a complex plane including a circle that forms a closed curve and a unit circle from the closed curve drawing unit 502. Furthermore, the adjustment unit 400 calculates the open-loop frequency characteristic H(jω) using the input / output gain (amplitude ratio) and phase delay output from the frequency characteristic calculation unit 300, and draws a Nyquist locus created from the open-loop frequency characteristic H(jω) on the complex plane including the circle that forms the closed curve and the unit circle. Fig. 7 is a diagram showing the Nyquist locus, the unit circle, and the circle that passes through the gain margin and the phase margin, drawn on the complex plane. The adjustment unit 400 adjusts the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficient ω of the transfer function of the filter 130 so that the Nyquist locus does not pass through the inside of the circle. c , τ, and δ are adjusted.
[0034] The method by which the adjustment section 400 creates the Nyquist locus will be described below. The speed feedback loop is composed of a subtractor 110 and an open loop circuit with a transfer function H. As already explained, the open loop circuit is composed of the speed control section 120, filter 130, current control section 140, and motor 150 shown in FIG. 1. When the input / output gain of the speed feedback loop at a certain frequency ω0 is c and the phase delay is θ, the closed loop frequency characteristic G(jω0) is c·e jθ The closed loop frequency characteristic G(jω0) can be expressed as G(jω0)=H(jω0) / (1+H(jω0)) using the open loop frequency characteristic H(jω0). Therefore, the open loop frequency characteristic H(jω0) at a certain frequency ω0 is H(jω0)=G(jω0) / (1-G(jω0))=c·e jθ / (1-c·e jθ ) can be calculated.
[0035] The adjustment unit 400 adjusts the integral gain K1v, the proportional gain K2v, and the coefficient ω c The servo control unit 100 is driven using a speed command (sine wave) whose frequency changes based on τ, τ, and δ, and the input / output gain and phase delay obtained are obtained from the frequency characteristic calculation unit 300. When the changing frequency is ω, the open-loop frequency characteristic H(jω) can be calculated by the relational expression H(jω)=G(jω) / (1-G(jω)), as described above. The adjustment unit 400 uses the input / output gain and phase delay obtained from the frequency characteristic calculation unit 300 to plot the open-loop frequency characteristic H(jω) on a complex plane, thereby creating a Nyquist locus. The initial Nyquist locus is determined by the integral gain K1v, proportional gain K2v, and coefficient ω c , τ, and δ, the servo control unit 100 is driven using a velocity command (sine wave). The Nyquist locus during adjustment is obtained by adjusting the integral gain K1v, the proportional gain K2v, and / or the coefficient ω c , τ, and δ, and drive the servo control unit 100 using a speed command (sine wave).
[0036] Next, the operation of the setting support unit 500 in this embodiment will be described with reference to the flowchart of FIG.
[0037] In step S11, the reference stability margin setting unit 501 sets stability margins (gain margin and phase margin) that serve as reference stability criteria.
[0038] In step S12, the closed curve drawing unit 502 draws a unit circle on the complex plane whose circumference passes through (-1,0), and draws a circle that is a closed curve that intersects with the unit circle based on the stability margins (gain margin and phase margin) that serve as the reference stability criterion.
[0039] In step S13, the display unit 503 displays on the display screen of the display unit 503 a complex plane showing the unit circle and the circle that becomes the closed curve drawn by the closed curve drawing unit 502, a slide bar, and an end button.
[0040] In step S14, the closed curve scaling unit 505 determines whether the slider of the slide bar has been moved by the user. If the slider of the slide bar has moved, the closed curve enlargement / reduction unit 505 proceeds to a discard step S15, and if the slider of the slide bar has not moved, the closed curve enlargement / reduction unit 505 proceeds to step S16.
[0041] In step S15, the closed curve scaling unit 505 scales or reduces the diameter of a circle that is centered on the reference point and passes through the stability margins (gain margin and phase margin) in accordance with the amount of movement of the slider, and then the process returns to step S12. In step S16 , the closed curve drawing unit 502 outputs, to the adjustment unit 400 , image data relating to a complex plane including the circle that forms the closed curve and the unit circle.
[0042] According to the present embodiment described above, the stability margin can be easily changed by using the stability margin change section such as a slide bar in the setting support section.
[0043] <Modification in which the adjustment unit is a machine learning unit> FIG. 9 is a block diagram showing a modified example of the control system in which adjustment unit 400 shown in FIG. 1 is replaced with machine learning unit 400A. The control system 10A is the same as the control system 10 shown in FIG. 1 except that the adjustment unit 400 uses a machine learning unit 400A. In the following explanation, we will explain the case where machine learning unit 400A performs reinforcement learning, but the learning performed by machine learning unit 400A is not limited to reinforcement learning, and the present invention is also applicable to cases where supervised learning is performed, for example.
[0044] Before describing each functional block included in the machine learning unit 400A, the basic mechanism of reinforcement learning will be described. An agent (corresponding to the machine learning unit 400A in this embodiment) observes the state of the environment, selects an action, and the environment changes based on that action. As the environment changes, some kind of reward is given, and the agent learns to select a better action (decision-making). While supervised learning provides a perfect answer, the rewards in reinforcement learning are often fractional values based on partial changes in the environment, so the agent learns to choose actions that maximize the total reward over the future.
[0045] In this way, reinforcement learning learns appropriate actions based on the interaction of the actions with the environment, i.e., it learns a learning method to maximize future rewards. In this embodiment, this means that it is possible to acquire actions that will have an impact on the future, such as selecting action information to suppress vibrations at the machine end.
[0046] Any learning method can be used for reinforcement learning, but the following explanation will use Q-learning, a method for learning the value Q(S,A) of selecting action A under a certain environmental state S, as an example. In Q-learning, the goal is to select the optimal action A with the highest value Q(S,A) from among the possible actions A when in a certain state S.
[0047] However, when Q-learning first begins, the correct value Q(S, A) for a combination of state S and action A is completely unknown. Therefore, the agent selects various actions A in a certain state S, and learns the correct value Q(S, A) by selecting the better action based on the reward given for each action A at that time.
[0048] Also, we want to maximize the total rewards we can get in the future, so we ultimately want Q(S,A)=E[Σ(γ t )r t ] where E[] represents the expected value, t is the time, γ is a parameter called the discount rate, which will be described later, and r t is the reward at time t, and Σ is the sum at time t. The expected value in this equation is the expected value when the state changes according to the optimal action. However, since it is unknown what the optimal action is in the process of Q-learning, reinforcement learning is performed by exploring by taking various actions. The update equation for such value Q(S,A) can be expressed, for example, by the following equation 2 (hereinafter referred to as equation 2).
[0049]
number
[0050] In the above formula 2, S t represents the state of the environment at time t, and A t represents the action at time t. Action A t Therefore, the status is S t+1 It changes to r t+1 represents the reward obtained by the change in the state. Also, the term with max represents the reward obtained by the change in the state S t+1 It is calculated by multiplying the Q value of the action A with the highest Q value known at that time by γ. Here, γ is a parameter with a range of 0<γ≦1 and is called the discount rate. Also, α is a learning coefficient with a range of 0<α≦1.
[0051] The above-mentioned formula 2 is tAs a result, the reward returned is r t+1 Based on this, state S t Action A in t The value of Q(S t ,A t ) is updated. This update formula is t Action A in t The value of Q(S t ,A t ) rather than Action A t Next state S t+1 The value of the best action in max a Q(S t+1 ,A) is larger, then Q(S t ,A t ) is increased, and conversely, if it is small, Q(S t ,A t ) is reduced. In other words, the value of a certain action in a certain state is brought closer to the value of the best action in the next state. However, the difference is determined by the discount rate γ and the reward r t+1 This varies depending on the state of affairs, but basically, the value of the best action in a certain state is propagated to the value of the action in the state immediately before it.
[0052] One method of Q-learning is to create a table of Q(S,A) for all state-action pairs (S,A) and then perform learning. However, there are too many states to calculate the Q(S,A) values for all state-action pairs, and it may take a long time for Q-learning to converge.
[0053] Therefore, a well-known technology called DQN (Deep Q-Network) may be used. Specifically, the value function Q may be configured using an appropriate neural network, and the value Q(S, A) may be calculated by approximating the value function Q with an appropriate neural network by adjusting the parameters of the neural network. By using DQN, it is possible to shorten the time required for Q-learning to converge. Note that DQN is described in detail in, for example, the following non-patent document:
[0054] <Non-patent literature> "Human-level control through deep reinforcement learning", Volodymyr Mnih1 [online], [searched on January 17, 2017], Internet <URL: http: / / files.davidqiu.com / research / nature14236.pdf>
[0055] The Q-learning described above is performed by the machine learning unit 400A. Specifically, the machine learning unit 400A calculates the integral gain K1v and proportional gain K2v of the speed control unit 120, and the coefficients ω of the transfer function of the filter 130. c , τ, δ, and the input / output gain (amplitude ratio) and phase delay output from the frequency characteristic calculation unit 300 are taken as state S. The integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficients ω of the transfer function of the filter 130 related to the state S are c , τ, δ is adjusted as an action A to learn the value Q.
[0056] The machine learning unit 400A calculates the integral gain K1v and proportional gain K2v of the speed control unit 120, and the coefficients ω of the transfer function of the filter 130. c , τ, and δ, the servo control unit 100 is driven using a speed command that is a sine wave with the aforementioned frequency changing, and state information S including input / output gain and phase delay for each frequency obtained from the frequency characteristic calculation unit 300 is observed, and an action A is determined. The machine learning unit 400A receives a reward every time it performs action A. The machine learning unit 400A searches for an optimal action A by trial and error, for example, so that the total reward over the future will be maximized. By doing so, the machine learning unit 400A calculates the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficients ω of the transfer function of the filter 130. c, τ, δ, the servo control unit 100 is driven using a speed command that is a sine wave whose frequency changes. The speed command is obtained from the frequency characteristic calculation unit 300, and an optimal action A (i.e., the integral gain K1v and proportional gain K2v of the speed control unit 120 and the optimal coefficient ω of the transfer function of the filter 130) is calculated for the state S including the input / output gain and phase delay for each frequency. c , τ, δ) can be selected.
[0057] That is, based on the value function Q learned by the machine learning unit 400A, the integral gain K1v and proportional gain K2v of the speed control unit 120 relating to a certain state S, and the coefficients ω of the transfer function of the filter 130 are calculated. c , τ, and δ, the behavior A that maximizes the value of Q is selected. This selects the behavior A that maximizes the value of Q (i.e., the integral gain K1v and proportional gain K2v of the velocity control unit 120, and / or the coefficients ω of the transfer function of the filter 130) of the servo control unit 100 that are generated by executing a program that generates a sinusoidal signal with a variable frequency. c , τ, δ) can be selected.
[0058] FIG. 10 is a block diagram showing the configuration of machine learning unit 400A. 10, in order to perform the above-described reinforcement learning, machine learning unit 400A includes state information acquisition unit 401, learning unit 402, behavior information output unit 403, value function storage unit 404, and optimized behavior information output unit 405. Learning unit 402 includes reward output unit 4021, value function update unit 4022, and behavior information generation unit 4023.
[0059] The state information acquisition unit 401 acquires the integral gain K1v and proportional gain K2v of the speed control unit 120, and the coefficients ω of the transfer function of the filter 130. c, τ, and δ, the state S including the input / output gain (amplitude ratio) and phase delay obtained by driving the servo control unit 100 using a speed command (sine wave) is acquired from the frequency characteristic calculation unit 300. This state information S corresponds to the environmental state S in Q-learning. In addition, the state information acquisition unit 401 acquires image data on a complex plane including a circle forming a closed curve and a unit circle from the closed curve drawing unit 502. The state information acquisition unit 401 outputs the acquired state information S and image data relating to a complex plane including the circle forming the closed curve and the unit circle to the learning unit 402.
[0060] Note that when Q learning is first started, the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficients ω of the transfer function of the filter 130 are c In this embodiment, the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficients ω of the transfer function of the filter 130 are generated by the user. c The initial setting values of τ and δ are adjusted to the optimum values by reinforcement learning. In addition, the integral gain K1v, proportional gain K2v, and coefficient ω c , τ, and δ may be machine-learned using the adjusted values as initial values if the operator has adjusted the machine tool in advance.
[0061] The learning unit 402 is a part that learns the value Q(S, A) when a certain action A is selected under a certain environmental state S.
[0062] First, the reward output unit 4021 of the learning unit 402 will be described. The reward output unit 4021 is a part that calculates a reward when an action A is selected under a certain state S.
[0063] The reward output unit 4021 acquires image data regarding the complex plane including a circle that forms a closed curve and the unit circle from the state information acquisition unit 401. The reward output unit 4021 creates a Nyquist locus by drawing on the complex plane where the open-loop frequency characteristic H(jω) is acquired by using the input / output gain and phase delay obtained from the state information acquisition unit 401. Since the method of creating the Nyquist locus has already been described in the operation explanation of the adjustment unit 400, it is omitted here. In this way, a complex plane showing the Nyquist locus, the unit circle, and the circle passing through the gain margin and phase margin as shown in FIG. 7 is obtained. The Nyquist locus in the initial state is obtained by driving the servo control unit 100 using a speed command (sine wave) based on the integral gain K1v, the proportional gain K2v, and the coefficients ω c , τ, and δ set by the user. The Nyquist locus in the process of Q learning is obtained by modifying the integral gain K1v, the proportional gain K2v, and / or the coefficients ω c , τ, and δ and driving the servo control unit 100 using a speed command (sine wave).
[0064] In the following description, the radius of the circle will be described as radius r, and the shortest distance between the circle and the Nyquist locus will be described as shortest distance d. Here, the shortest distance d is the shortest distance between the center of the circle and the Nyquist locus, but it is not limited to this. For example, it may be the shortest distance between the outer circumference of the circle and the Nyquist locus.
[0065] When the shortest distance d is smaller than the radius r (d < r) and the Nyquist locus passes inside the closed curve, the reward output unit 4021 gives a negative reward. On the other hand, when the shortest distance d is equal to or greater than the radius r (d ≧ r) and the Nyquist locus does not pass inside the circle, the reward output unit 4021 gives a zero reward.
[0066] By giving the reward as described above, the machine learning unit 400A tries to explore the integral gain K1v, the proportional gain K2v of the speed control unit 120, and the coefficients ω c , τ, and δ of the transfer function of the filter 130 so that the Nyquist locus does not pass inside the circle and the gain margin and phase margin become equal to or greater than the values set by the user.
[0067] In the example described above, whether or not the Nyquist locus passes inside the circle that is a closed curve is determined based on the shortest distance between the circle and the Nyquist locus. However, this method is not limiting and other methods may be used. For example, the determination may be made based on whether or not the Nyquist locus is tangent to or intersects with the periphery of the circle that is a closed curve.
[0068] (Example considering response speed) When the Nyquist locus passes on a circle (d=r) or outside the circle (d>r), the gain margin and phase margin increase as the Nyquist locus moves away from the circle, increasing the stability of the servo system, but decreasing the feedback gain and response speed. Therefore, it is desirable that the reward output unit 4021 give a reward so that the feedback gain is as large as possible, at or above the gain margin and phase margin determined by the user. Below, three examples of methods for the reward output unit 4021 to determine a reward so that the feedback gain is as large as possible, at or above the gain margin and phase margin determined by the user, will be described.
[0069] (1) A method for determining rewards based on cutoff frequency The reward output unit 4021 outputs the integral gain K1v, the proportional gain K2v, and the coefficient ω c , τ, and δ, a Bode diagram is created from the input / output gain (amplitude ratio) and phase delay of the closed loop obtained by driving the servo control unit 100 using a speed command (sine wave). Figure 11 shows an example of a Bode diagram of a closed loop. The cutoff frequency is, for example, the frequency at which the gain characteristic of the Bode diagram is −3 dB or the frequency at which the phase characteristic is −180 degrees. In Fig. 11, the frequency at which the gain characteristic is −3 dB is set as the cutoff frequency.
[0070] The reward output unit 4021 determines the reward so that the cutoff frequency becomes larger. Specifically, the reward output unit 4021 outputs the integral gain K1v and the proportional gain K2v, and / or the coefficient ω c, τ, δ are corrected, and the reward is determined based on whether the cutoff frequency fcut increases, remains the same, or decreases when the state changes from the state S before correction to state S'. In the following explanation, the cutoff frequency fcut when in state S is written as fcut(S), and the cutoff frequency fcut when in state S' is written as fcut(S').
[0071] When the state changes from S to S', if the cutoff frequency fcut increases, the reward output unit 4021 gives a positive reward, assuming that the cutoff frequency fcut(S') is greater than the cutoff frequency fcut(S). When the state changes from S to S', if the cutoff frequency fcut does not change, the reward output unit 4021 gives a reward of zero, assuming that the cutoff frequency fcut(S')=the cutoff frequency fcut(S). When the state changes from S to S' and the cutoff frequency fcut becomes smaller, the reward output unit 4021 determines that the cutoff frequency fcut(S') is smaller than the cutoff frequency fcut(S) and gives a negative reward.
[0072] By determining the reward as described above, when the Nyquist locus passes on or outside the circle, the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficient ω of the transfer function of the filter 130 are set so that the cutoff frequency fcut becomes large. c , τ, δ are searched for by trial and error. As the cutoff frequency fcut increases, the feedback gain increases and the response speed becomes faster.
[0073] (2) A method for determining rewards based on closed-loop characteristics The reward output unit 4021 outputs the integral gain K1v, the proportional gain K2v, and the coefficient ω c , τ, and δ, the closed-loop transfer function G(jω) is calculated from the input / output gain (amplitude ratio) and phase delay of the closed-loop obtained by driving the servo control unit 100 using a speed command (sine wave). The reward output unit 4021 calculates the evaluation function f in a preset frequency domain as f=Σ|1-G(jω)| 2can be applied. The reward output unit 4021 determines the reward so that the value of the evaluation function f becomes small. Specifically, the reward output unit 4021 outputs the integral gain K1v and the proportional gain K2v, and / or the coefficient ω c When τ, δ are corrected and the state changes from the state S before correction to state S', the reward is determined based on whether the value of the evaluation function f becomes smaller, the same, or larger. In the following explanation, the value of the evaluation function f when the state is S is written as f(S), and the value of the evaluation function f when the state is S' is written as f(S'). If the value of the evaluation function f becomes smaller, the cutoff frequency of the closed loop Bode diagram shown in FIG. 11 becomes larger.
[0074] When the state changes from S to S' and the value of the evaluation function f becomes smaller, the reward output unit 4021 gives a positive reward, assuming that the value of the evaluation function f(S')<the value of the evaluation function f(S). When the state changes from S to S', if the value of the evaluation function f does not change, the reward output unit 4021 gives a reward of zero, assuming that the value of the evaluation function f(S')=the value of the evaluation function f(S). When the state changes from S to S', if the value of the evaluation function f increases, the reward output unit 4021 gives a negative reward, assuming that the evaluation function value f(S')>the evaluation function value f(S).
[0075] By determining the reward as described above, when the Nyquist locus passes on or outside the circle, the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficient ω of the transfer function of the filter 130 are set so that the value of the evaluation function f becomes small. c , τ, δ are searched for by trial and error. As the value of the evaluation function f decreases, the feedback gain increases and the response speed becomes faster.
[0076] (3) A method for determining the reward so that the shortest distance d approaches the radius r When the Nyquist locus passes on the circle (d=r) or outside the circle (d>r), the reward is determined so that the Nyquist locus approaches a closed curve. Specifically, the reward output unit 4021 outputs the integral gain K1v and the proportional gain K2v, and / or the coefficient ω c When τ, δ are corrected and the state changes from the state S before correction to state S', the reward is determined based on whether the shortest distance d between the center of the circle and the Nyquist locus becomes smaller, the same, or larger. In the following explanation, the shortest distance d when in state S is referred to as d(s), and the shortest distance d when in state S' is referred to as d(s').
[0077] When the state changes from S to S' and the shortest distance d becomes small, the reward output unit 4021 determines that shortest distance d(S')<shortest distance d(S) and gives a positive reward. When the state changes from S to S', if the shortest distance d remains unchanged, the reward output unit 4021 gives a reward of zero, assuming that the shortest distance d(S')=the shortest distance d(S). When the state changes from S to S', if the shortest distance d becomes large, the reward output unit 4021 determines that shortest distance d(S')>shortest distance d(S) and gives a negative reward.
[0078] By determining the reward as described above, the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficient ω of the transfer function of the filter 130 can be adjusted so that the Nyquist locus passes through a circle or approaches the outer periphery of the circle. c , τ, δ are searched for by trial and error. As the Nyquist locus passes through a circle or approaches the periphery of the circle, the feedback gain increases and the response speed becomes faster. The method of determining the reward based on the information of the shortest distance d is not limited to the above method, and other methods can be applied.
[0079] (Example considering resonance) Even when the Nyquist locus passes on the circle (d=r) or outside the circle (d>r), the input / output gain may increase due to resonance at the machine end of the machine to be controlled. Therefore, it is desirable that the reward output unit 4021 determines a reward so as to suppress resonance at or above the gain margin and phase margin determined by the user. Below, a method for determining a reward by comparing the open loop characteristics with a reference model will be described.
[0080] Hereinafter, the operation of the reward output unit 4021 to give a negative reward when the input / output gain for each frequency in the created frequency characteristics is greater than the input / output gain of the reference model will be described with reference to FIGS.
[0081] The reward output unit 4021 stores a reference model of input / output gain. The reference model is a model of a servo control unit having ideal characteristics without resonance. The reference model is, for example, a model of the inertia Ja and torque constant K of the model shown in FIG. t , proportional gain K p , integral gain K I , differential gain K D The inertia Ja is the sum of the motor inertia and the mechanical inertia.
[0082] Fig. 13 is a characteristic diagram showing the frequency characteristics of the input / output gain of the servo control unit of the reference model and the servo control unit 100 before and after learning. As shown in the characteristic diagram of Fig. 13, the reference model has a region FA, which is a frequency region where the ideal input / output gain is equal to or greater than a certain input / output gain, for example, -20 dB or greater, and a region FB, which is a frequency region where the input / output gain is less than the certain input / output gain. In the region FA of Fig. 13, the ideal input / output gain of the reference model is shown by a curve MC1 (thick line). In the region FB of Fig. 13, the ideal virtual input / output gain of the reference model is shown by a curve MC 11 The input / output gain of the reference model is constant and the line MC 12 In the areas FA and FB in Fig. 13, the curves of the input / output gains with the servo control unit before and after learning are shown as curves RC1 and RC2, respectively.
[0083] In the region FA, when the curve RC1 of the input-output gain for each frequency in the created frequency characteristics exceeds the curve MC1 of the ideal input-output gain of the canonical model, the reward output unit 4021 gives a negative reward. In the region FB beyond the frequency where the input-output gain becomes sufficiently small, even if the curve RC1 of the input-output gain before learning exceeds the curve MC 11 of the ideal virtual input-output gain of the canonical model, the impact on stability becomes small. Therefore, in the region FB, as described above, the input-output gain of the canonical model is not the curve MC 11 of the ideal gain characteristic, but a straight line MC 12 of a constant input-output gain (for example, -20 dB). However, if the measured curve RC1 of the input-output gain before learning exceeds the straight line MC 12 of the constant input-output gain, there is a possibility of instability, so a negative value is given as a reward.
[0084] When adjusting the gain of the input-output gain, the integral gain K1v and proportional gain K2v of the speed control unit 120, and / or the coefficients ω c , τ, δ of the transfer function of the filter 130 are adjusted. The characteristics of the filter 130 change in gain and phase depending on the bandwidth fw of the filter 130, and also change in gain and phase depending on the attenuation coefficient k of the filter 130. Therefore, the gain of the input-output gain can be adjusted by adjusting the coefficients of the filter 130.
[0085] When the shortest distance d is smaller than the radius r (d < r) and the Nyquist locus passes inside the closed curve, and a negative reward is given, the reward output unit 4021 outputs this negative reward to the value function update unit 4022. When the shortest distance d is equal to or greater than the radius r (d ≧ r) and the Nyquist locus does not pass inside the circle, and a positive reward is given, the reward output unit 4021 outputs this positive reward to the value function update unit 4022. When the reward output unit 4021 gives a reward in three examples considering the response speed or an example considering resonance, the reward output unit 4021 outputs the total reward obtained by adding the positive reward given when the Nyquist locus does not pass inside the circle to this reward to the value function update unit 4022.
[0086] When adding rewards, weights may be applied to the rewards. For example, when emphasis is placed on the stability of the servo system, a positive reward given when the Nyquist locus does not pass through the inside of the circle can be weighted to be more important than the rewards given in the three examples taking response speed into consideration or the example taking resonance into consideration. The reward output unit 4021 has been described above.
[0087] The value function update unit 4022 updates the value function Q stored in the value function memory unit 404 by performing Q-learning based on the state S, the action A, the state S' when the action A is applied to the state S, and the reward calculated as described above. The value function Q may be updated by online learning, batch learning, or mini-batch learning. Online learning is a learning method in which a certain action A is applied to the current state S, and the value function Q is updated immediately each time the state S transitions to a new state S'. Batch learning is a learning method in which a certain action A is applied to the current state S, and the state S transitions to a new state S', thereby collecting learning data and updating the value function Q using all of the collected learning data. Mini-batch learning is a learning method that is intermediate between online learning and batch learning, and updates the value function Q each time a certain amount of learning data is accumulated.
[0088] The behavior information generation unit 4023 selects behavior A in the Q-learning process for the current state S. In the Q-learning process, the behavior information generation unit 4023 selects the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or each coefficient ω of the transfer function of the filter 130. c , τ, and δ (corresponding to action A in Q-learning), the behavior information A is generated and output to the behavior information output unit 403. More specifically, the behavior information generating unit 4023 generates, for example, the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficients ω of the transfer function of the filter 130 included in the state S. c , τ, δ, the integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficients ω of the transfer function of the filter 130 included in the action A c , τ, δ may be added or subtracted incrementally.
[0089] The integral gain K1v and proportional gain K2v of the speed control unit 120 and the coefficients ω of the filter 130 are c , τ, δ may all be modified, or some of the coefficients may be modified. c When correcting τ, δ, for example, it is easy to find the center frequency fc that causes resonance, and it is easy to specify the center frequency fc. Therefore, the behavior information generating unit 4023 temporarily fixes the center frequency fc, corrects the bandwidth fw and the attenuation coefficient δ, that is, corrects the coefficient ω c In order to fix (=2πfc) and perform an operation to modify the coefficient τ (=fw / fc) and the attenuation coefficient δ, behavior information A may be generated and output to the behavior information output unit 403.
[0090] In addition, the behavior information generation unit 4023 may take a strategy of selecting behavior A' using a known method, such as a greedy method that selects behavior A' with the highest value Q(S, A) from the currently estimated values of behavior A, or an ε-greedy method that randomly selects behavior A' with a certain small probability ε and otherwise selects behavior A' with the highest value Q(S, A).
[0091] The behavior information output unit 403 is a part that transmits the behavior information A output from the learning unit 402 to the speed control unit 120 and the filter 130. As described above, the filter 130 calculates the current state S, that is, the currently set integral gain K1v and proportional gain K2v of the speed control unit 120, and / or each coefficient ω c, τ, and δ, a transition to the next state S′ (that is, the corrected integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficients of the filter 130) occurs.
[0092] Value function storage unit 404 is a storage device that stores value function Q. Value function Q may be stored as a table (hereinafter referred to as an action value table) for each state S and action A, for example. Value function Q stored in value function storage unit 404 is updated by value function update unit 4022. Furthermore, value function Q stored in value function storage unit 404 may be shared with other machine learning units 400A. If value function Q is shared by multiple machine learning units 400A, reinforcement learning can be performed in a distributed manner among the machine learning units 400A, thereby improving the efficiency of reinforcement learning.
[0093] The optimization behavior information output unit 405 generates behavior information A (hereinafter referred to as "optimization behavior information") for causing the speed control unit 120 and the filter 130 to perform an operation that maximizes the value Q(S, A) based on the value function Q updated by the value function update unit 4022 through Q-learning. More specifically, the optimization behavior information output unit 405 acquires the value function Q stored in the value function storage unit 404. This value function Q is updated by the value function update unit 4022 performing Q-learning as described above. Then, the optimization behavior information output unit 405 generates behavior information based on the value function Q and outputs the generated behavior information to the filter 130. This optimization behavior information includes, like the behavior information output by the behavior information output unit 403 in the Q-learning process, the integral gain K1v and proportional gain K2v of the speed control unit 120, and / or each coefficient ω of the transfer function of the filter 130. c , τ, δ.
[0094] In the speed control section 120, the integral gain K1v and the proportional gain K2v are corrected based on this behavioral information, and in the filter 130, each coefficient ω of the transfer function is corrected based on this behavioral information. c , τ, δ are modified. Through the above operation, the machine learning unit 400A calculates the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficients ω of the transfer function of the filter 130. c , τ, and δ are optimized, and the servo control unit 100 can be operated so that the stability margin is equal to or greater than a predetermined value. Furthermore, by the above operation, the integral gain K1v and proportional gain K2v of the speed control unit 120 and / or the coefficients ω of the transfer function of the filter 130 c , τ, and δ are optimized so that the stability margin of the servo control unit 100 is equal to or greater than a predetermined value, and the feedback gain is increased to increase the response speed and / or suppress resonance. As described above, by using the machine learning unit 400A of the present disclosure, it is possible to simplify the adjustment of the gain of the speed control unit 120 and the parameters of the filter 130.
[0095] The control systems 10 and 10A, and the functional blocks included in the setting support unit 500 and machine learning unit 400A have been described above. To realize these functional blocks, the control system 10 or 10A, or the setting support unit 500 or the machine learning unit 400A includes a processing unit such as a CPU (Central Processing Unit). The control systems 10 and 10A also include a secondary storage device such as an HDD (Hard Disk Drive) that stores various control programs such as application software or an OS (Operating System), and a main storage device such as a RAM (Random Access Memory) that stores data temporarily required for the processing unit to execute a program.
[0096] In the control system 10 or 10A, or the setting support unit 500 or machine learning unit 400A, the arithmetic processing unit reads application software or an OS from the auxiliary storage device, and executes arithmetic processing based on the application software or OS while loading the loaded application software or OS into the main storage device. Furthermore, based on the results of this calculation, various pieces of hardware provided in each device are controlled. This realizes the functional blocks of this embodiment. In other words, this embodiment can be realized by the cooperation of hardware and software.
[0097] Since the amount of calculations involved in machine learning for machine learning unit 400A is large, high-speed processing can be achieved by, for example, equipping a personal computer with a GPU (Graphics Processing Unit) and using the GPU for the calculations involved in machine learning using a technology called GPGPU (General-Purpose Computing on Graphics Processing Units).Furthermore, in order to achieve even faster processing, a computer cluster can be constructed using multiple computers equipped with such GPUs, and parallel processing can be performed on the multiple computers included in this computer cluster.
[0098] Each component included in the control system 10 or 10A can be realized by hardware, software, or a combination thereof. Furthermore, the setting support method performed by the cooperation of each component included in the setting support device can also be realized by hardware, software, or a combination thereof. Here, "realized by software" means that the method is realized by a computer reading and executing a program.
[0099] The program can be stored and supplied to a computer using various types of non-transitory computer readable media. Non-transitory computer readable media include various types of tangible storage media. Examples of non-transitory computer readable media include magnetic recording media (e.g., hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer readable media.
[0100] The above-described embodiment is a preferred embodiment of the present invention, but the scope of the present invention is not limited to the above-described embodiment alone, and the present invention can be implemented in various modified forms within the scope that does not deviate from the gist of the present invention.
[0101] In the above-described embodiment, the case where one filter is provided has been described, but the filter 130 may be configured by connecting in series a plurality of filters each corresponding to a different frequency band. FIG. 14 is a block diagram showing an example of a filter configured by directly connecting a plurality of filters. In FIG. 14, when there are m resonance points (m is a natural number of 2 or more), the filter 130 is configured by connecting m filters 130-1 to 130-m in series. The coefficients ω of the m filters 130-1 to 130-m are c The optimal values for τ and δ are found using machine learning.
[0102] In addition to the configurations shown in FIGS. 1 and 9, the control system may have the following configurations. In the following description, a modified example of the configuration of FIG. 1 will be described, but the configuration of FIG. 9 can also be modified in a similar manner. <Modification in which the machine learning unit is provided outside the servo control unit> Fig. 15 is a block diagram showing another example of the configuration of a control device. Control device 10B shown in Fig. 15 differs from control system 10 shown in Fig. 1 in that n (n is a natural number of 2 or more) servo control units 100-1 to 100-n are connected to n adjustment units 400-1 to 400-n via network 600, and each includes a frequency generation unit 200 and a frequency characteristic calculation unit 300. The n adjustment units 400-1 to 400-n are connected to n setting support units 500-1 to 500-n. The setting support units 500-1 to 500-n have the same configuration as the setting support unit 500 shown in Fig. 2. The servo control units 100-1 to 100-n correspond to servo control devices, and the setting support units 500-1 to 500-n correspond to setting support devices, respectively. Of course, one or both of the frequency generation unit 200 and the frequency characteristic calculation unit 300 may be provided outside the servo control units 100-1 to 100-n.
[0103] Here, servo control unit 100-1 and adjustment unit 400-1 are paired one-to-one and connected to be able to communicate. Servo control units 100-2 to 100-n and adjustment units 400-2 to 400-n are connected in the same manner as servo control unit 100-1 and adjustment unit 400-1. In FIG. 15, n pairs of servo control units 100-1 to 100-n and adjustment units 400-1 to 400-n are connected via network 600, but the servo control unit and machine learning unit of each of the n pairs of servo control units 100-1 to 100-n and adjustment units 400-1 to 400-n may be directly connected via a connection interface. These n pairs of servo control units 100-1 to 100-n and adjustment units 400-1 to 400-n may be installed in the same factory, for example, or may be installed in different factories.
[0104] The network 600 may be, for example, a local area network (LAN) established within a factory, the Internet, a public telephone network, or a combination of these. There are no particular limitations on the specific communication method of the network 600, whether it is a wired connection or a wireless connection, etc.
[0105] <Flexibility of system configuration> In the above-described embodiment, the servo control units 100-1 to 100-n and the adjustment units 400-1 to 400-n are connected in one-to-one pairs so as to be able to communicate with each other, but for example, one adjustment unit may be connected to multiple servo control units so as to be able to communicate with each other via the network 600. In this case, the functions of one adjustment unit may be distributed to multiple servers as needed to form a distributed processing system.Furthermore, the functions of one adjustment unit may be realized using a virtual server function on the cloud.
[0106] Furthermore, when there are n servo control units 100-1 to 100-n of the same model name, the same specifications, or the same series and corresponding n adjustment units 400-1 to 400-n, respectively, the adjustment units 400-1 to 400-n may be configured to share the learning results of the adjustment units 400-1 to 400-n. This makes it possible to build a more optimal model.
[0107] The setting support device, control system, and setting support method according to the present disclosure can take on a variety of embodiments having the following configurations, including the above-described embodiments. (1) A setting support device (for example, a setting support unit 500) that supports a user in setting a stability margin of a servo control device (for example, a servo control unit 100), a closed curve drawing unit (for example, the closed curve drawing unit 502) that draws a closed curve on the complex plane that includes (-1, 0) on the complex plane and passes through the gain margin and the phase margin; a change unit (for example, a stability margin change unit 504) that changes the gain margin and the phase margin based on a user operation; A setting support device comprising: a closed curve scaling unit (e.g., a closed curve scaling unit 505) that expands or reduces the closed curve in conjunction with the amount of change in the gain margin and the phase margin by the change unit. According to this setting support device, when a user sets a stability margin of a servo control device, the user can easily set and change the stability margin.
[0108] (2) The setting assistance device according to (1) above, further comprising a reference stability margin setting unit that sets a reference stability margin that is the gain margin and the phase margin and outputs the set reference stability margin to the closed curve drawing unit.
[0109] (3) The setting support device according to (1) or (2), wherein the change unit is a slide bar.
[0110] (4) The setting support device according to any one of (1) to (3) above, wherein the closed curve is a circle.
[0111] (5) The machine learning device according to any one of (1) to (4) above, wherein the closed curve scaling unit scales or reduces the closed curve based on a point on the real axis.
[0112] (6) The setting support device according to any one of (1) to (5) above, wherein the closed curve drawing unit draws a plurality of closed curves. According to this setting support device, a closed curve suited to each mode can be used for a plurality of modes of the servo control device, for example, cutting feed and rapid feed.
[0113] (7) A setting support device according to any one of (1) to (4) above; A control system provided in the servo control device, the control system including an adjustment unit for adjusting at least one filter coefficient and feedback gain. According to this control system, when a user sets a stability margin for a servo control device, the user can easily set and change the stability margin.
[0114] (8) The adjustment unit a state information acquiring unit that acquires state information including at least one of the coefficient of the filter and the feedback gain, and an input / output gain and an input / output phase delay of the servo control device; a behavior information output unit that outputs behavior information including adjustment information for at least one of the coefficient and the feedback gain included in the state information; a reward output unit that calculates and outputs a reward based on whether a Nyquist locus calculated from the input / output gain and the input / output phase delay passes inside the closed curve output from the setting assistance device; and a value function update unit that updates a value function based on the reward value output by the reward output unit, the state information, and the behavior information; The control system according to (7) above, wherein the machine learning unit is provided with:
[0115] (9) The computer A process of setting a first gain margin and a first phase margin, which are reference stability margins; A process of drawing a closed curve on the complex plane that includes (-1, 0) on the complex plane and passes through the first gain margin and the first phase margin; a process of changing the first gain margin and the first phase margin to a second gain margin and a second phase margin based on a user operation; expanding or contracting the closed curve in conjunction with an amount of change from the first gain margin and the first phase margin to the second gain margin and the second phase margin; A configuration assistance method for performing the above. According to this setting support method, when a user sets a stability margin of a servo control device, the user can easily set and change the stability margin. [Explanation of symbols]
[0116] 10, 10A, 10B control system 100, 100-1 to 100-n Servo control unit 110 Subtractor 120 Speed control section 130 filters 140 Current control section 150 motor 200 Frequency generation unit 300 Frequency characteristic calculation section 400, 400-1~400-n adjustment section 400A Machine Learning Department 401 Status information acquisition unit 402 Learning Department 403 Behavioral Information Output Unit 404 Value Function Memory Unit 405 Optimization behavior information output unit 500, 500-1~500-n Setting support department 600 Network
Claims
1. A setting support device that supports a user in setting a stability margin of a servo control device, a closed curve drawing unit that draws a closed curve on a complex plane that includes (-1, 0) on the complex plane and passes through a gain margin and a phase margin; a change unit that changes the gain margin and the phase margin based on a user operation; a closed curve scaling unit that scales or shrinks the closed curve in conjunction with the amount of change in the gain margin and the phase margin made by the change unit.
2. 2. The setting assistance device according to claim 1, further comprising a reference stability margin setting unit that sets reference stability margins that are the gain margin and the phase margin and outputs the set reference stability margins to the closed curve drawing unit.
3. The setting support device according to claim 1 , wherein the change unit is a slide bar.
4. The setting support device according to claim 1 , wherein the closed curve is a circle.
5. The setting assistance device according to claim 1 , wherein the closed curve enlarging / reducing unit enlarges or reduces the closed curve with respect to a point on a real axis.
6. The setting support device according to claim 1 , wherein the closed curve drawing unit draws a plurality of closed curves.
7. The setting support device according to any one of claims 1 to 4, A control system including an adjustment unit for adjusting a coefficient and a feedback gain of at least one filter, the adjustment unit being provided in the servo control device.
8. The adjustment unit a state information acquiring unit that acquires state information including at least one of the coefficient of the filter and the feedback gain, and an input / output gain and an input / output phase delay of the servo control device; a behavior information output unit that outputs behavior information including adjustment information for at least one of the coefficient and the feedback gain included in the state information; a reward output unit that calculates and outputs a reward based on whether a Nyquist locus calculated from the input / output gain and the input / output phase delay passes inside the closed curve output from the setting assistance device; a value function update unit that updates a value function based on the reward value output by the reward output unit, the state information, and the behavior information; The control system according to claim 7, wherein the machine learning unit comprises:
9. The computer A process of setting a first gain margin and a first phase margin, which are reference stability margins; A process of drawing a closed curve on a complex plane that includes (−1, 0) on the complex plane and passes through the first gain margin and the first phase margin; a process of changing the first gain margin and the first phase margin to a second gain margin and a second phase margin based on a user operation; expanding or contracting the closed curve in conjunction with an amount of change from the first gain margin and the first phase margin to the second gain margin and the second phase margin; A configuration assistance method for performing the above.
Citation Information
Patent Citations
Setting support apparatus, setting support method, information processing program, and record medium
JP2017167607A
Control system design support device, control system design support method, and control system design support program
JP2019168777A
Design support device, design support method, and design support program
JP2020120573A
Design assistance device, design assistance method, and design assistance program
WO2020153488A1