Output device, control device and method for outputting a learning parameter
The output device converts machine-learned parameters into user-friendly information, allowing users to understand and optimize servo control devices by displaying physical quantities and frequency responses.
Patent Information
- Application Number
- DE102019216081
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-25
- Filing Date
- 2019-10-18
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2039-10-18
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
BACKGROUND OF THE INVENTION Area of the invention
[0001] The present invention relates to an output device, a control device and a method for outputting learning parameters, and in particular to an output device which receives parameters (called learning parameters) from a machine learning device that performs machine learning on a servo control device for controlling a servo motor, parameters that are machine learned or have been machine learned, and which outputs information from the learning parameters, which is easily understood by a user, such as an operator, to a control device which includes an output device, and to a method for outputting the learning parameters. State of the art
[0002] As a technique relating to the present invention, patent document 1 discloses, for example, a signal converter comprising an output unit which uses a multiplication coefficient pattern mastering method with machine learning means to determine a predetermined multiplication coefficient pattern, which uses the multiplication coefficient pattern to perform a digital filter operation, and which displays an output of a digital filter.
[0003] In particular, patent document 1 discloses that the signal converter includes a signal input unit, an operation processing unit which has the function of characterizing the signal data based on the input signal data, and the output unit which outputs a result from the operation processing unit, that the operation processing unit includes an input file, learning means, a digital filter and parameter setting means, and that the learning means uses the multiplication coefficient pattern mastering method with machine learning means to determine the intended multiplication coefficient pattern.
[0004] Patent document 1: Japanese unexamined patent application, publication no. JP H11-31139A WO 2018 / 151 215 A1 DE 10 2018 203 702 A1 DE 10 2018 205 015 A1 US 2017 / 0 262 573 A1 SUMMARY OF THE INVENTION
[0005] Although patent document 1 displays the output from the operational processing unit, it disadvantageously does not output a pattern learned by machine learning, thus preventing a user, such as an operator, from verifying the progress or outcome of the machine learning. Similarly, when the control parameters of a servo control device that controls a servo motor driving the axis of a machine tool, robot, or industrial machine are machine learned using a machine learning device, the user cannot verify the progress or outcome of the machine learning because the learning parameters and an evaluation function value used in the machine learning device are generally not displayed.Even when the learning parameters or the evaluation function value are displayed, the user has difficulty understanding from the learning parameters how the characteristic curve of the servo control device is optimized.
[0006] It is an object of the present invention to provide an output device that acquires the learning parameters and outputs from the learning parameters information that is easily understood by a user, such as an operator, a control device that includes such an output device, and a method for outputting the learning parameters.
[0007] The problem is solved by the subject matter of the independent patent claims.
[0008] (1) An output device (e.g., an output device 200, 200A, 210, which will be described later) according to the present invention comprises: an information acquisition unit (e.g., an information acquisition unit 211, which will be described later) that receives a parameter or a first physical quantity of a constituent element of the servo control device, which is being learned or has been learned, from a machine learning device (e.g., a machine learning device 200, 210, which will be described later) that performs machine learning on a servo control device (e.g., a servo control device 300, 310, which will be described later) for controlling a servo motor (e.g., a servo motor 400, 410, which will be described later) that drives the axis of a machine tool, robot, or industrial machine; and an output unit (e.g.,a control unit 215 and a display unit 219, a control unit 215 and a storage unit 216, which are described later), which outputs at least one of any of the detected first physical quantity and a second physical quantity determined from the detected parameter, a characteristic curve of the time response of the constituent element of the servo control device and a characteristic curve of the frequency response of the constituent element of the servo control device, wherein the characteristic curve of the time response and the characteristic curve of the frequency response are determined with the parameter, the first physical quantity or the second physical quantity.
[0009] (2) In the output device described above according to (1), the output unit may include a display unit which displays on a display screen the first physical quantity, the second physical quantity, the time response characteristic or the frequency response characteristic.
[0010] (3) In the output device described above according to (1) or (2), an instruction to set the parameter or the first physical quantity of the constituent element of the servo control device based on the first physical quantity, the second physical quantity, the time response characteristic or the frequency response characteristic can be provided to the servo control device.
[0011] (4) In the output device described above according to one of (1) to (3), a machine learning instruction can be provided to the machine learning device to perform machine learning of the parameter or first physical quantity of the constituent element of the servo control device based on the first physical quantity, the second physical quantity, the time response characteristic or the frequency response characteristic by changing or selecting a learning area.
[0012] (5) In the output device described above according to one of (1) to (4), an evaluation function value, which is used in the machine learning of the device, can be output.
[0013] (6) In the output device described above according to one of (1) to (5), the information about a position error that can be output by the servo control device is output.
[0014] (7) In the output device described above according to one of (1) to (6), the parameter of the constituent element of the servo control device may be a parameter of a model of a mathematical formula or a filter.
[0015] (8) In the output device described above according to one of (1) to (7) the model of a mathematical formula or the filter may be included in a velocity feedforward processing unit or a position feedforward processing unit and the parameter may include a coefficient in a transfer function of the filter.
[0016] (9) A control device according to the present invention comprises: the output device described above according to one of (1) to (8); the servo control device which controls the servo motor which drives the axis of the machine tool, robot or industrial machine; and the machine learning device which performs the machine learning on the servo control device.
[0017] (10) In the control device described above according to (9), the output device may be included in one of the servo control devices, the machine learning device and the numerical control device.
[0018] (11) A method for outputting a learning parameter of an output device according to the present invention is a method for outputting a parameter that is machine-learned in a machine learning device for a servo control device that controls a servo motor for driving the axis of a machine tool, robot or industrial machine, wherein the method comprises: acquiring from the machine learning device a parameter or a first physical quantity of a constituent element of the servo control device that is being learned or has been learned;Output at least one of the measured first physical quantities and a second physical quantity determined from the measured parameter, a characteristic curve of the time response of the constituent element of the servo control device, and a characteristic curve of the frequency response of the constituent element of the servo control device; and determine the characteristic curve of the time response and the characteristic curve of the frequency response with the parameter, the first physical quantity, or the second physical quantity.
[0019] According to the present invention, the parameters that are machine-learned or have been machine-learned are captured and changed into information that is easily understood by a user, such as an operator, and the information can be output. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a block diagram showing an example of the configuration of a control device according to a first embodiment of the present invention; Fig. Figure 2 is a block diagram showing the overall configuration of a control device and the configuration of a servo control device in a first example; Fig. Figure 3 is a graphical representation showing a velocity command, which serves as an input signal, and a detection velocity, which serves as an output signal; Fig. Figure 4 is a graphical representation showing the frequency response of an amplitude ratio between the input signal and the output signal and a phase delay; Fig. 5 is a block diagram showing a machine learning device according to the first embodiment of the present invention; Fig. Figure 6 is a graphical representation showing a reference model of a servo control device that has an ideal characteristic curve without resonance; Fig. Figure 7 is a characteristic map showing the frequency response of the input / output gains of the servo control device of the reference model and the servo control device before and after learning; Fig. Figure 8 is a block diagram showing an example of the configuration of an output device included in the control device according to the first example of the present invention; Fig. 9A is a characteristic map showing a machine-learned evaluation function value and the progress of the minimum value of the evaluation function value, and a graphical representation showing an example of a display screen when the values of the control parameters being learned are displayed; Fig. Figure 9B is a graphical representation showing an example of a display screen when the physical quantities of the control parameters relating to a state S are displayed in a display unit while the machine learning is performed to correspond to the progress of the machine learning; Fig. Figure 10 is a flowchart showing the operation of the control device from the start of the machine learning until the completion of the machine learning, while the attention is focused on the output device in the first example of the present invention; Fig. Figure 11 is a block diagram showing the overall configuration of a control device and the configuration of a servo control device according to a second example of the present invention; Fig. Figure 12 is a graphical representation showing a case in which a machined shape specified by a learning processing program is an octagon; Fig. Figure 13 is a graphic representation showing a case in which the processed shape is a shape in which the corners of an octagon are alternately replaced by arcs; Fig. Figure 14 is a graphical representation showing a complex plane that indicates the search area of a pole and a zero point; Fig. Figure 15 is a graphical representation showing a frequency response characteristic of a velocity feedforward processing unit and a characteristic of a position error; Fig. Figure 16 is a flowchart showing the operation of an output device after an instruction to complete the machine learning in the second example of the present invention; Fig. Figure 17 is a graphical representation showing a frequency response characteristic of the velocity feedforward processing unit and a characteristic of the position error when a center frequency is changed; Fig. Figure 18 is a graphical representation showing a case in which the velocity feedforward processing unit is designed with a motor blocking characteristic, a notch filter and a low-pass filter; Fig. Figure 19 is a block diagram showing an example of the configuration of a control device according to a second embodiment of the present invention; Fig. Figure 20 is a block diagram showing an example of the configuration of a control device according to a third embodiment of the present invention; and Fig. Figure 21 is a block diagram showing a control device that has a different configuration. DETAILED DESCRIPTION OF THE INVENTION
[0020] The embodiments of the present invention are described in detail below with reference to the drawings. (First embodiment)
[0021] Fig. Figure 1 is a block diagram showing an example of the configuration of a control device according to a first embodiment of the present invention. The diagram shown in Fig. The control device 10 shown in Figure 1 includes a machine learning device 100, an output device 200, a servo control device 300, and a servo motor 400. The machine learning device 100 receives control commands from the output device 200, such as a position command and a speed command, which are input to the servo control device 300, and servo information, such as a position error, which is output by the servo control device 300, or information used in machine learning, such as information (e.g., an input / output gain and a phase delay) obtained from the control commands and the servo information. Fig. Figure 1 shows an example in which the machine learning device 100 acquires the control commands and servo information. The machine learning device 100 also acquires the parameters of a mathematical formula model or the parameters of a filter output from the servo control device 300, which are then transmitted to the output device 200. Based on the input information, the machine learning device 100 learns the parameters of the mathematical formula model or filter from the servo control device 300 in order to output these learned parameters to the output device 200. Examples of such learned parameters include the coefficients of a notch filter in the servo control device 300 or the coefficients of a velocity feedforward processing unit.Although in the description of the present embodiment the machine learning device 100 performs reinforcement learning, the learning performed by the machine learning device 100 is not specifically limited to reinforcement learning, and the present invention can also be applied, for example, to a case in which supervised learning is performed.
[0022] The output device 200 acquires the learning parameters of the model of a mathematical formula or filter, which are being or have been machine-learned in the device 100 for machine learning, and outputs information from the learning parameters that specifies a physical quantity, a time response, or a frequency response, which is easily understood by a user, such as an operator. Examples of the output method include a screen display in a liquid crystal display, printing on paper using a printer or the like, storage in a memory unit, and outputting an external signal through a communication unit.
[0023] If the learning parameters of the model are, for example, the coefficient of the notch filter or the coefficient of the velocity feedforward processing unit, it is difficult to grasp the characteristic curve of the notch filter or the velocity feedforward processing unit, and it is also difficult to grasp how the characteristic curve is optimized by learning with the machine learning device, even if the operator can see it. If the machine learning device 100 performs reinforcement learning, it is difficult to grasp how the parameters are optimized using only the evaluation function value, although an evaluation function value can be output to the output device 200 to provide compensation. Consequently, the output device 200 outputs information that is easily understood by the user and that includes the physical quantities of the parameters (e.g.,Specify a center frequency, a bandwidth fw, and a damping coefficient (damping), or the time response or frequency response of the model of a mathematical formula or filter. As described above, the output device 200 provides information that is easily understood by the user, so that the operator can easily understand the progress and result of the machine learning.If the learning parameter itself, output by the machine learning device 110, is a physical quantity easily understood by the user, the output device outputs the information about it. However, if the learning parameter is information not easily understood by the user, the output device converts the information into a physical quantity, time response, or frequency response of the model, a mathematical formula, or a filter that is easily understood by the user, and outputs it. The physical quantity is any of, for example, inertia, mass, viscosity, stiffness, a resonance frequency, a damping center frequency, a damping rate, a damping frequency range, a time constant, a cutoff frequency, or a combination thereof.
[0024] The output device 200 also functions as a setting device, which performs the forwarding of information (such as control commands, control parameters and servo information) between the machine learning device 100 and the servo control device 300 and the control of an operation between the machine learning device 100 and the servo control device 300.
[0025] The servo control device 300 outputs a current command based on control commands, such as the position command and the speed command, to control the rotation of the servo motor 400. The servo control device 300 contains the speed feedforward processing unit, which is represented, for example, by the notch filter or a model of a mathematical formula. The servo motor 400 is contained, for example, in a machine tool, a robot, or an industrial machine. The control device 10 can be contained in a machine tool, a robot, an industrial machine, or the like. The servo motor 400 outputs a detection position and / or a detection speed as feedback information to the servo control device 300.
[0026] The following describes a specific configuration of the control device of the first embodiment based on the first to fourth examples. <Erstes Beispiel>
[0027] The present example is an example where a machine learning device 110 learns the coefficient of a filter contained in a servo control device 310 and where an output device 210 displays the progress of the frequency response of the filter in a display unit. Fig. Figure 2 is a block diagram showing the overall configuration of a control device and the configuration of the servo control device in the first example. The control device 11 includes the machine learning device 110, the output device 210, the servo control device 310, and a servo motor 410. The machine learning device 110, the output device 210, the servo control device 310, and the servo motor 410, which are in Fig. Figure 2 shows the machine learning device 100, the output device 200, the servo control device 300 and the servo motor 400 in Fig. 1. One or all of the machine learning device 110 and the output device 210 may be provided within the servo control device 310.
[0028] The servo control device 310 comprises as its constituent elements a subtractor 311, a speed control unit 312, a filter 313, a current control unit 314, and a measuring unit 315. The measuring unit 315 can be located outside the servo control device 310. The subtractor 311, the speed control unit 312, the filter 313, the current control unit 314, and the servo motor 410 configure a speed control loop.
[0029] The subtractor 311 determines the difference between an input speed command and a detected speed subjected to speed feedback, and outputs the difference as a speed error to the speed control unit 312. A sine wave signal with a variable frequency is input as the speed command to the subtractor 311 and the measuring unit 315. Although the variable frequency sine wave signal is input from a higher-level device, the servo control device 310 may include a frequency generation unit that produces the variable frequency sine wave signal.The speed control unit 312 adds a value obtained by multiplying the speed error by an integral gain K1v and integrating the resulting value, and a value obtained by multiplying the speed error by a proportional gain K2v, and outputs the obtained value as a torque command to the filter 313.
[0030] Filter 313 is a filter that attenuates a specific frequency component, for example, using a notch filter. In a machine, such as a machine tool driven by a motor, a resonance point exists, and this resonance can be amplified in the servo control device 310. The notch filter is used, thereby reducing the resonance. An output from filter 313 is sent as the torque command to the current control unit 314. Expression 1 (hereafter represented mathematically as 1) specifies a transfer function G(s) of filter 313. The control parameters are the coefficients a0, a1, a2, b0, b1, and b2. If the filter is the notch filter, then b0 = a0 and b2 = a2 = 1. G(s)=b0+b1s+bss2a0+a1s+a2s2
[0031] In the following description, it is assumed that the filter is the notch filter, that in expression 1 b0 = a0 and b2 = a2 = 1, and that the machine learning device 110 learns the coefficients a0, a1, and b1 by machine.
[0032] The current control unit 314 generates the current command to drive the servomotor 410 based on the torque command and outputs the current command to the servomotor 410. The angular position of the servomotor 410 is detected by a rotary encoder (not shown) provided in the servomotor 410, with a velocity detection value being input to the subtractor 311 as velocity feedback.
[0033] The sine wave signal, whose frequency is changed, is input into the measuring unit 315 as the speed command. The measuring unit 315 uses the speed command (the sine wave), which serves as an input signal, and the detection speed (sine wave), which serves as an output signal, from the (not shown) rotary encoder, to determine, for each frequency specified by the speed command, an amplitude ratio (an input / output gain) between the input signal and the output signal, and a phase delay. Fig. Figure 3 is a graphical representation showing the velocity command, which serves as the input signal, and the detection velocity, which serves as the output signal. Fig. Figure 4 is a graphical representation showing the frequency response of the amplitude ratio between the input signal and the output signal and the phase delay. Although the servo control device 310 is configured as described above, the control device 11 further includes the machine learning device 110 and the output device 210 to machine learn the optimal parameters for the filter in order to output the frequency response of the parameters.
[0034] The machine learning device 110 uses the input / output gain (amplitude ratio) and phase delay output by the output device 210 to machine learn the coefficients a0, a1, and b1 in the transfer function of the filter 313 (hereinafter referred to as learning). Although the learning by the machine learning device 110 is performed prior to delivery, relearning can be performed after delivery.
[0035] The configuration and details of the operation of the machine learning device 110 are described below. Before describing the individual functional blocks contained in the machine learning device 110, the basic mechanism of reinforcement learning is described first. An agent (corresponding to the machine learning device 110 in the present embodiment) observes the state of an environment and selects a specific action, consequently modifying the environment based on the action. When the environment is modified, some form of reward is provided, thereby enabling the agent to learn to select (decide on) a better action. While supervised learning provides a perfect response, the reward in reinforcement learning is often a fragmentary value based on the modification of a part of the environment.Consequently, the agent performs the learning process to select an action that maximizes the total amount of future compensation.
[0036] In this way, reinforcement learning involves learning the action itself, and consequently, learning a procedure for learning a suitable action, taking into account any reciprocal effect provided by the action in the environment. That is, a procedure for performing the learning to maximize future rewards is learned. In the present embodiment, this indicates, for example, that the action information for reducing the vibration of a machine end is selected; that is, an action that influences the future can be captured.
[0037] Although any learning method can be used here as reinforcement learning, the following discussion describes as an example a case in which Q-learning, a method for learning the value Q(S, A) of selecting an action A according to the state S of a given environment, is used. One goal of Q-learning is to select, from among the actions A that can be performed in a given state S, an action A whose value Q(S, A) is the highest, as the optimal action.
[0038] However, if Q-learning is initiated first, a suitable value Q(S, A) is not found at all in any combination of state S and action A. Consequently, the agent selects various actions A according to the given state S, choosing a better action based on the compensation provided for action A at that time, in order to learn the suitable value Q(S, A).
[0039] Because it is desired that the total sum of compensation that can be received in the future be maximized, the final goal is Q(S, A) = E [Σ(γ t )r t ] to achieve. Here, E[ ] represents an expected value, t represents time, γ represents a parameter called a discount rate, which will be described later, r represents tA is a compensation at time t, and Σ represents a total sum at time t. The expected value in this formula is the expected value if the state is changed according to the optimal action. However, because it is not clear what the optimal action is in the Q-learning process, various actions are performed, consequently reinforcement learning is carried out while a search is performed. The formula for updating the value Q(S, A), as described above, can be represented, for example, by the expression 2 below (which is represented mathematically as 2 in the following). Q(St+1,At+1)←Q(St,At)+α(rt+1+γmaxA Q(St+1,A)−Q(St,At))
[0040] In the expression 2 described above, S represents t the state of the environment at time t, while A t The action at time t is represented. The state is determined by the action A. t into the t + 1 changed. rt + 1 represents a compensation that can be obtained by changing the state. A term containing max is obtained by multiplying a Q-value when an action A, whose Q-value at that time according to state S, is performed. t + 1 The highest value is selected, obtained with y. Here, γ is a parameter with 0 < γ ≤ 1, where it is called the discount rate. α is a learning coefficient, where it is assumed to fall within the range 0 < α ≤ 1.
[0041] The expression 2 described above gives a procedure for updating the value Q(S). t , A t ) of the plot A t in state S t based on the remuneration r t + 1 on, which as a result of an attempt A t is returned. This update formula indicates that if the value is max a Q(S t + 1 , A) the best action in the subsequent state S t + 1 through action A thigher than the value Q (S t , A t ) of the plot A t in state S t is, Q (S t , A t ) is increased, whereas if the value max a Q(S t + 1 , A) less than the value Q (S t , A t ) is, Q (S t , A t ) is reduced. In other words, the value of a particular action in a given state is brought close to the value of the best action in the resulting subsequent state. Although the difference therefrom is determined by the discount rate γ and the compensation r t + 1 When a change occurs, a mechanism is formed such that the value of the best action in a given state is generally spread to the value of the action in the previous state.
[0042] In Q-learning, a procedure is used where a table of Q(S, A) is generated for all state-action pairs (S, A), and the learning process is then performed on this table. However, because there is a large number of states, it is likely that Q-learning will take a long time to complete, so the values of Q(S, A) for all state-action pairs are determined.
[0043] Consequently, a well-known technique called DQN (deep Q-network) can be used. Specifically, a value function Q is configured with a suitable neural network, a parameter for the neural network is set, and the value function Q is then approximated by the suitable neural network, resulting in the computation of the value Q(S, A). Using DQN makes it possible to reduce the time required to complete Q-learning. Details of DQN are disclosed, for example, in the non-patent document below. <Nicht Patentdokument>
[0044] “Human-level control through deep reinforcement learning”, written by Volodymyr Mnihl [online], [searched on January 17, 2017], Internet<URL: http: / / files.davidqiu.com / research / nature14236.pdf>
[0045] The Q-learning described above is performed by the machine learning device 110. Specifically, the machine learning device 110 learns the value Q, in which the values of the individual coefficients a0, a1, and b1 in the transfer function of the filter 313, the input / output gain (the amplitude ratio) output by the output device 120, and the phase delay are assumed to be in state S, and in which the setting of the values of the individual coefficients a0, a1, and b1 in the transfer function of the filter 313 in state S is selected as action A.
[0046] The machine learning device 110 uses the previously described velocity command, a sine wave whose frequency is varied, to drive the servo control device 310, based on the individual coefficients a0, a1, and b1 in the transfer function of the filter 313. This command drives the servo control device 310 and thereby observes the state information S received from the output device 210, which includes the input / output gain (amplitude ratio) and phase delay for each frequency, in order to determine the action A. Each time the action A is performed, the machine learning device 110 returns a reward. The machine learning device 110 searches for the optimal action A, for example, through trial and error, so that the total sum of rewards is maximized in the future.In this way, the machine learning device 110 uses the velocity command, which is a sine wave whose frequency is varied, to drive the servo control device 310 based on the individual coefficients a0, a1 and b1 in the transfer function of the filter 313, thereby enabling it to select the optimal action A (i.e., the optimal coefficients a0, a1 and b1 in the transfer function of the filter 313) for the state S obtained from the output device 210, which includes the input / output gain (amplitude ratio) and phase delay for each frequency.
[0047] In other words, based on the value function Q learned by the machine learning device 110, an action A is selected from among the actions A applied to the individual coefficients a0, a1 and b1 in the transfer function of the filter 313 in the given state S in order to maximize the value of Q, whereby it is consequently possible to select such an action A (i.e., the individual coefficients a0, a1 and b1 in the transfer function of the filter 313) to minimize the vibration of the machine end caused by the execution of a machining program.
[0048] Fig. Figure 5 is a block diagram showing the machine learning device 110 according to the first embodiment of the present invention. To perform the reinforcement learning described above, the machine learning device 110 includes a state information acquisition unit 111, a learning unit 112, an action information output unit 113, a value function storage unit 114, and an optimization action information output unit 115, as shown in Figure 5. Fig. Figure 5 is shown. Learning unit 112 contains a compensation output unit 1121, a value function update unit 1122, and an action information generation unit 1123.
[0049] The state information acquisition unit 111 acquires the state S from the output device 210. This state S is obtained based on the individual coefficients a0, a1, and b1 in the transfer function of the filter 313, using the velocity command (the sine wave) to drive the servo motor 410. It includes the input / output gain (the amplitude ratio) and the phase delay. The state information S corresponds to an ambient state S in the Q-learning process. The state information acquisition unit 111 outputs the acquired state information S to the learning unit 112.
[0050] The individual coefficients a0, a1, and b1 in the transfer function of filter 313 are predefined by the user when Q-learning is initiated. In this example, the initial user-defined values of the individual coefficients a0, a1, and b1 in the transfer function of filter 313 are set to be optimal values for reinforcement learning. If the machine tool is predefined by the operator, machine learning with respect to the coefficients a0, a1, and b1 can be performed using these predefined values as the initial values.
[0051] Learning unit 112 is a unit that learns the value Q(S, A) when a specific action A is selected according to a specific environmental state S.
[0052] The compensation output unit 1121 is a unit that calculates compensation when action A is selected according to a given condition S. The compensation output unit 1121 compares an input / output gain Gs, measured when the individual coefficients a0, a1, and b1 in the transfer function of filter 313 are corrected to the input / output gain Gb of a given reference model for each frequency. If the measured input / output gain Gs is greater than the input / output gain Gb of the reference model, the compensation output unit 1121 provides negative compensation.On the other hand, if the measured input / output gain Gs is equal to or less than the input / output gain Gb of the reference model, the compensation output unit 1121 provides positive compensation if the phase delay is reduced, the compensation output unit 1121 provides negative compensation if the phase delay is increased, or the compensation output unit 1121 provides zero compensation if the phase delay remains the same.
[0053] An operation in which the compensation output unit 1121 provides a negative compensation when the measured input / output gain Gs is greater than the input / output gain Gb of the reference model is considered to have negative compensation when... Fig. 6 and Fig. 7 described. The compensation output unit 1121 stores the reference model of the input / output gain. The reference model is a model of the servo control device that has an ideal characteristic curve without resonance. The reference model can be calculated, for example, from the inertia Ja, a torque constant Kt, a proportional gain K P , an integral gain K I and a differential gain K D one in Fig. The inertia can be determined using the model shown in section 6. The inertia (Yes) is the sum of the motor inertia and the machine inertia. Fig. Figure 7 is a characteristic map showing the frequency response of the input / output gains of the servo control device of the reference model and the servo control device 310 before and after learning. As shown in the characteristic map after Fig. As shown in Figure 7, the reference model includes: a region A, which is a frequency range where an ideal input / output gain, which is a constant input / output gain or greater, e.g., -20 dB or greater, is provided; and a region B, which is a frequency range where an input / output gain, which is less than the constant input / output gain, is provided. In region A according to Fig. 7 The ideal input / output gain of the reference model is given by a curve MC1 (thick line). In the region B according to Fig. 7 is an ideal virtual input / output gain of the reference model by a curve MC. 11 (dashed thick line) is indicated, assuming that the input / output gain of the reference model is constant, so that it is represented by a straight line MC 12 (thick line) is indicated. In areas A and B according to Fig. Figure 7 shows the input / output gain curves of a servo control unit before and after learning, represented by curves RC1 and RC2 respectively.
[0054] In region A, the compensation output unit 1121 outputs an initial negative compensation when the RC1 curve of the measured input / output gain before learning exceeds the MC1 curve of the ideal input / output gain of the reference model. In region B, which exceeds a frequency at which the input / output gain is sufficiently small, the impact on stability is reduced, even if the RC1 curve of the input / output gain before learning exceeds the MC curve. 11 exceeds the ideal virtual input / output gain of the reference model. Consequently, in region B, as described above, the input / output gain of the reference model is replaced by the curve MC. 11 , which has an ideal gain characteristic, the straight line MC12 , which has a constant input / output gain (e.g., -20 dB). However, if the RC1 curve of the measured input / output gain before learning is the straight line MC 12 If the constant input / output gain is exceeded, a first negative value is provided as compensation because instability may occur.
[0055] Then, an operation is described in which, if the measured input / output gain Gs is equal to or less than the input / output gain Gb of the reference model, the compensation output unit 1121 determines the compensation based on the phase delay information. In the following description, the phase delay, which is a state variable related to the state information S, is represented by D(S), while the phase delay, which is a state variable related to a state S' into which state S is changed by the action information A (the correction of the individual coefficients a0, a1, and b1 in the transfer function of filter 313), is represented by D(S').
[0056] As a method where the compensation output unit 1121 determines the compensation based on the phase delay information, there is, for example, a method described below. The method for determining the compensation based on the phase delay information is not particularly restricted to the method described below. When state S is changed to state S', the compensation is determined depending on which of the following cases applies: a case where the frequency at which the phase delay is 180 degrees is increased, a case where the frequency is decreased, and a case where the frequency is the same. Although the case where the phase delay is 180 degrees is described here, there is no particular restriction to 180 degrees, and a different value can be applied. If the phase delay in the phase diagram is... Fig. As shown in Figure 4, and if state S is changed to state S', the phase delay is increased, for example, if a curve is changed such that the frequency at which the phase delay is 180 degrees (towards the direction towards X2 in Figure 4) is increased. Fig. 4) is reduced. On the other hand, if state S is changed to state S', the phase delay is reduced if the curve is changed such that the frequency at which the phase delay is 180 degrees (towards the direction towards X1 in Fig. 4) is enlarged.
[0057] When state S is changed to state S', it is defined that the phase delay D(S) is less than the phase delay D(S') if the frequency at which the phase delay is 180 degrees is decreased, with the compensation output unit 1121 setting the compensation value to a second negative value. The absolute value of the second negative value is defined as less than the first negative value. Conversely, when state S is changed to state S', it is defined that the phase delay D(S) is greater than the phase delay D(S') if the frequency at which the phase delay is 180 degrees is increased, with the compensation output unit 1121 setting the compensation value to a positive value.When state S is changed to state S', it is defined that the phase delay D(S) = the phase delay D(S') if the frequency at which the phase delay is 180 degrees remains the same, with the compensation output unit 1121 setting the value of the compensation to zero.
[0058] A negative value, defined as the phase delay D(S') in state S' after action A is greater than the phase delay D(S) in the preceding state S, can be increased according to their ratio. In the method described above, for example, the negative value is preferably increased by the degree to which the frequency is reduced. Conversely, a positive value, defined as the phase delay D(S') in state S' after action A is less than the phase delay D(S) in the preceding state S, can be increased according to their ratio. In the first method described above, the positive value is preferably increased by the degree to which the frequency is increased.
[0059] The value function update unit 1122 performs Q-learning based on the state S, the action A, the state S' when the action A has been applied to state S, and the value of a compensation calculated as described above, in order to update the value function Q stored in the value function storage unit 114. Updating the value function Q can be performed through online learning, batch learning, or mini-batch learning. Online learning is a learning method in which a specific action A is applied to the current state S, and consequently, each time state S changes to the new state S', the value function Q is immediately updated.Batch learning is a learning method in which a specific action A is applied to the current state S, and consequently, state S is repeatedly changed to the new state S', thus collecting data for learning. All collected data is then used to update the value function Q. Furthermore, mini-batch learning is an intermediate learning method between online learning and batch learning, in which the value function Q is updated each time a certain amount of data is stored for learning.
[0060] The action information generation unit 1123 selects action A in the Q-learning process for the current state S. In the Q-learning process, the action information generation unit 1123 generates the action information A for performing an operation (corresponding to action A in the Q-learning) to correct the individual coefficients a0, a1, and b1 in the transfer function of filter 313 and outputs the generated action information A to the action information output unit 113. More specifically, the action information generation unit 1123 incrementally adds or subtracts the individual coefficients a0, a1, and b1 in the transfer function of filter 313, which are contained in action A, from or to the individual coefficients a0, a1, and b1 in the transfer function of filter 313 that are contained in state S.
[0061] When the action information generation unit 1123 applies the increase or decrease of the individual coefficients a0, a1 and b1 in the transfer function of the filter 313, the state is changed to state S', with a positive compensation being returned, and the action information generation unit 1123 can select the subsequent action A', such as to incrementally perform the addition or subtraction on the individual coefficients a0, a1 and b1 in the transfer function of the filter 313 in the same way as the previous action, so that the measured phase delay is smaller than the previous phase delay.
[0062] If, on the other hand, a negative compensation is returned, the action information generation unit 1123 can select the subsequent action A', such as incrementally performing the subtraction or addition to the individual coefficients a0, a1 and b1 in the transfer function of the filter 313, in a manner opposite to the previous action, so that if the measured input / output gain is greater than the input / output gain of the reference model, the difference in input / output gain is reduced compared to the previous action, or the measured phase delay is smaller than the previous phase delay.
[0063] The action information generation unit 1123 can select the action A' by a known procedure, such as a greedy procedure of selecting the action A' whose value Q(S, A) is the highest in the value of the action A currently being estimated, or an ε-greedy procedure of randomly selecting the action A' with a small probability ε and otherwise selecting the action A' whose value Q(S, A) is the highest.
[0064] The action information output unit 113 is a unit that sends the action information A output by the learning unit 112 to the filter 313. As previously described, based on the action information, the filter 313 slightly corrects the current state S, i.e., the individual coefficients a0, a1, and b1 that are currently set, in order to change the state to the subsequent state S' (i.e., the individual coefficients of the filter 313 that are corrected).
[0065] The value function storage unit 114 is a storage device that stores the value function Q. The value function Q can be stored, for example, for each state S or each action A in a table (hereinafter referred to as an action-value table). The value function Q stored in the value function storage unit 114 is updated by the value function update unit 1122. The value function Q stored in the value function storage unit 114 can be shared with another machine learning device 110. When the value function Q is shared between several machine learning devices 110, reinforcement learning can be distributed across the individual machine learning devices 110, thereby improving the efficiency of reinforcement learning.
[0066] The optimization action information output unit 115 generates the action information A (hereinafter referred to as the "optimization action information") based on the value function Q, which is updated in the result of the value function update unit 1122 that performs Q learning. This action information causes the filter 313 to perform an operation to maximize the value Q(S, A). More specifically, the optimization action information output unit 115 captures the value function Q stored in the value function storage unit 114. As described above, this value function Q is updated in the result of the value function update unit 1122, which performs Q learning. The optimization action information output unit 115 then generates the action information based on the value function Q and outputs the generated action information to the filter 313.The optimization action information described above, like the action information output by the action information output unit 113 in the Q-learning process, contains the information for correcting the individual coefficients a0, a1 and b1 in the transfer function of the filter 313.
[0067] In filter 313, the individual coefficients a0, a1, and b1 in the transfer function are corrected based on this action information. The machine learning device 110 performs the operation described above to optimize the individual coefficients a0, a1, and b1 in the transfer function of filter 313, thereby reducing the vibration of the machine end.
[0068] As described above, the machine learning device 110 is used according to the present example, making it possible to simplify the parameter setting of the filter 313. Although the embodiment discussed above describes a case in which there is one resonance point in the machine driven by the servomotor 410, there may be multiple resonance points in the machine. If there are multiple resonance points in the machine, multiple filters are provided to correspond to the individual resonance points, connected in series, thus making it possible to dampen all resonances. The machine learning device sequentially determines the optimal values for damping the resonance points for the individual coefficients a0, a1, and b1 in the filters.
[0069] Then the output device 210 is described. Fig. Figure 8 is a block diagram showing an example of the configuration of the output device contained in the control device according to the first example of the present invention. As shown in Fig. As shown in Figure 8, the output device 210 comprises an information acquisition unit 211, an information output unit 212, a character unit 213, an actuation unit 214, a control unit 215, a storage unit 216, an information acquisition unit 217, an information output unit 218, a display unit 219, and an operation unit 220. The information acquisition unit 211 serves as an information acquisition unit that acquires the learning parameters from the device 110 for machine learning. The control unit 215 and the display unit 219 serve as an output unit that outputs the physical quantities of the learning parameters. A liquid crystal display device, a printer, or the like can be used as the display unit 219 of the output unit. The output unit contains the memory in the storage unit 216, in which case the output unit is the control unit 215 and the storage unit 216.The output device 210 has an output function of displaying as a figure the physical quantities of the control parameters (learning parameters) that are or have been machine-learned in the machine learning device 110, and the frequency response, which is determined by the physical quantities, such as the center frequency (also referred to as the attenuation center frequency), the bandwidth and the attenuation coefficient in the transfer function G(s) of the filter, and the frequency response of the filter. The output device 210 also has a setting function that controls the forwarding of information (e.g., the input / output gain and the phase delay) between the servo control device 310 and the machine learning device, and the forwarding of information (e.g.,The information for correcting the coefficients of filter 313 is transmitted between machine learning device 1100 and servo control device 310. Control (e.g., fine-tuning filter 313) is performed in servo control device 310, and control (e.g., sending an instruction to start a learning program to the machine learning device) is performed at the machine learning device 1100. The information acquisition units 211 and 217 and the information output units 212 and 218 perform the information acquisition units 211 and 217, respectively.
[0070] First, a case is considered in which the output device 210 outputs the physical quantities of the control parameters, which are learned by machine, with respect to the Fig. 9A and Fig. 9B described. Fig. 9A is a characteristic map showing a machine-learned evaluation function value and the progress of the minimum value of the evaluation function value, and a graphical representation showing an example of the display screen when the values of the learned control parameters are displayed. Fig. Figure 9B is a graphical representation showing an example of the display screen when the physical quantities of the control parameters related to state S are displayed in the display unit 219, so that they correspond to the progress of the machine learning while the machine learning is running. Even if, as in Fig. As shown in Figure 9A, the machine-learned evaluation function value, the minimum value of the evaluation function value, and the coefficients a0, a1, a2, b0, b1, and b2 in the transfer function of expression 1 are displayed on the screen of display unit 219. However, the user does not understand the physical meanings of the evaluation function and the control parameters, making it difficult to understand the learning progress and the resulting characteristic curve of the servo control device. Consequently, in the present example, as described below, the control parameters are modified to be easily understood by the user, such as the operator, for output. Similarly, in the second through fourth examples, the control parameters are modified to be easily understood by the user, such as the operator, for output.By pressing the "Change" button on the screen. Fig. The display screen shown in 9A can, for example, be the one in Fig. The display screen shown in 9B will be shown so that the information is displayed in a way that is easily understood by the user. As in Fig. As shown in Figure 9B, for example, column P1 of a setup sequence on the display screen P of display unit 219 shows the selection elements for axis selection, parameter check, program editing, program start, machine learning, and completion. A column P2 is displayed on the display screen P, showing, for example, a setting goal such as the filter, a status such as the data being collected, the number of trials (the total number of trials to date with respect to the specified number of trials (hereinafter also referred to as the "maximum number of trials") until the completion of machine learning), and a button for selecting to interrupt the learning process.On the display screen P, a column P3 is shown containing the transfer function G(s) of the filter, a table of the center frequency fc, the bandwidth fw, and the attenuation coefficient R in the transfer function G(s) of the filter, and a figure showing the frequency response characteristic of the current filter and the best frequency response characteristic of the filter during the learning process. A column P4 is also displayed, containing a figure showing the progress of the center frequency (attenuation center frequency) fc over the learning steps. The information displayed on the display screen P is an example; some of the information can be displayed, e.g., only the figure showing the frequency response characteristic and the best frequency response characteristic of the filter during the learning process, or additional information can be added.
[0071] If the user, such as the operator, uses the actuation unit 214, such as a mouse or a keyboard, the “machine learning” in column P1 of the “setting process” on the in Fig. When the display screen shown in 9B is selected in the display unit 219, such as a liquid crystal display device, the control unit 215 executes an instruction to output the coefficients a0, a1 and b1, which are related to the state S, which is associated with the number of trials, the information about the setting target (learning target) of the machine learning, the number of trials, the information containing the maximum number of trials and the like, through the information output unit 212 of the machine learning device 110.
[0072] When the information acquisition unit 211 receives from the machine learning device 110 the coefficients a0, a1 and b1, which are related to the state S associated with the number of trials, the information about the setting target (learning target) of the machine learning, the number of trials, the information containing the maximum number of trials and the like, the control unit 215 stores the received information in the storage unit 216, thereby transmitting the control to the operation unit 220.
[0073] The operation unit 220 determines, from the control parameters learned by machine learning in the device 110, the specific control parameters (e.g., the coefficients a0, a1, and b1 described above, which are related to state S) at the time of amplification learning or after amplification learning, the characteristic curves (the center frequency fc, the bandwidth fw, and the attenuation coefficient R) of the filter 313 and the frequency response of the filter 313. The center frequency fc, the bandwidth fw, and the attenuation coefficient R are second physical quantities that are determined from the coefficients a0, a1, and b1.To determine the center frequency fc, the bandwidth fw and the attenuation coefficient (the damping) R from the coefficients a0, a1 and b1, a center angular frequency ωn, a fractional bandwidth ζ and the attenuation coefficient R are determined from expression 3, where the center frequency fc and the bandwidth fw are further determined from ωn = 2πfc and ζ = fw / fc. a0+b1s+s2a0+a1s+s2=ωc2+2δτωcs+s2ωc2+2τωcs+s2
[0074] Consequently, the center frequency fc, the bandwidth fw and the damping coefficient R can be determined by expression 4. fc=a02π, fw=a12π, δ=b1a1
[0075] The center frequency fc, the bandwidth fw, and the attenuation coefficient R can be calculated using ωn = 2πfc and ζ = fw / fc from the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R, which are determined by assuming that the transfer function on the right-hand side of expression 3 is the transfer function of filter 313, and the parameters of the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R are machine-learned in the device 110. In this case, the center frequency fc, the bandwidth fw, and the attenuation coefficient R are the first physical quantities. The first physical quantities can be transformed into the second physical quantities, which can then be displayed.When the operation unit 220 calculates the center frequency fc, the bandwidth fw, and the attenuation coefficient R, and the transfer function, which includes the center angular frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R, is determined on the right-hand side of expression 3, the control is transferred to the control unit 215. Although a case where the filter is a notch filter is described here, the center frequency fc, the bandwidth fw, and the attenuation coefficient R can be determined even if the filter is in the form of a general formula, as given in expression 1, because the filter has a gain valley. In general, it is equally possible to determine one or more of the attenuated center frequency fc, bandwidth fw, and attenuation coefficient R, regardless of the filter order.
[0076] The control unit 215 stores the physical quantities of the center frequency fc, the bandwidth fw, and the attenuation coefficient R, as well as the transfer function containing the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R, in the memory unit 216, and transmits the processing to the drawing unit 213. The drawing unit 213 determines the frequency response of the filter 313 from the transfer function containing the coefficients a0, a1, and b1, which are related to the state S associated with the number of trials; the transfer function containing the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R (which are the first physical quantities); or the transfer function containing the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R (which are the second physical quantities), which is determined from the coefficients a0, a1, and b1. become,It generates a frequency-gain characteristic map, performs the processing to add the most prominent frequency response curve of the filter during learning to the frequency-gain characteristic map, generates the graphic information of the frequency-gain characteristic map to which the most prominent frequency response curve has been added, furthermore generates a graph showing the progress of the center frequency (attenuation center frequency) fc for the learning steps, generates the graphic information of the graph, and transmits the control to the control unit 215. The frequency response of the filter 313 can be determined from the transfer function on the right-hand side of expression 3. Software that can analyze the frequency response from the transfer function is known, for example, the following can be used: https: / / jp.mathworks.com / help / signal / ug / frequency-renpose.htm1; https: / / jp.mathworks.com / help / signal / ref / freqz.html; https: / / docs.scipy.org / doc / scipy~0.19.1 / reference / generated / scipy.signal.freqz.html; https: / / wiki.octave.org / Control_package; and the like.
[0077] The control unit 215 displays the frequency-gain characteristic map (which indicates the characteristic curve of the frequency response), the table formed with the center frequency fc, the bandwidth fw and the attenuation coefficient (damping) R (which are the second physical quantities), and the figure showing the progress of the transfer function G(s) of the filter and the center frequency (attenuation center frequency) fc for the learning steps, as in Fig. Figure 9B shows that although the center frequency fc, the bandwidth fw, and the attenuation coefficient (damping) R, which are the second set of physical quantities, as well as the frequency-gain map, which specifies the frequency response characteristic, are shown here, any one of them can be displayed. Instead of the frequency-gain map, which specifies the frequency response characteristic, or together with the frequency-gain map, which specifies the frequency response characteristic, a time-gain map, which specifies the time response characteristic, can be displayed. This is the same as in the second to fourth examples, which are described later. The control unit 215 displays "Notch Filter" in the setting target element of column P2 on the [unclear - possibly "displaying the frequency response characteristic"]. Fig. The display screen P shown in Figure 9B, for example, is based on information indicating that the notch filter is the setting target. If the number of trials does not reach the maximum number of trials, the status element on the display screen shows "Data is being collected." The control unit 215 also displays a ratio of the number of trials to the maximum number of trials in the trial count element on the display screen.
[0078] The in Fig. The display screen shown in 9B is an example, and the present invention is not limited to this display screen. Information other than the elements illustrated above can be displayed. The display of information from some of the elements illustrated above can be omitted. Although in the description above, the control unit 215 stores the information received from the machine learning device 110 in the storage unit 216 and displays, for example, the information about the frequency response of the filter 213, related to state S associated with the number of trials, in real time on the display unit 219, there is no limitation to this configuration. The examples in which a display is not generated in real time include the following. Variation 1: When the operator provides a display instruction, the information shown in Fig. The information shown in 9B is displayed. Variation 2: When the total number of trials (after the start of learning) reaches a predetermined number, the information shown in Fig. The information shown in 9B is displayed. Variation 3: If machine learning is interrupted or completed, the information shown in Fig. Information shown in 9B is displayed.
[0079] Even in variations 1 to 3 described above, the control unit 215 stores the received information in the storage unit 216, as in the operation of the real-time display described above, when the information acquisition unit 211 receives from the machine learning device 110 the coefficients a0, a1, and b1, which are related to the state S associated with the number of trials, the information about the setting goal (learning objective) of the machine learning, the number of trials, and information including the maximum number of trials, and the like. Subsequently, the control unit 215 transmits the control to the operation unit 220 and the display unit 213 when, in variation 1, the operator provides a display instruction, when, in variation 2, the total number of trials reaches a predetermined number, or when, in variation 3, the machine learning is interrupted or completed.
[0080] Then the output function and the setting function of the output device 210 are described. Fig. Figure 10 is a flowchart showing the operation of the control device from the start of machine learning until its completion, while attention is focused on the output device. In step S31, the control unit 215 outputs to the output device 210 when the operator, using the actuation unit 214, such as a mouse or keyboard, signals the "program start" in column P1 of the "setting sequence" on the display screen of the display unit 219, which is located in Fig. As shown in Figure 9, the system selects an instruction to start the learning program, which is then sent by the information output unit 212 to the machine learning device 110. The output device 210 then sends a learning program start instruction message to the servo control device 310, reporting the output of the learning program start instruction to the machine learning device 110. In step S32, the output device 210 provides a sine wave output instruction to a high-level device, which outputs a sine wave to the servo control device 310. Step S32 can be executed before or concurrently with step S31. Upon receiving the sine wave output instruction, the high-level device outputs a sine wave signal, the frequency of which is varied, to the servo control device 310.In step S21, the machine learning device 110 begins machine learning when the machine learning device 110 receives the learning program start instruction.
[0081] In step S11, the servo control device 310 controls the servo motor 410 to output parameter information, the input / output gain and phase delay, and the information contained in the coefficients a0, a1, and b1 in the transfer function of the filter 313 (which serve as the parameter information) to the output device 210. The output device 210 then outputs the parameter information, input / output gain, and phase delay to the machine learning device 110.
[0082] The machine learning device 110 outputs the information contained in the coefficients a0, a1, and b1 in the transfer function of filter 313, which are related to the state S associated with the number of trials used in the compensation output unit 2021 while the machine learning operation is performed in step S21, the maximum number of trials, the number of trials, and the correction information (serving as the parameter correction information) of the coefficients a0, a1, and b1 in the transfer function of filter 313 to the output device 210. In step S33, when the “machine learning” in column P1 of the “setting process” on the in Fig. When the display screen shown in Figure 9B is selected, the output device 210 uses the output function described above to convert the correction information of the coefficients in the transfer function of the filter 313, which are machine-learned in the device 110, into a graphical representation that shows the progress of the physical quantities (the center frequency fc, the bandwidth fw, and the attenuation coefficient R), which are easily understood by the user, such as the operator, and the center frequency (attenuation center frequency) fc with respect to the learning steps, and a frequency response map, outputting it to the display unit 219. In step S33, or before or after step S33, the output device 210 feeds the coefficients in the transfer function of the filter 313 to the servo control device 310.Step S11, step S21 and step S33 are repeated until the machine learning is complete.
[0083] Although the case described here is in which the physical quantities (the center frequency fc, the bandwidth fw and the attenuation coefficient R) of the coefficients in the transfer function of the filter 313, which are related to the control parameters that are machine learned, and the information related to the frequency response map are output to the display unit 219 in real time, in the cases of variations 1 to 3, which have already been described as examples of a case in which a display is not generated in real time, the physical quantities of the coefficients in the transfer function of the filter 313 and the information related to the frequency response map can be output to the display unit 219 in real time.
[0084] In step S34, the output device 210 determines whether the number of trials has reached the maximum. If the maximum number of trials is reached, the output device 210 sends a termination instruction to the machine learning device 210 in step S35. If the maximum number of trials is not reached, the process returns to step S33. When the machine learning device 210 receives the termination instruction in step S35, it completes the machine learning process. The first example of the output device and the control device in the first embodiment has been described above, followed by a description of the second example. <Zweites Beispiel>
[0085] The present example is an example where the machine learning device 110 learns the coefficients of a velocity feedforward processing unit contained in a servo control device 320, and where the output device 210 displays the progress of the frequency response and position error of the velocity feedforward processing unit in the display unit. Fig. Figure 11 is a block diagram showing the overall configuration of a control device and the configuration of the servo control device according to the second example of the present invention. The control device of the present example differs from that described in Fig. The control device shown in Figure 1 is configured as follows: the servo control device and the machine learning device and output device are configured as follows: The configurations of the machine learning device and output device in this example are the same as those of the machine learning device and output device in the first example, which differs with respect to the Fig. 5 and Fig. 8 has been described.
[0086] As in Fig. As shown in Figure 11, the servo control device 320 comprises as its constituent elements a subtractor 321, a position control unit 322, an adder 323, a subtractor 324, a speed control unit 325, an adder 326, an integrator 327, the speed feedforward processing unit 328, and a position feedforward processing unit 329. The adder 326 is connected to the servo motor 410 via a current control unit, which is not shown. The speed feedforward processing unit 328 includes a double differentiator 3281 and an IIR filter 3282.Although the position feedforward processing unit 329 does not contain the IIR filter, the IIR filter is provided, with the coefficients of the IIR filter being learned as in the velocity feedforward processing unit. As described later, the output device 210 can be used to output information on the frequency response of the IIR filter, the time response of a position error, a frequency response, and the like. In other words, the output device 210 can be used to output information on the frequency response of the velocity feedforward processing unit 328 and / or the position feedforward processing unit 329, the time response of the position error, the frequency response, and the like.
[0087] A position command is issued to the subtractor 321, the velocity feedforward processing unit 328, the position feedforward processing unit 329, and the output device 210. The subtractor 321 determines a difference between a position command value and a detection position that has been subjected to position feedback and outputs the difference as the position error to the position control unit 322 and the output device 210.
[0088] The position command is generated by the high-level device based on a program for operating the servomotor 410. The servomotor 410 is, for example, contained in a machine tool. When a table with a workpiece mounted on it moves in an X-axis and a Y-axis direction in a machine tool, the servo control device 320 and the servomotor 410, which are located in Fig. Figure 11 illustrates the arrangement in the X-axis and Y-axis directions. When the table is moved in the directions of three or more axes, the servo control device 320 and the servo motor 410 are positioned in the respective axis directions. The position command specifies a feed rate such that a machined shape, as specified by the machining program, is created.
[0089] The position control unit 322 outputs a value obtained by multiplying a position gain Kp by the position error as a velocity command value to the adder 323.
[0090] The adder 323 adds the velocity command value and the output value (the position feedforward term) of the position feedforward processing unit 329 and outputs the result as a forward-controlled velocity command value to the subtractor 324. The subtractor calculates the difference between the output of the adder 323 and a feedback velocity detection value and outputs the difference as a velocity error to the velocity control unit 325.
[0091] The speed control unit 325 adds a value obtained by multiplying and integrating an integral gain K1v with the speed error and a value obtained by multiplying a proportional gain K2v with the speed error and outputs an addition result as a torque command value to the adder 326.
[0092] The adder 326 adds the torque command value and an output value (the velocity forward feedback term) of the velocity forward feedback processing unit 328 and outputs the addition result as a forward-controlled torque command value via a current control unit (not shown) to the servo motor 410 to drive the servo motor 410.
[0093] The angular position of the servomotor 410 is detected by a rotary encoder, which serves as a position detection unit associated with the servomotor 410. A velocity detection value is input to the subtractor 324 as velocity feedback. The velocity detection value is integrated by the integrator 327 so that it becomes a position detection value, which is then input to the subtractor 102 as position feedback.
[0094] The double differentiator 3281 of the velocity feedforward processing unit 328 differentiates the position command value twice and multiplies one differentiation result by a constant β. The IIR filter 3282 performs an IIR filtering process, defined by a transfer function VFF(z) in expression 5 (specified by Math. 1 below), on the output of the double differentiator 3281 and outputs the processing result as a velocity feedforward term to the adder 326. The coefficients c1, c2, d0 to d2 in expression 5 are the coefficients of the transfer function of the IIR filter 3282. Although the denominator and numerator of the transfer function VFF(z) are quadratic functions in this example, they are not specifically restricted to being quadratic; they can be cubic or higher-order functions. VFF(z)=b0+b1z−1+b2z−21+a1z−1+a2z−2
[0095] The position feedforward processing unit 329 differentiates the position command value and multiplies the differentiation result by a constant α, outputting the processing result as a position feedforward term to the adder 323. The servo control device 320 is configured in this way.
[0096] The machine learning device 110 executes a predefined machining program (hereinafter also referred to as a "learning machining program") to learn the coefficients in the transfer function of the IIR filter 3282 of the velocity feedforward processing unit 328. Here, a machining shape designated by the learning machining program is an octagon or a shape in which the vertices of an octagon are alternately replaced by arcs. Here, the machining shape designated by the learning machining program is not limited to these machining shapes, but can be other machining shapes.
[0097] Fig. Figure 12 is a graphical representation for describing an operation of a motor when a machining shape is an octagon. Fig. Figure 13 is a graphical representation for describing an operation of a motor, where a machining mode is a shape in which the corners of an octagon are alternately replaced by arcs. In the Fig. 12 and Fig. 13 assumes that a table is moved in the X and Y axis directions so that a workpiece (workpiece) is machined in the clockwise direction.
[0098] If the processing shape is an octagon, as in Fig. As illustrated in Figure 12, the rotational speed of a motor moving the table in the Y-axis direction decreases at corner position A1, whereas the rotational speed of a motor moving the table in the X-axis direction increases. At corner position A2, the direction of rotation of the motor moving the table in the Y-axis direction is reversed, with the motor moving the table in the X-axis direction rotating at the same speed and in the same direction from position A1 to position A2 and from position A2 to position A3. The rotational speed of the motor moving the table in the Y-axis direction increases at corner position A3, whereas the rotational speed of a motor moving the table in the X-axis direction decreases.At position A4 of a corner, the direction of rotation of the motor that moves the table in the X-axis direction is reversed, while the motor that moves the table in the Y-axis direction rotates at the same speed in the same direction from position A3 to position A4 and from position A3 to the next corner position.
[0099] If the processing shape is a shape in which the corners of an octagon are alternately replaced by arcs, as in Fig. As illustrated in Figure 13, the rotational speed of a motor moving the table in the Y-axis direction decreases at corner position B1, whereas the rotational speed of a motor moving the table in the X-axis direction increases. At position B2 of an arc, the direction of rotation of the motor moving the table in the Y-axis direction is reversed, while the motor moving the table in the X-axis direction rotates from position B1 to position B3 in the same direction and at a constant speed. This differs from the case where the machining shape is an octagon, which is Fig. As illustrated in Figure 12, the rotational speed of the motor moving the table in the Y-axis direction gradually decreases as it approaches position B2, with the rotation stopping at position B2, and the rotational speed gradually increases as it moves away from position B2, so that the machining shape of an arc is formed before and after position B2.
[0100] The rotational speed of the motor that moves the table in the Y-axis direction increases at corner position B3, whereas the rotational speed of the motor that moves the table in the X-axis direction decreases. The direction of rotation of the motor that moves the table in the X-axis direction is reversed at arc position B4, causing the table to move in a linearly reversed direction in the X-axis direction. Furthermore, the motor that moves the table in the Y-axis direction rotates at the same speed and in the same direction from position B3 to position B4 and from position B4 to the next corner position.The rotational speed of the motor that moves the table in the X-axis direction gradually decreases as it approaches position B4, stopping the rotation at position B4, and the rotational speed gradually increases as it moves away from position B4, so that a machining shape of an arc is formed before and after position B4.
[0101] In the present embodiment, vibration is evaluated when the rotational speed is changed between position A1 and position A3 and between position B1 and position B3 in the linear control of the machined shape specified by the machine learning program. The influence of the position error is checked, and consequently, machine learning is performed to optimize the coefficients in the transfer function of the IIR filter 3282 of the velocity feed-through unit 328, as described above. The machine learning for optimizing the coefficients in the transfer function of the IIR filter is not specifically limited to the velocity feed-through processing unit; it can also be performed, for example, on other components.to the position feedforward processing unit containing the IIR filter, or to a current feedforward processing unit provided when current feedforward is performed in the servo control device, and which contains the IIR filter.
[0102] The following section describes the machine learning device 110 in more detail. An example of machine learning is described assuming that the machine learning device 110 of the present embodiment performs reinforcement learning when optimizing the coefficients in the transfer function of the IIR filter 3282 of the velocity feedforward processing unit 328. Moreover, the machine learning in the present invention is not limited to reinforcement learning; the present invention can also be applied to a case of performing other machine learning (e.g., supervised learning).
[0103] The machine learning device 110 learns a value Q of selecting an action A of setting the coefficients a1, a2 and b0 to b2 of the transfer function VFF(z) of the IIR filter 3282, which are associated with a state S, wherein the state S is a servo state, such as the commands and feedbacks, which contains the coefficients a1, a2 and b0 to b2 of the transfer function VFF(z) of the IIR filter 3282 of the velocity feedforward processing unit 328 and the position error information and the position command value of the servo control device 320, which are acquired by executing the machine learning processing program (hereinafter referred to as learning).Specifically, the machine learning device 100, according to the embodiment of the present invention, determines the coefficients of the transfer function VFF(z) of the IIR filter 3282 by searching within a predetermined range for a radius r and an angle θ, which represent a zero point and a pole of the transfer function VFF(z) in polar coordinates, respectively, in order to learn the radius r and the angle θ. The pole is the value of z at which the transfer function VFF(z) is infinite, while the zero point is the value of z at which the transfer function VFF(z) is 0. Consequently, the coefficients in the numerator of the transfer function VFF(z) are modified as follows. b0+b1z−1+b2z−2=b0(1+(b1 / b0)z−1+(b2 / b0)z−2)
[0104] In the following, (b1 / b0) and (b2 / b0) are referred to as b1' and b2', respectively, unless otherwise specified. The machine learning device 110 learns the radius r and the angle θ, which minimize the positional error, in order to determine the coefficients a1, a2, b1', and b2' of the transfer function VFF(z). The coefficient b0 can be obtained, for example, by performing the machine learning after the radius r and the angle θ have been set to the optimal values r0 and θ0. The coefficient b0 can be learned simultaneously with the angle θ. Furthermore, the coefficient b0 can be learned simultaneously with the radius r.
[0105] The machine learning device 110 observes the state information S, which contains the servo state, such as the commands and feedback, the position commands, and the position error information of the servo control device 320 at positions A1 and A3 and / or positions B1 and B3 of the machining form. It does this by executing the learning machining program based on the values of the coefficients a1, a2, and b0 to b2 of the transfer function VFF(z) of the IIR filter 3282 to determine the action A. The machine learning device 110 receives a reward whenever the action A is performed. The machine learning device 110 searches for the optimal action A in a trial-and-error manner, such that the total sum of the rewards is maximized over the course of the future. In this way, the machine learning device 110 can determine an optimal action A (i.e.,, select the values of the optimal zero point and the optimal pole of the transfer function VFF(z) of the IIR filter 3282 with respect to state S, which contains the servo state, such as the commands and feedback, which contain the position commands and the position error information of the servo control device 320, which are acquired by executing the learning processing program based on the values of the coefficients calculated based on the values of the zero point and pole of the transfer function VFF(z) of the IIR filter 3282. The direction of rotation of the servo motor in the X-axis and Y-axis directions does not change positions A1 and A3 and positions B1 and B3, so that the machine learning device 110 can learn the values of the zero point and pole of the transfer function VFF(z) of the IIR filter 3282 during linear operation.
[0106] That is, it is possible to select an action A (i.e., the values of the zero point and the pole of the transfer function VFF(z) of the IIR filter 3282) that minimizes the position error detected by executing the learning processing program, by selecting an action A that maximizes the value of Q, from among the actions A applied to the transfer function VFF(z) of the IIR filter 3282 associated with a particular state S, based on the value function Q learned by the machine learning device 110.
[0107] The following describes a method for learning the radius r and the angle θ, which represent the zero point and the pole of the transfer function VFF(z) of the IIR filter 3282, which minimize the position error, in polar coordinates to obtain the coefficients a1, a2, b1' and b2' of the transfer function VFF(z), and a method for obtaining the coefficient b0.
[0108] The machine learning device 110 defines a pole, which is z, where the transfer function VFF(z) in expression 5 is infinite, and a zero point, which is z, where the transfer function VFF(z) is 0, which are detected by the IIR filter 3282. The machine learning device 110 multiplies the denominator and the numerator in expression 5 by z. 2 , in order to obtain expression 6 (which is hereafter referred to as Math. 6) in order to obtain the pole and the zero point. VFF(Z)=b0(z2+b1'z+b2')z2+a1z+a2
[0109] The pole is the z where the denominator of the expression is 6 0 (i.e., z 2 + a1z + a2 = 0), while the zero point is the z where the numerator of the expression 6 is 0 (i.e., z 2 + b1' z + b2' = 0).
[0110] In the present embodiment, the pole and the zero point are represented in polar coordinates, and the searches for the pole and the zero point are also represented in polar coordinates. The zero point is important when suppressing an oscillation, and the machine learning device 110 first determines the pole and the coefficients b1' (=-re iθ - re -iθ ) and b2' (= r 2 ), which are calculated when z = re iθ and the complex conjugate z* = re -iθ in the numerator (z 2 + b1' z + b2') the zero point are (the angle θ is in a given range where 0 ≤ r ≤ 1), as the coefficients of the transfer function VFF(z) determine the zero point re iθThe values of the optimal coefficients b1' and b2' are determined using polar coordinates. The radius r depends on a damping factor, while the angle θ depends on a vibration suppression frequency. The zero point can then be set to an optimal value, allowing the value of the coefficient b0 to be learned. Subsequently, the pole of the transfer function VFF(z) is represented in polar coordinates, where the value of re is determined. iθThe pole, represented in polar coordinates, is sought using a method similar to that used for the zero point. In this way, it is possible to learn the values of the optimal coefficients a1 and a2 in the denominator of the transfer function VFF(z). Once the pole is determined and the coefficients in the numerator of the transfer function VFF(z) are known, it is sufficient to suppress gain on the high-frequency side, where the pole corresponds, for example, to a second-order low-pass filter. A transfer function of a second-order low-pass filter is represented, for example, by the expression 7 (hereafter denoted as Math. 7). ω is a peak gain frequency of the filter. 1s² + 2ωs + ω²
[0111] If the pole is a third-order low-pass filter, the third-order low-pass filter can be formed by providing three first-order low-pass filters whose transfer function is represented by 1 / (1 + Ts) (where T is a time constant of the filter). It can be formed by combining the first-order low-pass filter with the second-order low-pass filter in expression 5. The transfer function in the z-domain is obtained using a bilinear transformation of the transfer function in the s-domain.
[0112] Although the pole and zero point of the transfer function VFF(z) can be sought simultaneously, it is possible to reduce the amount of machine learning and shorten the learning time if the pole and zero point are sought and learned separately.
[0113] The search regions for the pole and the origin can be defined by specifying the radius r, e.g., in the range 0 ≤ r ≤ 1 in a complex plane. Fig. 14 and the definition of the angle θ in a relevant frequency range of a velocity loop is restricted to predefined search ranges, which are indicated by the hatched areas. The upper limit of the frequency range can, for example, be set to 110 Hz because the oscillation generated due to the resonance of the velocity loop is approximately 110 Hz. Although the search range is determined by the resonance properties of a control target, such as a machine tool, a search range of the angle θ is defined as in the complex plane according to Fig. The result is obtained because at approximately 250 Hz the angle θ corresponds to 90° when the sampling period is 1 ms, provided the upper limit of the frequency range is 110 Hz. By narrowing the search range to a predetermined area in this way, it is possible to reduce the amount of machine learning required and shorten the settling time of the machine learning process.
[0114] When searching for the origin in polar coordinates, the coefficient b0 is first set to, for example, 1, while the radius r is set to any value within the range 0 ≤ r ≤ 1, where the angle θ is in the range specified in Fig. The search area illustrated in Figure 14 is determined using a trial-and-error method in order to find the coefficients b1' (=-re jθ - re -jθ ) and b2' (= r 2 ) to determine such that z and the complex conjugate z* are the zero point of (z 2 + b1' z + b2'). The initial setting value of the angle θ is in the Fig. The search area is defined as illustrated in Figure 14. The machine learning device 110 transfers the setting information of the obtained coefficients b1' and b2' to the IIR filter 3282 as action A, setting the coefficients b1' and b2' in the numerator of the transfer function VFF(z) of the IIR filter 3282. The coefficient b0 is, for example, set to 1, as described above. When such an ideal angle θ0, maximizing the value of the value function Q, is determined by the machine learning device 110, which performs the learning to search for the angle θ, the angle θ is set to the angle θ0, varying the radius r to thereby determine the coefficients b1' (=-re jθ - re -jθ ) and b2' (= r 2) in the numerator of the transfer function VFF(z) of the IIR filter 3282. By learning to search for the radius r, such an optimal radius r0, which maximizes the value of the value function Q, is determined. The coefficients b1' and b2' are determined using the angle θ0 and the radius r0, and then the learning process with respect to b0 is performed, thereby determining the coefficients b0, b1' and b2' in the numerator of the transfer function VFF(z).
[0115] When searching for the pole in polar coordinates, the learning process can be performed similarly to the denominator of the transfer function VFF(z). First, the radius r is set to a value within the range (e.g., 0 ≤ r ≤ 1), while the angle θ is searched for within the same range, similar to searching for the origin, to determine an ideal angle θ of the pole of the transfer function VFF(z) of the IIR filter 3282 through learning. Then, the angle θ is set to the desired angle, while the radius r is searched for and learned, thus determining the ideal angle θ and the ideal radius r of the pole of the transfer function VFF(z) of the IIR filter 3282. In this way, the optimal coefficients a1 and a2, corresponding to the ideal angle θ and the ideal radius r of the pole, are determined. As described above, the radius r depends on the damping factor, while the angle θ depends on a vibration suppression frequency.Therefore, it is preferable to learn the angle θ before the radius r in order to suppress the oscillation.
[0116] In this way, by searching within a given range for the radius r and the angle θ, which represent the zero point and the pole of the transfer function VFF(z) of the IIR filter 3282 in polar coordinates, respectively, so that the position error is minimized, it is possible to perform the optimization of the coefficients a1, a2, b0, b1' and b2' of the transfer function VFF(z) more efficiently than by directly learning the coefficients a1, a2, b0, b1' and b2'.
[0117] When the coefficient b0 of the transfer function VFF(z) of the IIR filter 3282 is learned, the initial value of the coefficient b0 is set to 1, and subsequently the coefficient b0 of the transfer function VFF(z), which is contained in action A, is incrementally increased or decreased. The initial value of the coefficient b0 is not restricted to 1. The initial value of the coefficient b0 can be set to any value. The machine learning device 110 sends back compensation based on a positional error whenever action A is performed and adjusts the coefficient b0 of the transfer function VFF(z) to an ideal value that maximizes the value of the value function Q, thus maximizing future total compensation, through reinforcement learning of the search for the optimal action A in a trial-and-error manner.Although in this example the learning of the coefficient b0 is performed after the learning of the radius r, the coefficient b0 can be learned concurrently with the angle θ. While the radius r, the angle θ, and the coefficient b0 can be learned simultaneously, it is possible to reduce the amount of machine learning required and shorten the settling time of the machine learning process by learning these coefficients separately.
[0118] The configuration of the device 110 for machine learning in Fig. 11 is the same as the one in Fig. 5 configuration shown, with the following a description regarding Fig. 5 is given. The state information acquisition unit 111 acquires the state S from the servo control device 320, which contains a servo state, such as the commands and feedback, the position commands, and the position error information of the servo control device 132. This is acquired by executing the learning processing program based on the values of the coefficients a1, a2, and b0 to b2 of the transfer function VFF(z) of the IIR filter 3282 of the velocity feedforward processing unit 328 of the servo control device 320. The state information S corresponds to a state S of the environment during Q-learning. The state information acquisition unit 111 outputs the acquired state information S to the learning unit 112.Furthermore, the state information acquisition unit 111 acquires the angle θ and the radius r, which represent the zero point and the pole in polar coordinates, and the corresponding coefficients a1, a2, b1' and b2' from the action information generation unit 1123 in order to store them therein, outputting the angle θ and the radius r, which represent the zero point and the pole, corresponding to the coefficients a1, a2, b1' and b2' acquired by the servo control device 320, in polar coordinates to the learning unit 112.
[0119] The initial settings of the transfer function VFF(z) of the IIR filter 3282 at the start of Q-learning are predetermined by a user. In the present embodiment, the user-defined initial settings of the coefficients a1, a2, and b0 to b2 of the transfer function VFF(z) of the IIR filter 3282 are then optimized by gain learning by searching for the radius r and the angle θ, which represent the origin and the pole in polar coordinates, as described above. The coefficient α of the double differentiator 3281 of the velocity feedforward processing unit 328 is, for example, set to a fixed value, such as α = 1. Furthermore, the initial settings of the denominator of the transfer function VFF(z) of the IIR filter 3282 are set to those specified in Math.Figure 5 (of the transfer function implemented by a bilinear transformation in the z-domain) illustrates this. Furthermore, with respect to the initial setpoint values of the coefficients b0 to b2 in the numerator of the transfer function VFF(z), b0 = 1, the radius r can be set to a value within the range 0 ≤ r ≤ 1, while the angle θ can be set to a value within the specified search range. Additionally, with respect to the coefficients a1, a2, and b0 to b2, and the coefficients c1, c2, and d0 to d2, if an operator pre-sets the machine tool, machine learning can be performed using the values of the radius r and the angle θ, which represent the zero point and the pole of the set transfer function in polar coordinates, as the initial values.
[0120] Learning Unit 112 is a unit that learns the value Q(S, A) when a specific action A is selected according to a given state S of the environment. Here, action A sets the coefficient b0 to, for example, 1, calculating the correction information of the coefficients b1' and b2' in the numerator of the transfer function VFF(z) of the IIR filter 3282 based on the correction information of the radius r and the angle θ, which represent the zero point of the transfer function VFF(z) in polar coordinates. The following description provides an example of a case where the coefficient b0 is initially set to, for example, 1, and the action information A consists of the correction information of the coefficients b1' and b2'.
[0121] The compensation output unit 1121 is a unit that calculates compensation when action A is selected according to a specific state S. Here, a set (a position error set) of position errors that are the state variables of state S is denoted by PD(S), while a position error set consisting of the state variables related to the state information S' that is modified by state S due to action information A is denoted by PD(S'). Furthermore, the position error value in state S is a value calculated based on a predefined evaluation function f(PD(S)). The functions that can be used as the evaluation function f include: A function that calculates an integrated value of an absolute value of a positional error, ∫|e|dt. A function that calculates an integrated value by weighting an absolute value of a positional error over time, ∫tleldt. A function that calculates an integrated value of a 2n-th power (n is a natural number) of an absolute value of a positional error, ∫e 2n dt (n is a natural number). A function that calculates a maximum value of an absolute value of a positional error. Max{|e|}.
[0122] In this case, the compensation output unit 1121 sets the value of a compensation to a negative value if the position error value f(PD(S')) of the servo control device 320, which is operated based on the velocity forward feedback processing unit 328, after the correction based on the state information S' corrected by the action information A, is greater than the position error value f(PD(S)) of the servo control device 320, which is operated based on the velocity forward correction processing unit 328, before the correction based on the state information S before it is corrected by the action information A.
[0123] On the other hand, the compensation output unit 1121 sets the value of the compensation to a positive value if the position error value f(PD(S')) of the servo control device 320, which is operated based on the velocity forward feedback processing unit 328, after the correction based on the state information S' corrected by the action information A, is smaller than the position error value f(PD(S)) of the servo control device 320, which is operated based on the velocity forward correction processing unit 328, before the correction based on the state information S before it is corrected by the action information A.Furthermore, the compensation output unit 1121 can set the compensation value to zero if the position error value f(PD(S')) of the servo control device 320, which is operated based on the velocity forward coupling processing unit 328, after the correction based on the state information S' corrected by the action information A, is equal to the position error value f(PD(S)) of the servo control device 320, which is operated based on the velocity forward correction processing unit 328, before the correction based on the state information S before it is corrected by the action information A.
[0124] If, after the execution of action A, the position error value f(PD(S')) in state S' becomes greater than the position error value f(PD(S)) in the preceding state S, the negative value can be increased according to the ratio. That is, the negative value can be increased according to the degree of increase in the position error value. Conversely, if, after the execution of action A, the position error value f(PD(S')) in state S' becomes less than the position error value f(PD(S)) in the preceding state S, the positive value can be increased according to the ratio. That is, the positive value can be increased according to the degree of decrease in the position error value.
[0125] The value function update unit 1122 updates the value function Q, which is stored in the value function storage unit 114, by performing Q-learning based on the state S, the action A, the state S' when the action A has been applied to state S, and the value of the compensation calculated as above. Updating the value function Q can be performed through online learning, batch learning, or mini-batch learning.
[0126] The action information generation unit 1123 selects action A in the Q-learning process with respect to the current state S. The action information generation unit 1123 generates the action information A and outputs the generated action information A to the action information output unit 113 to perform an operation (corresponding to action A of Q-learning) of correcting the coefficients b1' and b2' of the transfer function VFF(z) of the IIR filter 3282 of the servo control device 320 in the Q-learning process, based on the radius r and the angle θ, which represent the origin in polar coordinates. More specifically, the action information generation unit 1123 increases or decreases the angle θ received from the state information acquisition unit 111 within the range specified in the Q-learning unit 1123. Fig. 14 illustrated search range in a state where the coefficients a1, a2 and b0 of the transfer function VFF(z) in expression 6 are fixed, and the radius r received by the state information acquisition unit 111, while they determine the zero point of z in the numerator (z 2 + b1' z + b2') as re iθ This is used, for example, to search for the origin in polar coordinates. Furthermore, z, which serves as the origin, and its complex conjugate z* are determined using the fixed radius z and the enlarged or reduced angle θ, with the new coefficients b1' and b2' being calculated based on the origin.
[0127] If state S transitions to state S' by increasing or decreasing the angle θ and redefining the coefficients b1' and b2' of the transfer function VFF(z) of the IIR filter 3282, and a positive compensation is offered in return, the action information generation unit 1123 can select a strategy where an action A' that further reduces the value of the position error, such as by increasing or decreasing the angle θ similarly to the previous action, is selected as the next action A'.
[0128] If, in contrast, a negative compensation is offered in return, the action information generation unit 1123 can select a strategy where an action A' that results in the value of the position error becoming smaller than the previous value, such as by decreasing or increasing the angle θ opposite to the previous action, is selected as the next action A'.
[0129] If, through learning with the aid of the optimization action information (which will be described later) the search for the angle θ continues from the optimization action information output unit 115 and an ideal angle θ0 maximizing the value of Q is determined, the action information generation unit 1123 sets the angle θ to the angle θ0 in order to search for the radius r within the range 0 ≤ r ≤ 1, setting the coefficients b1' and b2' in the numerator of the transfer function VFF(z) of the IIR filter 3282 similarly to the search for the angle θ. When, through learning with the help of the optimization action information (which will be described later) the search for the radius r continues from the optimization action information output unit 115 and an ideal radius r0 that maximizes the value of Q is determined, the action information generation unit 1123 determines the optimal coefficients b1' and b2' in the numerator.Then, the optimal values of the coefficients in the numerator of the transfer function VFF(z) are learned by learning the coefficient b0, as described above.
[0130] The action information generation unit 1123 then searches for the coefficients of the transfer function, which are related to the numerator of the transfer function VFF(z), based on the radius r and the angle θ, which represent the pole in polar coordinates, as described above. The learning process sets the radius r and the angle θ, which represent the pole in polar coordinates, by reinforcement learning, similar to the case of the numerator of the transfer function VFF(z) of the IIR filter 3282. In this case, the radius r is learned similarly to the case of the numerator of the transfer function VFF(z) after learning the angle θ. Because the learning procedure is similar to the case of finding the zero point of the transfer function VFF(z), its detailed description is omitted.
[0131] The action information output unit 113 is a unit that sends the action information A output by the learning unit 112 to the servo control device 320. As described above, the servo control device 320 fine-tunes the current state S (i.e., the currently set radius r and the currently set angle θ, which represent the zero point of the transfer function VFF(z) of the IIR filter 3282 in polar coordinates) based on the action information in order to transition to the next state S' (i.e., the coefficients b1' and b2' of the transfer function VFF(z) of the IIR filter 3282, which correspond to the corrected zero point).
[0132] The value function storage unit 114 is a storage device that stores the value function Q. The value function Q can be stored, for example, as a table (referred to below as an action-value table) for each state S and each action A. The value function Q stored in the value function storage unit 114 is updated by the value function update unit 1122. Furthermore, the value function Q stored in the value function storage unit 114 can be shared with other machine learning devices 110. When the value function Q is shared by multiple machine learning devices 110, it is possible to improve the efficiency of reinforcement learning because reinforcement learning can be performed in a distributed manner across the respective machine learning devices 110.
[0133] The optimization action information output unit 115 generates the action information A (hereinafter referred to as the "optimization action information"), which causes the velocity feedforward processing unit 328 to perform an operation to maximize the value Q(S, A) based on the value function Q updated by the value function update unit 1122, which performs Q-learning. More specifically, the optimization action information output unit 115 captures the value function Q, which is stored in the value function storage unit 114. As described above, the value function Q is updated by the value function update unit 1122, which performs Q-learning.The optimization action information output unit 115 generates the action information based on the value function Q and outputs the generated action information to the servo control device 320 (the IIR filter 3282 of the velocity feedforward processing unit 328). The optimization action information contains the information that corrects the coefficients of the transfer function VFF(z) of the IIR filter 3282 by learning the angle θ, the radius r, and the coefficient b0, similar to the action information that the action information output unit 113 outputs in the Q-learning process.
[0134] In the servo control device 320, the coefficients of the transfer function, which are related to the numerator of the transfer function VFF(z) of the IIR filter 3182, are corrected based on action information derived from the angle θ, the radius r, and the coefficient b0. After the optimization of the coefficients in the numerator of the transfer function VFF(z) of the IIR filter 3282 has been performed using the operations described above, the machine learning device 110 performs the optimization of the coefficients in the denominator of the transfer function VFF(z) of the IIR filter 3282 by learning the angle θ and the radius r in a manner similar to the optimization. As described above, using the machine learning device 110 according to the present invention, it is possible to simplify the setting of the parameters in the velocity feedforward processing unit 328 of the servo control device 320.
[0135] In the present embodiment, the compensation output unit 1121 calculates the compensation value by comparing the evaluation function value f(PD(S)) of the position error in state S, calculated based on the predefined evaluation function f(PD(S)) using the position error PD(S) in state S as an input, with the evaluation function value f(PD(S')) of the position error in state S', calculated based on the evaluation function f(PD(S')) using the position error PD(S') in state S' as an input. However, the compensation output unit 1121 can add another element other than the position error when calculating the compensation value. The machine learning device 110 can, for example,at least one position-forward controlled velocity command issued by adder 323, one difference between a velocity feedback and a position-forward controlled velocity command, and one position-forward controlled torque command issued by adder 326, in addition to the position error issued by subtractor 102.
[0136] Although the output device 210 is then described, because its configuration is the same as that of the output device 210 of the first example, which is described in Fig. Figure 8 shows only the difference in operation. The display screen of display unit 219 in the first example is the same as the display screen after. Fig. 9B, which is shown in the first example, except that the details (such as the frequency response characteristic of the filter) of column P3 on the display screen P after Fig. 9B, which are shown in the first example, by the frequency response characteristic of a velocity feedforward processing unit, which is in Fig. Figure 15 is shown, and a graphical representation showing the characteristic curve of the position error has been replaced.
[0137] In the present example, the output device 210 outputs the servo state of the commands, feedback, and the like, which contains the coefficients a1, a2, and b0 to b1 in the transfer function VFF(z) of the IIR filter 3282 of the velocity feedforward processing unit 328, the position error of the servo control device 320, and the position command, to the machine learning device 110. Here, the control unit 215 stores the position error output by the subtractor 321 together with the time information in the memory unit 216.
[0138] If the operator uses the operating unit 214, such as a mouse or a keyboard, to activate the “machine learning” feature in column P1 of the “setting process” on the [document / tablet / etc.] Fig. When the display screen shown in 9B is selected in the display unit 219, the control unit 215 executes an instruction to output the coefficients a0, a1 and b0 to b2, which are related to the state S, which is associated with the number of trials, the information about the setting target (learning target) of the machine learning, the number of trials, the information containing the maximum number of trials, the evaluation function value and the like, through the information output unit 212 of the machine learning device 110.
[0139] When the information acquisition unit 211 receives from the machine learning device 110 the coefficients a0, a1 and b0 to b2, which are related to the state S, which is associated with the number of trials, the information about the setting target (learning target) of the machine learning, the number of trials, the maximum number of trials, the information including the evaluation function value and the like, the control unit 215 stores the received information in the storage unit 216, transmitting the control to the operation unit 220.
[0140] The operation unit 220 determines, from the control parameters that are machine-learned in the device 110 for machine learning, specifically the control parameters (e.g. the coefficients a0, a1 and b0 to b2 described above in the transfer function VFF(z) of expression 6, which is related to the state S) at the time of gain learning or after gain learning, the properties (the center frequency fc, the bandwidth fw and the attenuation coefficient R) of the IIR filter 3282 of the velocity feedforward processing unit 328.It is possible to determine the center frequency fc, the bandwidth fw and the attenuation coefficient (the damping) R from the zero point and the pole of the transfer function VFF(z), wherein, when the operation unit 220 calculates the center frequency fc, the bandwidth fw and the attenuation coefficient R to determine the transfer function VFF(z) which contains the center frequency fc, the bandwidth fw and the attenuation coefficient R, the operation unit 220 transmits the control to the control unit 215.
[0141] The control unit 215 stores the parameters of the center frequency fc, the bandwidth fw and the attenuation coefficient R and the transfer function VFF(z), which contains the center angular frequency ωn, the fractional bandwidth ζ and the attenuation coefficient R, in the storage unit 216 and transfers the processing to the character unit 213.
[0142] As described in the first example, the drawing unit 213 determines the frequency response of the IIR filter 3282 from the transfer function, which contains the center frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R, to generate the frequency-gain characteristic map. The same procedure as in the first example can be used to determine the frequency response of the IIR filter 3282 from the transfer function. The drawing unit 213 then uses the individual values of the center frequency fc, the bandwidth fw, and the attenuation coefficient (damping) R to form a table, combining them with the frequency-gain characteristic map. This yields the information about the VFF(z) according to Fig. 15. The drawing unit 213 determines the frequency response of the position error based on the position error and the position command stored in the memory unit 216 to generate a frequency-position error characteristic map. The drawing unit 213 also generates a time-response characteristic map of the position error based on the position error and its timing information. Then, the root mean square (RMS) value of a position error at each sampling point, an error peak frequency (which is a frequency peak when the position error is observed in a frequency range), and the weighting function are combined with the frequency-position error characteristic map and the time-response characteristic map of the position error. This results in the position error information. Fig. 15. The root mean square (RMS) value of the position error per sampling time and the error peak frequency can be determined using the operation unit 220. The drawing unit 213 generates the image information, in which the information about the VFF(z) and the information about the position error are combined, and transmits the control to the control unit 215.
[0143] The control unit 215 shows in column P3 after Fig. 9B the information about the VFF(z) and the information about the position error in Fig. 15. The control unit 215 displays the speed feedforward processing unit, e.g., based on the information indicating that the speed feedforward processing unit is the setting target, in the setting target element on the display screen, as in Fig. Figure 9B shows that if the number of attempts does not reach the maximum number of attempts, the status column on the display screen shows "Data is being collected". The control unit 215 also displays the ratio of the number of attempts to the maximum number of attempts in the number of attempts column on the display screen.
[0144] Even if the machine learning device 110 performs, for example, learning on the coefficients a1, a2, and b0 to b2, and the evaluation function value is not changed, the time response of the position error or the frequency response can be altered by a post-stop vibration, even in a state where the machine tool is stopped after machining. After learning, the output device 210 issues an instruction to adjust the coefficients in the speed feedforward processing unit, or an instruction to perform relearning, via an instruction from the operator who displays the screen of the display unit. Fig. 15 sees and observes a variation in the time behavior of the position error or the frequency response, the device 110 is ready for machine learning.
[0145] Fig. Figure 16 is a flowchart showing the operation of the output device after an instruction to complete the machine learning in the second example of the present invention. A flowchart in the present example showing the operation of the control device after the start of the machine learning until the instruction to complete the machine learning, while attention is focused on the output device, is the same as the one in Figure 16. Fig. 10. The sequence shown from step S31 to step S35, except that the state information is not the coefficients of the input / output gain, phase delay and notch filter, but the coefficients of the position command, position error and velocity feedforward processing unit, and that the action information is the correction information of the coefficients in the velocity feedforward processing unit.
[0146] The time-response characteristic curve of the position error and the frequency-position error characteristic curve in Fig. Figure 15 shows a case in which the position error is increased by the oscillation after the stop. Fig. 15. If the operator selects the "Setting" control panel, the individual values of the center frequency fc, the bandwidth fw, and the attenuation coefficient (damping) R in the table can be changed. If the operator selects the time-response characteristic of the position error and the frequency-response characteristic in Fig. When the operator sees 15, they change the center frequency fc in the table from 480 Hz to 500 Hz. Then, in step S36, they determine... Fig. 16. The control unit 215 is instructed that the setting is the instruction, wherein the control unit 215 in step S37 outputs a correction instruction containing correction parameters (change values of the coefficients a1, a2 and b0 to b2) of the IIR filter 3282 to the servo control device 310. The servo control device 310 returns to step S11, drives the machine tool with the changed coefficients a1, a2 and b0 to b2 and outputs the position error to the output device 210. In step S38, the output device 210 determines the frequency response of the filter 3282 based on the changed center frequency fc, displaying the frequency-gain characteristic on the display screen of the display unit 219 and the time response characteristic and the frequency-position error characteristic, which show the characteristic curve of the time response of the position error and the frequency-position error characteristic, on the display screen of the display unit 219, as shown in Fig. 17 is shown.
[0147] In this way, the operator observes the frequency response of the IIR filter 3282, the time response of the position error and the frequency response, changing one or more of the center frequency fc, the bandwidth fw and the damping coefficient (damping) R as needed, and thereby fine-tuning the characteristic curve of the frequency response of the IIR filter 3282, the characteristic curve of the time response of the position error and the characteristic curve of the frequency response.
[0148] On the other hand, if the operator presses the "Release" button, which is located in Fig. 15 is shown, in step S36 in Fig. When step 16 is selected, the control unit 215 determines that the instruction is relearning, in order to instruct the machine learning device 110 in step S39 to perform relearning at 480 Hz. The machine learning device 110 returns to step S21 to perform relearning at 480 Hz. Here, the in Fig. The search range shown in Figure 14 is changed to a range around 480 Hz or selected from a wide range to a narrow range. In step S40, the output device 210 determines the frequency response of the IIR filter 3282 based on the control parameters supplied by the machine learning device, displaying the frequency gain characteristic on the display screen of the display unit 219 and showing the time response characteristic and the frequency position error characteristic, which show the time response of the position error characteristic and the frequency position error characteristic, as shown in Figure 14. Fig. 17 is shown.
[0149] In this way, the operator observes the frequency response of the IIR filter 3282, the time response of the position error, and the frequency response in order to perform relearning with the machine learning device 110, thereby enabling the operator to relearn the characteristic curves of the frequency response of the IIR filter 3282, the time response of the position error, and the frequency response. The second example of the output device of the control device in the first embodiment has been described above, followed by a third example. <Drittes Beispiel>
[0150] In the present example, the coefficients of the speed feedforward processing unit of the control device in the second example are converted into values that can be understood by the user and that have physical meanings, specifically the coefficients in a motor blocking characteristic, a notch filter and a low-pass filter, which are in Fig. Figure 18 shows and serves as an expression model, and specifically implements the inertia J, a center angular frequency (notch frequency) ωn, a fractional bandwidth (notch attenuation), a damping coefficient (notch depth) R, and a time constant τ, outputting them. The configuration of an output device in the present example is the same as that shown in Figure 18. Fig. Output device 210 shown in 8. Although in the second example the learning is performed using polar coordinates, in the present example the learning is performed without the use of polar coordinates as in the first example.
[0151] The transfer function F(s) of the velocity feedforward processing unit 328 can be represented by expression 8 with a motor blocking characteristic 3281A, a notch filter 3282A and a low-pass filter 3283A, which serve as a model of a mathematical formula. F(s)=∑j=04bjsj∑i=04aisi =Js2(1+τs)2s2+2Rζωns+ωn2s2+2ζwns+ωn2
[0152] It follows from expression 8 that b4 = J, b3 = 2JRζω n , b1 = 0, b0 = 0, a4 = τ 2 , a3 = (2ζω n τ 2 + 2τ), a2 = (ω n 2 τ 2 + 4ζω n τ +1), a1 = (2ζω n 2 + 2ζω n ) and a0 = ω n 2Here, the damping center frequency ω is... n represented by a formula as follows. ωn=b2b4=a0
[0153] The fractional bandwidth (notch damping), the damping coefficient (notch depth) R and the time constant τ are calculated in the same way.
[0154] In this way, the output device 210 determines the physical quantities from the coefficients in the transfer function F(s). These quantities are easily understood by the user, such as the operator, and can display them on the screen of the display unit 219. A frequency response characteristic is determined from the fractional bandwidth (notch attenuation), the attenuation coefficient (notch depth) R, and the time constant τ, and can be displayed on the screen. The third example of the output device and the control device in the first embodiment has been described above, followed by a fourth example. <Viertes Beispiel>
[0155] Although the first to third examples describe the case in which the transfer function of the constituent elements of the servo control device is characterized as indicated by expressions 1, 5, and 8, the present embodiment can also be applied to a case in which the transfer function of the constituent elements of the servo control device is a transfer function of a general formula represented by expression 10 (n is a natural number). The constituent element of the servo control device is, for example, a velocity feedforward processing unit, a position feedforward processing unit, or a current feedforward processing unit. The machine learning device 110 determines, for example, the optimal coefficients a. i and b j through machine learning, so that the positional error is reduced. F(s)=∑j=0nbjsj∑i=0naisi
[0156] Then, based on the determined coefficients a i and b j or a transfer function F(s) that determines the coefficients a i and b j It contains the physical quantities, which are easily understood by the user, and the information specifying a time response or frequency response, which is output by the output device 210. When the frequency response is determined, known software that can analyze the frequency response of a transfer function is used, consequently determining the frequency response of the transfer function F(s), which contains the determined coefficients a i and b jThe output device 210 can display a characteristic curve of the frequency response on the display screen of the display unit 219. For example, the software described in the first example can be used to analyze the frequency response of a transmission file: https: / / jp.mathworks.com / help / signal / ug / frequency-renpose.htm1; https: / / jp.mathworks.com / help / signal / ref / freqz.html; https: / / docs.scipy.org / doc / scipy~0.19.1 / reference / generated / scipy.signal.freqz.html; and https: / / wiki.octave.org / Control_package.
[0157] The first to fourth examples of the output device and the control device in the first embodiment of the present invention have been described above, followed by a second embodiment and a third embodiment. (Second embodiment)
[0158] In the first embodiment, the output device 200 is connected to the servo control device 300 and the machine learning device 100, and performs the forwarding of information between the machine learning device 100 and the servo control device 300 and the control of the operations of the servo control device 300 and the machine learning device 100. The present embodiment describes a case in which the output device is connected only to the machine learning device. Fig. Figure 19 is a block diagram showing an example of the configuration of a control device according to the second embodiment of the present invention. The control device 10A includes the machine learning device 100, an output device 200A, the servo control device 300, and the servo motor 400. Compared to the one in Fig. The output device 200 shown in Figure 8 does not include the information acquisition unit 217 and the information output unit 218.
[0159] Because the output device 200A is not connected to the servo control device 300, the output device 200A does not perform the forwarding of information between the machine learning device 100 and the servo control device 300, nor does it send and receive information with the servo control device 300. Specifically, in step S31, the output device 200A executes the learning program start instruction, in step S33 the output of the physical parameters, and in step S35 the relearning instruction, which is described in Fig. 10 are shown, but they do not include the other operations (e.g., steps S32 and S34) that are in Fig. Figure 10 does not execute. Because the output device 200A is not connected to the servo control device 300, the operation of the output device 200A is reduced in this way, and consequently the configuration of the device can be simplified. (Third embodiment)
[0160] Although in the first embodiment the output device 200 is connected to the servo control device 300 and the machine learning device 100, in the present embodiment a case is described in which an adjustment device is connected to the machine learning device 100 and the servo control device 300, and in which the output device is connected to the adjustment device. Fig. Figure 20 is a block diagram showing an example of the configuration of a control device according to the third embodiment of the present invention. The control device 10B includes the machine learning device 100, an output device 200A, the servo control device 300, and the adjustment device 500. Although the in Fig. The output device 200A shown in section 20 has the same configuration as the one in Fig. In the output device 200A shown in Figure 19, the information acquisition unit 211 and the information output unit 212 are not connected to the machine learning device 100, but to the setting device 700. The setting device 500 is configured such that the drawing unit 213, the actuation unit 214, the display unit 219, and the operation unit 200 are located in the output device 200 according to Fig. 8 are omitted.
[0161] Although the in Fig. 20 output device 200A shown, as in the one in Fig. The output device 200A of the second embodiment shown in Figure 19 not only displays the learning program start instruction in step S31, the output of the physical parameters in step S33, and the fine-tuning instructions for the parameters in step S34, which are shown in Fig. In addition to the relearning instruction shown in Figure 10, which is also executed in step S35, these operations are performed by the setting device 700. The setting device 500 performs the forwarding of information between the machine learning device 100 and the servo control device 300. The setting device 500 forwards the learning program start instruction and the like to the machine learning device 100, which is executed by the output device 200A, and outputs the start instructions to the machine learning device 100. In this way, compared to the first embodiment, the function of the output device 200 is divided between the output device 200A and the setting device 500, consequently reducing the operation of the output device 200A and thus simplifying the device configuration.
[0162] Although the embodiments and examples according to the present invention have been described above, the servo control unit of the servo control device described above, the components contained in the machine learning device and the output device can be implemented by hardware, software, or a combination thereof. The servo control program executed by the cooperation of the components contained in the servo control device described above can also be implemented by hardware, software, or a combination thereof. Here, they are implemented by software means, which are realized when a computer reads and executes a program.
[0163] Programs can be stored on various types of non-transient computer-readable storage media and can be fed into a computer. Non-transient computer-readable media include various types of physical storage media. Examples of non-transient computer-readable media include magnetic recording media (e.g., a flexible disk and a hard disk drive), magneto-optical recording media (e.g., a magneto-optical disk), CD-ROM (read-only memory), CD-R, CD-R / W, semiconductor memory (e.g., a mask ROM, a PROM (programmable ROM), an EPROM (erasable PROM), a flash ROM, and RAM (read-write memory)).
[0164] The embodiment and example described above are a preferred embodiment and example of the present invention. However, the scope of protection of the present invention is not limited to this embodiment and example; rather, the present invention can be embodied in various modifications without departing from the inventive concept of the present invention. Although Fig. 9B shows the frequency response of the notch filter and the Fig. 15 and Fig. Figure 16 shows the characteristic curve of the frequency response of the IIR filter and the like. For example, the time response of the notch filter and the characteristic curve of the time response of the IIR filter and the like may be shown. The time response examples include a step response when a step-like input is given, an impulse response when an impulse-like input is given, and a slope response when an input is applied from a state where it is not changing to a state where it is changing at a constant rate. The step response, the impulse response, and the slope response can be determined from the transfer function, which includes the center angular frequency ωn, the fractional bandwidth ζ, and the attenuation coefficient R. <Variationen, wo die Ausgabevorrichtung in der Servosteuerungsvorrichtung oder der Vorrichtung zum maschinellen Lernen enthalten ist>
[0165] In the embodiments discussed above, the example where the machine learning device 100, the output device 200 or 200A, and the servo control device 300 are configured as the control device, and the example where the output device 200 is divided into the output device 200A and the adjustment device 500 and is provided in the control device, are described. Although in these examples the machine learning device 100, the output device 200 or 200A, the servo control device 300, and the adjustment device 500 are configured with separate devices, one of these devices can be integrated with another device. For example, part or all of the function of the output device 200 or 200A can be implemented with the machine learning device 100 or the servo control device 300.The output device 200 or 200A can be provided outside the control device, which is designed with the machine learning device 100 and the servo control device 3. <Freiheit in der Systemkonfiguration>
[0166] Fig. 21 is a block diagram illustrating a control device according to a further embodiment of the present invention. As in Fig. As shown in Figure 21, the control device 10C contains n machine learning devices 100-1 to 100-n, output devices 200-1 to 200-n, n servo control devices 300-1 to 300-n, servo motors 400-1 to 400-n, and a network 600. n is an arbitrarily chosen natural number. Each of the n machine learning devices 100-1 to 100-n corresponds to the one shown in Figure 21. Fig. 5 illustrated device 100 for machine learning. The output devices 200-1 to 200-n correspond to the one in Fig. Output device 210 shown in 8 or the one in Fig. 19 shown 200A. Each of the n servo control devices 300-1 to 300-n corresponds to the one in Fig. 2 or Fig. 11 servo control device 300 shown. The output device 200A and the adjustment device 500, which are in Fig. The 20 shown correspond to output devices 200-1 to 200-n.
[0167] Here, the output device 200-1 and the servo control device 300-1 are configured as a one-to-one pair, connected so that they can communicate with each other. The output devices 200-2 to 200-n and the servo control devices 300-2 to 300-n are also connected to the output device 200-1 and the servo control device 300-1. Although in Fig. Since 21 pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n are connected via network 600, the output device and the servo control device in each of the n pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n can be directly connected via a connection interface. For example, several pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n can be installed in the same factory, or they can each be installed in different factories.
[0168] The Network 600 is, for example, a LAN (local area network) set up within a factory, the internet, a public telephone network, or a combination thereof. The specific communication method used within the Network 600, whether a wired or wireless connection is used, and similar aspects are not particularly restricted.
[0169] Although in the control device described above after Fig.21. Since the output devices 200-1 to 200-n and the servo control devices 300-1 to 300-n are connected as one-to-one pairs so that they can communicate with each other, a configuration can be used, for example, in which one output device 200-1 is connected via network 600 to several servo control devices 300-1 to 300-m (m < n or m = n) so that they can communicate with each other, and in which a machine learning device connected to the one output device 200-1 performs the machine learning on the individual servo control devices 300-1 to 300-m. In this case, a distributed processing system can be used in which the respective functions of the machine learning device 100-1 are, if necessary, distributed across several servers.The machine learning functions of device 100-1 can be implemented using a virtual server or similar cloud-based solution. If there are multiple machine learning devices 100-1 to 100-n, each corresponding to multiple servo control devices 300-1 to 300-n of the same type, specification, or series, these devices can be configured to share the learning results. In this way, another optimal model can be constructed. EXPLANATION OF REFERENCE SYMBOLS 10, 10A, 10B Control device 100, 110 Machine learning device 200, 200A, 210 Output device 211 Information Acquisition Unit 212 Information output unit 213 character units 214 Actuating unit 215 Control unit 216 storage units 217 Information acquisition unit 218 Information output unit 219 Display unit 220 operating unit 300, 310 servo control device 400, 410 servo motor 500 Adjustment device 600 network
Claims
[1] Output device (200, 200A, 210) comprising the following: an information acquisition unit (211) that is provided by a machine learning device (200, 210), which applies machine learning to a servo control device (300, 310) for controlling a servo motor (400, 410) that drives an axis of a machine tool, a robot or a drive, executes, a parameter or a first physical quantity of a constituent element (311, 312, 313, 314, 315) of the servo control device (300, 310) which is being learned or has been learned; an output unit (216, 219) which outputs at least one of any of the detected first physical quantities and a second physical quantity determined from the detected parameter, a characteristic curve of the time response of the constituent element of the servo control device (300, 310) and a characteristic curve of the frequency response of the constituent element of the servo control device (300, 310), wherein the time response characteristic and the frequency response characteristic are determined with the parameter, the first physical quantity or the second physical quantity; wherein the output unit (216, 219) includes a display unit (219) which, on a display screen, shows the progress of the machine learning of the first physical quantity or the second physical quantity while the machine learning is being performed, the time response characteristic or the frequency response characteristic; and where the first and second physical quantities are any of an inertia, mass, viscosity, stiffness, resonance frequency, damping center frequency, damping rate, damping frequency range, time constant, cutoff frequency or a combination thereof. [2] Output device (200, 200A, 210) according to claim 1, wherein an instruction to set the parameter or the first physical quantity of the constituent element of the servo control device (300, 310) based on the first physical quantity, the second physical quantity, the time response characteristic or the frequency response characteristic is provided to the servo control device (300, 310). [3] Output device (200, 200A, 210) according to one of claims 1 to 2, wherein a machine learning instruction is provided to the machine learning device (200, 210) to perform machine learning of the parameter or first physical quantity of the constituent element of the servo control device (300, 310) based on the first physical quantity, the second physical quantity, the time response characteristic or the frequency response characteristic by changing or selecting a learning area. [4] Output device (200, 200A, 210) according to one of claims 1 to 3, wherein an evaluation function value used for machine learning during the learning of the device (200, 210) is output. [5] Output device (200, 200A, 210) according to one of claims 1 to 4, wherein the information about a position error which is output by the servo control device (300, 310) is output. [6] Output device (200, 200A, 210) according to any one of claims 1 to 5, wherein the parameter of the constituent element of the servo control device (300, 310) is a parameter of a model of a mathematical formula or a filter. [7] Output device (200, 200A, 210) according to claim 6, wherein the model of a mathematical formula or the filter is contained in a velocity feedforward processing unit (328) or a position feedforward processing unit (329) and the parameter contains a coefficient in a transfer function of the filter. [8] Control device comprising: the output device (200, 200A, 210) according to any one of claims 1 to 7; the servo control device (300, 310) that controls the servo motor (400, 410) that drives the axis of the machine tool, robot or industrial machine; and the machine learning device (200, 210) which performs the machine learning on the servo control device (300, 310). [9] Control device according to claim 8, wherein the output device (200, 200A, 210) is included in one of the servo control device (300, 310) and the machine learning device (200, 210). [10] Method for outputting a learning parameter of an output device (200, 200A, 210) which is machine learned in a device (200, 210) for machine learning for a servo control device (300, 310) which controls a servo motor (400, 410) for driving an axis of a machine tool, robot or industrial machine, wherein the method comprises: Acquisition by the device (200, 210) for machine learning of a parameter or a first physical quantity of a constituent element of the servo control device (300, 310) that is being learned or has been learned; Output at least one of any of the detected first physical quantity and a second physical quantity determined from the detected parameter, a characteristic curve of the time behavior of the constituent element of the servo control device (300, 310) and a characteristic curve of the frequency response of the constituent element of the servo control device (300, 310); Determining the characteristic curve of the time response and the characteristic curve of the frequency response with the parameter, the first physical quantity or the second physical quantity; and wherein the output includes displaying the progress of the machine learning of the first physical quantity or the second physical quantity while the machine learning is being performed, the time response characteristic or the frequency response characteristic; and where the first and second physical quantities are any of an inertia, mass, viscosity, stiffness, resonance frequency, damping center frequency, damping rate, damping frequency range, time constant, cutoff frequency or a combination thereof.
Citation Information
Patent Citations
machine learning apparatus, servo control apparatus, servo control system and machine learning method
DE102018203702A1
adjustment device and adjustment method
DE102018205015A1
Signal converter and signal conversion method using the same
JP1999031139A
Simulation device, simulation method, control program and recording medium
US20170262573A1
Control device and control method
WO2018151215A1