Output device, control device and method for outputting a rating function value

DE102019216190B4Active Publication Date: 2025-10-30FANUC LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
DE102019216190
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2018-10-29
Filing Date
2019-10-21
Publication Date
2025-10-30
Estimated Expiration
2039-10-21

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Output device (200, 200A) which features: an information acquisition unit (201) that is provided by a machine learning device (100) which performs machine learning on a servo control device (300) for controlling a servo motor (400) that drives an axis of a machine tool, robot or industrial machine, (i) a parameter or a first physical quantity of a component of the servo control device (300) that is or has been machine learned, and (ii) a valuation function value, and an output unit (205 and 209, 205 and 206) that outputs information indicating a relationship between (i) the detected parameter, the first physical quantity or a second physical quantity determined from the parameter, and (ii) the evaluation function value, wherein the output unit (205 and 209) includes a display unit (209) which displays on a screen a graph based on the information specifying the relationship between (i) the parameter, the first physical quantity or the second physical quantity and (ii) the evaluation function value, and wherein the parameter is a coefficient in a transfer function of the component of the servo control device (300) and an instruction to change the order of the coefficient is provided based on the information from the servo control device (300).
Need to check novelty before this filing date? Find Prior Art

Description

Background of the invention; Field of the invention

[0001] The present invention relates to an output device, a control device, and a method for outputting an evaluation function value, and in particular to an output device which receives from a machine learning device performing machine learning on a servo control device for controlling a servo motor a parameter or a first physical quantity of a component of the servo control device that is machine learned or has been machine learned, and an evaluation function value, and which outputs a relationship between the parameter, the first physical quantity, or a second physical quantity determined from the parameter, and the evaluation function value to a control device that includes such an output device, and to a method for outputting the evaluation function value. Related technology

[0002] As a technology related to the present invention, patent document 1 discloses, for example, a signal converter comprising an output unit which uses a method for learning a pattern of multiplication coefficients with a machine learning means to determine a predetermined pattern of multiplication coefficients, which uses the pattern of multiplication coefficients to perform a digital filtering operation, and which displays an output of a digital filter.

[0003] In particular, patent document 1 discloses that the signal converter includes a signal input unit, a process processing unit which has the function for characterizing signal data based on input signal data, and the output unit which displays an output from the process processing unit, that the process processing unit includes an input file, a learning means, a digital filter and a parameter setting means, and that in the learning means the method for learning a pattern of multiplication coefficients is used with the means of machine learning to determine the intended pattern of multiplication coefficients. Patent document 1: Unexamined Japanese patent application, publication no. JP H11-31139A JP 2004 - 322 224 A relates to a robot control device for allowing a tool held by a robot to follow a machining path with high accuracy. DE 10 2019 209 104 A1 relates to an output device, control device and output method for a valuation function value. DE 10 2019 216 081 A1 relates to an output device, control device and method for outputting a learning parameter. D4 DE 10 2018 205 015 A1 refers to adjusting device and adjusting procedure. D5 DE 692 20 561 T2 relates to the control of a servo motor used for a machine tool or the like. D6 WO 2014 / 194 000 A1 refers to machine learning and, in particular, a dynamic user interface for machine learning results. Overview of the invention

[0004] Although patent document 1 displays the output from the operation processing unit, a pattern learned by machine learning is unfortunately not output, and therefore a user, such as an operator, cannot verify the progress or result of the machine learning. If a parameter of a component of a servo control device that controls a servo motor driving the axis of a machine tool, robot, or industrial machine is machine learned using a machine learning device, an operator cannot verify the progress or result of the machine learning because the parameter and an evaluation function value used in the machine learning device are not publicly displayed.Even when the evaluation function value is displayed, it is difficult for the operator to understand a machine characteristic curve from the evaluation function value.

[0005] An objective of the present invention is to provide: an output device that captures a parameter or a first physical quantity of a component of a servo control device that has been learned, and an evaluation function value that can verify the progress or result of machine learning from information that indicates a relationship between the parameter, the first physical quantity or a second physical quantity determined from the parameter, and the evaluation function value, and that outputs information that makes it possible to understand a machine characteristic curve from a first or second physical quantity; a control device that includes such an output device; and a method for outputting the evaluation function value.

[0006] The problem is solved by the subject matter of the independent patent claims. (1) An output device (for example, an output device 200, 200A described below) according to the present invention comprises: an information acquisition unit (for example, an information acquisition unit 201 described below) which receives from a machine learning device (for example, a machine learning device 100 described below) performing machine learning on a servo control device (for example, a servo control device 300 described below) for controlling a servo motor (for example, a servo motor 400 described below) that drives the axis of a machine tool, robot, or industrial machine, a parameter or a first physical quantity of a component of the servo control device that is machine learned or has been machine learned, and an evaluation function value; and an output unit (for example, a control unit 205 and a display unit 209, a control unit 205 and a storage unit 206, which are described below) that outputs information indicating a relationship between the detected parameter, the first physical quantity or a second physical quantity determined from the parameter, and the evaluation function value. (2) In the output device of (1) described above, the output unit may include a display unit which displays on a screen the information indicating the relationship between the parameter, the first physical quantity or the second physical quantity and the evaluation function value. (3) In the output device of (1) or (2) described above, the parameter is a coefficient in a transfer function of the component of the servo control device, and the output device can provide an instruction to change the order of the coefficient based on the information from the servo control device. (4) In the output device of one of the above described (1) to (3), an instruction to change or select the parameter of the component of the servo control device or a search area of ​​the machine learning of the first physical quantity can be provided to the machine learning device based on the information. (5) In the output device of one of the above described (1) to (4), the parameter of the component of the servo control device may include a parameter of a mathematical formula model or a filter. (6) In the output device of the above described (5) the mathematical formula model or the filter may be included in a velocity feedforward processing unit or a position feedforward processing unit, and the parameter may include a coefficient in a transfer function of the filter. (7) A control device according to the present invention comprises: the output device of one of the above described (1) to (6); the servo control device that controls the servo motor that drives the axis of the machine tool, robot, or industrial machine; and the machine learning device that performs the machine learning on the servo control device. (8) In the control device described above (7), the output device may be included in the servo control device or in the machine learning device. (9) A method for outputting an evaluation function value from an output device according to the present invention, used in the machine learning of a machine learning device that performs machine learning on a servo control device for controlling a servo motor that drives the axis of a machine tool, robot or industrial machine, comprises: acquiring a parameter or a first physical quantity of a component of the servo control device that is being or has been machine learned, and the evaluation function value from the machine learning device; and outputting information indicating a relationship between the acquired parameter, the first physical quantity or a second physical quantity derived from the parameter, and the evaluation function value.

[0007] According to the present invention, it is possible to capture a learned parameter or first physical quantity and an evaluation function value, thereby verifying the progress or result of machine learning from information that indicates a relationship between the parameter, the first physical quantity, or a second physical quantity derived from the parameter, and the evaluation function value. Furthermore, it is possible to derive a machine characteristic curve from the first or the second physical quantity. Brief description of the drawings Fig. 1 is a block diagram that represents an example of the design of a control device according to a first embodiment of the present invention; Fig. Figure 2 is a block diagram illustrating the overall design of the control device and the design of a servo control device in the first embodiment of the present invention; Fig. Figure 3 is a diagram to illustrate the operation of an engine when one of the machining shapes is an octagon; Fig. Figure 4 is a diagram illustrating the operation of the motor when the machining mode is the mode in which the corners of an octagon are alternately replaced by arcs; Fig. Figure 5 is a block diagram representing a machine learning device according to the first embodiment of the present invention; Fig. 6 is a block diagram that provides an example of the design of an output device included in the control device according to the first embodiment of the present invention; Fig. Figure 7 is a diagram that provides an example of a screen displaying a characteristic curve diagram representing a relationship between a filter attenuation center frequency and an evaluation function value calculated from parameters related to a state S in a display unit in such a way as to correspond to the progress of a machine learning process as the machine learning is performed; Fig. Figure 8 is a diagram that provides another example of the characteristic curve diagram displayed on the screen of the display unit in the output device; Fig. Figure 9 is a frequency response diagram that represents a frequency gain characteristic curve, which is added to the screen of the display unit in the output device; Fig. Figure 10 is a three-dimensional diagram that represents a relationship between a damping center frequency, the weighting function value, and a filter attenuation rate; Fig. Figure 11 is a frequency-gain characteristic curve diagram intended to illustrate the filter attenuation rate and representing the depth of the valley of a curve; Fig. Figure 12 is a frequency-gain characteristic diagram intended to illustrate a filter band and showing the depth of the valley of the curve; Fig. Figure 13 are characteristic curve diagrams that represent the curves of the attenuation center frequency and the weighting function value when the filter attenuation rate (attenuation coefficient (attenuation)) is changed to three fixed values; Fig. Figure 14 is a three-dimensional diagram that shows a detailed relationship between the attenuation center frequency, the weighting function value, and the filter attenuation rate; Fig. 15 is a flowchart that describes the operation of the control device after the start of machine learning until the completion of machine learning, with a focus on the output device; Fig. 16 is a flowchart that describes the operation of the output device after an instruction to complete the machine learning; Fig. Figure 17 is a block diagram illustrating an example of the design of a control device according to a second embodiment of the present invention; Fig. 18 is a block diagram illustrating an example of the design of a control device according to a third embodiment of the present invention; and Fig. Figure 19 is a block diagram depicting a control device with a further design. Detailed description of the invention

[0008] Embodiments of the present invention are described in detail below with reference to the drawings. (First embodiment)

[0009] Fig. Figure 1 is a block diagram illustrating an example of the design of a control device according to a first embodiment of the present invention. The diagram shown in Figure 1 is a block diagram illustrating the design of a control device according to a first embodiment of the present invention. Fig. The control device 10 shown in Figure 1 includes a machine learning device 100, an output device 200, a servo control device 300, and a servo motor 400. The control device 10 drives a machine tool, robot, industrial machine, or the like. The control device 10 can be provided separately from a machine tool, robot, industrial machine, or the like, or it can be integrated into a machine tool, robot, industrial machine, or the like. The machine learning device 100 receives information from the output device 200 that is used in machine learning. Examples of this information include control commands, such as a position command and a speed command, which are input to the servo control device 300, and servo information, such as a position error, which is output by the servo control device 300.The machine learning device 100 also acquires parameters (for example, coefficients in the transfer function of a speed feedforward processing unit) from components of the servo control device 300 from the output device 200. Instead of the parameters of components of the servo control device 300, the machine learning device 100 can acquire physical quantities (for example, a damping center frequency, a bandwidth, and a damping coefficient (damping) associated with the parameters) (where the physical quantities correspond to the first set of physical quantities).The machine learning device 100 learns the parameters or physical quantities of components of the servo control device 300 based on the input information, in order to output the parameters or physical quantities that are being or have been machine learned, and an evaluation function value used in machine learning, to the output device 200.

[0010] The output device 200 receives the control commands, such as the position command and the speed command, that are input into the servo control device 300, and the servo information, such as the position error, that is output by the servo control device 300, and outputs this information to the machine learning device 100. The output device receives the parameters or physical quantities that are being or have been machine-learned from the machine learning device 100 and feeds them to the servo control device 300.The output device 200 receives the parameters or physical quantities that are being or have been machine-learned from the machine learning device 100 and outputs information indicating a relationship between the parameters (for example, coefficients in the speed feedforward processing unit) or values ​​(for example, the center frequency, bandwidth, and attenuation coefficient, which serve as secondary physical quantities) calculated from the parameters and the evaluation function value. Examples of the output method include a screen display in a liquid crystal display device, printing on paper using a printer or the like, storing in a storage unit such as memory, and outputting an external signal through a data transmission unit.A user, such as an operator, controls the output device 200 based on the information it outputs, for example, to change the order of the coefficients in the transfer function of the speed feedforward processing unit or the search range of the machine learning. To change the order of the coefficients in the transfer function of the speed feedforward processing unit or the search range of the machine learning, the output device 200 sends adaptation information to the servo control device 300 or the machine learning device 100.

[0011] As described above, the output device 200 has the function for passing information (such as control commands, parameters and servo information) between the machine learning device 100 and the servo control device 300, the function for outputting information that specifies the relationship between the parameters or the values ​​calculated from the parameters and the evaluation function value, and the adaptation function for outputting adaptation information to control the operations of the machine learning device 100 and the servo control device 300.

[0012] The servo control device 300 issues a current command based on control commands such as position and speed commands to control the rotation of the servo motor 400. For example, the servo control device 300 includes the speed feedforward processing unit, represented by the transfer function, which contains the coefficients that are machine-learned by the device 100 for machine learning. The servo motor 400 drives the axis of a machine tool, robot, or industrial machine. The servo motor 400 is, for example, integrated into a machine tool, robot, or industrial machine. The servo motor 400 outputs a position and / or speed detection value as feedback information to the servo control device 300.

[0013] The individual designs of the control device 10 in the first embodiment are described in more detail below.

[0014] Fig. Figure 2 is a block diagram showing the overall design of the control device 10 and the design of the servo control device 300 in the first embodiment.

[0015] First, the servo control device 300 is described. As in Fig. As shown in Figure 2, the servo control device 300 includes as components a subtractor unit 301, a position control unit 302, an adding unit 303, a subtractor unit 304, a speed control unit 305, an adding unit 306, an integrator 307, the speed feedforward processing unit 308 and a position feedforward processing unit 329.

[0016] The position command is issued to the subtractor 301, the speed feedforward processing unit 308, the position feedforward processing unit 309, and the output device 200. The position command is generated by a higher-level device based on a program for operating the servo motor 400. The servo motor 400 is, for example, included in a machine tool. When a table in the machine tool, on which a workpiece (a work) is mounted, is moved in the direction of an x-axis and in the direction of a y-axis, the servo control device 300 and the servo motor 400, which are located in Fig. The servo control device 300 and the servo motor 400 are provided separately for the x-axis and y-axis directions, respectively. When the table is moved in the directions of three or more axes, the servo control device 300 and the servo motor 400 are provided in the respective axis directions. A feed rate is specified in the position command to achieve a machining operation as defined by a machining program.

[0017] The subtractor 301 determines a difference between a position command value and the feedback detection position and outputs the difference as a position error to the position control unit 302 and the output device 200. The position control unit 302 outputs a value obtained by multiplying the position error by a position gain Kp as a velocity command value to the adder 303. The adder 303 adds the velocity command value and the output value (a position feedforward term) of the position feedforward processing unit 309 to output the resulting value as a forward-controlled velocity command value to the subtractor 304.The subtracting device 304 determines a difference between the output of the adding device 303 and a feedback velocity detection value and outputs the difference as a velocity error to the velocity control unit 305.

[0018] The speed control unit 305 adds a value obtained by multiplying the speed error by an integral gain K1v and integrating the resulting value, and a value obtained by multiplying the speed error by a proportional gain K2v, and outputs the resulting value as a torque command value to the adding unit 306. The adding unit 306 adds the torque command value and the output value (speed feedforward term) of the speed feedforward processing unit 308, outputs the resulting value via a current control unit (not shown) as a forward-controlled torque command value to the servo motor 400, and thereby drives the servo motor 400.

[0019] The angular position of the servomotor 400 is detected by a rotary encoder belonging to the servomotor 400, which serves as a position detection unit, and the speed detection value is input as speed feedback into the subtractor 304. The speed detection value is integrated with the integrator 307 as the position detection value, and the position detection value is input as position feedback into the subtractor 301.

[0020] The speed feedforward processing unit 308 performs speed feedforward processing on the position command and outputs the result of the processing as a speed feedforward term to the adding device 306. The transfer function of the speed feedforward processing unit 308 is a transfer function F(s), which is represented by an expression 1 (hereafter represented as mathematical formula 1). The optimal values ​​of the coefficients a i and b j (0≤i≤m, 0≤i≤n, where m and n are natural numbers) in expression 1 are machine learned in the machine learning device 100. F(s)=b0+b1s+b2s2+⋯+bnsna0+a1s+a2s2+⋯+amsm

[0021] The position feedforward processing unit 309 differentiates the position command value and multiplies the resulting value by a constant α, outputting the result of this processing as a position feedforward term to the adding device 303. The servo control device 300 is designed as described above. The device 100 for machine learning is described below.

[0022] The machine learning device 100 executes the specified machining program (hereinafter also referred to as the "learning machining program") to learn the coefficients in the transfer function of the speed feedforward processing unit 308. A machining shape specified by the learning machining program is, for example, an octagon or a shape in which the vertices of an octagon are alternately replaced by arcs. The machining shape specified by the learning machining program is not limited to these machining shapes and can be a different machining shape.

[0023] Fig. Figure 3 is a diagram illustrating the operation of a motor when the machining shape is an octagon. Fig. Figure 4 is a diagram illustrating the operation of the motor when the machining configuration is one in which the corners of an octagon are alternately replaced by arcs. Fig. 3 and Fig. 4. It is assumed that the table is moved in the direction of the x-axis and the direction of the y-axis, so that the workpiece (the work) is machined clockwise.

[0024] If it is, as in Fig. As shown in Figure 3, where the machining shape is an octagon, the rotational speed of the motor that moves the table in the y-axis direction is reduced at corner position A1, while the rotational speed of the motor that moves the table in the x-axis direction is increased. At corner position A2, the direction of rotation of the motor that moves the table in the y-axis direction is reversed, and the motor that moves the table in the x-axis direction rotates in the same direction and at a constant speed from position A1 to position A2 and from position A2 to position A3. At corner position A3, the rotational speed of the motor that moves the table in the y-axis direction is increased, while the rotational speed of the motor that moves the table in the x-axis direction is decreased.At corner position A4, the direction of rotation of the motor that moves the table in the direction of the x-axis is reversed, and the motor that moves the table in the direction of the y-axis is rotated in the same direction and at constant speed from position A3 to position A4 and from position A4 to the next corner position.

[0025] If the processing method is the method in which the corners of an octagon are alternately replaced by arcs, as in Fig. As shown in Figure 4, the rotational speed of the motor that moves the table in the y-axis direction is reduced at corner position B1, and the rotational speed of the motor that moves the table in the x-axis direction is increased. At corner position B2, the direction of rotation of the motor that moves the table in the y-axis direction is reversed, and the motor that moves the table in the x-axis direction rotates from position B1 to position B3 in the same direction and at a constant speed. This differs from the case described in Fig. Since the machining shape shown in Figure 12 is an octagon, the motor, which moves the table in the direction of the y-axis, is gradually reduced in speed towards position B2 so that the machining shape of the arc around position B2 is formed, the rotation is stopped at position B2 and the rotational speed is gradually increased after position B2.

[0026] At corner position B3, the rotational speed of the motor moving the table along the y-axis is increased, while the rotational speed of the motor moving the table along the x-axis is decreased. At arc position B4, the direction of rotation of the motor moving the table along the x-axis is reversed, and the table is moved in a linear inversion along the x-axis. The motor moving the table along the y-axis rotates in the same direction and at a constant speed from position B3 to position B4, and from position B4 to the next corner position. The motor moving the table along the x-axis gradually decreases in speed as it approaches position B4, forming the machining shape of the arc around position B4. It stops rotating at position B4 and then gradually increases in speed after reaching position B4.

[0027] In the present embodiment, as described above, vibrations are evaluated in the linear control device of the machining tool when the rotational speed is changed between position A1 and position A3 and between position B1 and position B3. These vibrations are determined by the machine learning program, and their influence on the position error is verified. This process also allows machine learning to optimize the coefficients in the transfer function of the speed feedforward processing unit 308. Although this method is not used in the present embodiment, a run-on delay (inertial operation) generated when the direction of rotation is reversed between position A2 and position A4 and between position B2 and position B4 of the machining tool is evaluated, and its influence on the position error can also be verified.The machine learning for optimizing the coefficients in the transfer function is not subject to any particular restriction to the speed feedforward processing unit and can, for example, also be applied to the position feedforward processing unit or to a current feedforward processing unit that is provided when current feedforward is performed on the servo control device.

[0028] The following section describes the machine learning device 100 in more detail. Although the following discussion describes a case in which the machine learning device 100 performs reinforcement learning, the learning performed by the machine learning device 100 is not specifically limited to reinforcement learning, and the present invention can, for example, also be applied to a case in which supervised learning is performed.

[0029] Before describing individual functional blocks included in the machine learning device 100, the basic mechanism of reinforcement learning will first be described. An agent (which in the present embodiment corresponds to the machine learning device 100) observes the state of an environment and selects a specific action, and the environment is modified based on this action. When the environment is modified, an arbitrary reward is provided, and in this way, the agent learns to select (decide on) a better action. While supervised learning provides a perfect response, in reinforcement learning, the reward is often a fragment value based on the modification of a part of the environment. Therefore, the agent learns to select an action that maximizes the total reward in the future.

[0030] In this way, during reinforcement learning, the action is learned, and consequently, a procedure for learning a suitable action with respect to an interaction provided by the action to the environment is learned; that is, a procedure for performing learning to maximize rewards obtained in the future is learned. In the present embodiment, this indicates, for example, that action information is selected to reduce the positional error; that is, an action that has an impact on the future can be captured.

[0031] Although any learning method can be used here as reinforcement learning, the following discussion describes as an example a case involving Q-learning, which is a method for learning a value Q(S,A) to select an action A for the state S of a given environment. One goal of Q-learning is to select the action A with the highest value Q(S,A) as the optimal action from among the actions A that can be taken in a given state S.

[0032] When Q-learning is first initiated, a correct value of Q(S, A) is not found at all in any combination of state S and action A. Consequently, the agent selects various actions A in a given state S and chooses a better action based on the reward provided for action A at that time in order to learn the correct value Q(S, A).

[0033] Since the total reward that can be achieved in the future is to be maximized, the final goal is to make Q (S, A) = E[Σ(γ t )r t ] to achieve. Here, E[] represents an expected value, t represents a time, γ represents a parameter described below called the discount rate, rt represents a reward at time t, and Σ represents a sum at time t. The expected value in this formula is the value expected when the state is changed according to the optimal action. However, since it is not clear in the Q-learning process which action is optimal, various actions are performed, and in this way, reinforcement learning takes place while a search is conducted. The formula for updating the value Q(S, A), as described above, can be represented, for example, by expression 2 below (hereafter referred to as mathematical formula 2). Q(St+1,At+1)←Q(St,At)+α(rt+1+γmaxAQ(St+1,A)−Q(St,At))

[0034] In expression 2 described above, St represents the state of the environment at time t, and A t represents an action at time t. The state is determined by the action A. t into the t+1 changed. r t+1 represents a reward that can be achieved by changing the state. A term that includes max is achieved by multiplying a Q-value when an action A, whose Q-value at that time is in state S, is performed. t+1 The highest value is selected by γ. Here, γ is a parameter 0 < γ ≤ 1 and is called the discount rate. α is a learning coefficient, and it is assumed to fall within the range 0 < α ≤ 1.

[0035] The expression 2 described above gives a procedure for updating the value Q (St, At) of action A t in state S t based on the reward r t+1on, which as a result of an attempt A t is returned. This update formula specifies that if the value max a Q (S t+1 , A) the best action in the subsequent state S t+1 through action A t greater than the value Q (S t , A t ) of Action A t in state S t is, Q (St, A t ) is increased, whereas Q (S t , A t ) is reduced when the value is max a Q (S t+1 , A) less than the value Q (S t , A t ). In other words, the value of a particular action in a given state is approximated to the value of the best action in the subsequent state that results from it. Although the difference between these two actions depends on the discount rate γ and the reward r. t+1When a change is made, a mechanism is designed such that the value of the best action in a given state is essentially propagated to the value of an action in the previous state.

[0036] In Q-learning, a procedure is used in which a table of Q(S, A) is generated for all state-action pairs (S, A), and this learning process is then performed. However, since there is a large number of states, and the values ​​of Q(S, A) need to be determined for all state-action pairs, it is likely that the Q-learning process will take longer to complete.

[0037] Therefore, a well-known technology called DQN (Deep-Q Network) can be used. Specifically, a value function Q is designed using a suitable neural network, a parameter for the neural network is adjusted, and the value function Q is approximated by the suitable neural network, resulting in the computation of the value Q(S, A). By using the DQN, it is possible to reduce the time required to complete Q-learning. Details about the DQN are disclosed, for example, in a non-patent document below. <nichtpatentdokument>

[0038] “Human-level control through deep reinforcement learning”, written by Volodymyr Mnihl [online], [searched on 17 January 2017], Internet <URL: http: / / files.davidqiu.com / research / nature14236.pdf>

[0039] The Q-learning described above is performed by the machine learning device 100. Specifically, the machine learning device 100 learns the value Q, assuming that it is located in the servo control device 300 in a servo state of commands, feedback, and the like, including the values ​​of the individual coefficients a. i and b j (0≤i≤m, 0≤j≤n, m and n are natural numbers) in the transfer function of the speed feedforward processing unit 308, the position error of the servo control device 300, which is detected by executing the learning processing program, and the position command is about the state S and where the adjustment of the values ​​of the individual coefficients a i and b j in the transfer function of the speed control processing unit 308, in state S, action A is selected.

[0040] The machine learning device 100 performs based on the individual coefficients a i and b j In the transfer function of the speed feedforward processing unit 308, the learning processing program is executed and, between position A1 and position A3 and between position B1 and position B3, observes state information S during the processing modes described above. This information includes the servo state of commands, feedback, and the like, including the position command and position error information of the servo control device 300, in order to determine the action A. The machine learning device 100 returns a reward each time action A is performed. For example, the machine learning device 100 searches for the optimal action A through trial and error, thus maximizing the total reward in the future. In this way, the machine learning device 100 can determine the optimal action A (that is, the optimal coefficients a). i and b j in the speed feedforward processing unit 308) for the state S, which includes the servo state of commands, feedbacks and the like, including the position command and the position error of the servo control device 300, which is determined by executing the learning processing program based on the individual coefficients a i and b j in the transfer function of the speed feedforward processing unit 308. Between position A1 and position A3 and between position B1 and position B3, the directions of rotation of the servo motors for the x-axis and y-axis are not changed, and in this way the device 100 for machine learning can determine the individual coefficients a i and b j in the transfer function of the speed feedforward processing unit 308 learn when a linear operation is performed.

[0041] In other words, based on the value function Q learned by the machine learning device 100, the actions A applied to the individual coefficients a are used to generate i and b j In the transfer function of the speed feedforward processing unit 308, in the specific state S, such an action A is selected to maximize the value of Q, and in this way it is possible to apply such an action A (that is, the individual coefficients a) i and b j in the transfer function of the speed feedforward processing unit 308) to be selected so that the position error detected by running the learning processing program is minimized.

[0042] Fig. Figure 5 is a block diagram illustrating the machine learning device 100 according to the first embodiment of the present invention. To perform the reinforcement learning described above, the machine learning device 100 includes, as shown in Figure 5, the following components: Fig. Figure 5 shows a state information acquisition unit 101, a learning unit 102, an action information output unit 103, a value function storage unit 104, and an optimization action information output unit 105. The learning unit 102 includes a reward output unit 1021, a value function update unit 1022, and an action information generation unit 1023.

[0043] The state information acquisition unit 101 acquires the state S from the servo control device 300, which includes the servo state of commands, feedback and the like, including the position command and the position error of the servo control device 300, which is determined by executing the learning processing program based on the individual coefficients a i and b j The state information S is acquired in the transfer function of the speed feedforward processing unit 308 in the servo control device 300. The state information corresponds to an environmental state S in Q-learning. The state information acquisition unit 101 outputs the acquired state information S to the learning unit 102.

[0044] The coefficients a i and b j In the speed feedforward processing unit 308, the initial values ​​of the coefficients a are generated in advance by a user when Q-Learning is first initiated. In the present embodiment, the initial setting values ​​of the coefficients a i and b j In the speed feedforward processing unit 308, the user-generated values ​​are adjusted by reinforcement learning to achieve optimal results. For example, in expression 1, the initial setpoint values ​​of the coefficients a i and b j in the speed feedforward processing unit 308 to a0 = 1, a1 = 0, a2 = 0, ..., a m = 0, b0 = 1, b1 = 0, b2 = 0, ... and b n = 0. The orders m and n of the coefficients a i and b j are predetermined. In particular, it is assumed that 0 ≤ i ≤ m for a i holds true and that 0≤j≤n for b j This applies. With regard to the coefficients a i and b j If the machine tool is pre-adjusted by the operator, machine learning can be performed using values ​​that have been adjusted as initial values.

[0045] Learning unit 102 is a unit that learns the value Q (S, A) when a specific action A is selected under a specific environmental condition S.

[0046] The reward output unit 1021 is a unit that calculates a reward when action A is selected in a given state S. Here, the position error set (position error set) of the position error, which is a state variable in state S, is represented by PD(S), and a position error set, which is a state variable related to state information S' into which state S is changed by the action information A (correction of the individual coefficients a). i and b j In the speed feedforward processing unit 308), the position error value is represented by PD(S'). It is assumed that the position error value in state S is a value calculated based on a predefined evaluation function f(PD(S)). For example, if the position error is represented by e, the following functions can be applied to the evaluation function f: a function for calculating an integrated value of absolute values ​​of position errors; ∫|e|dt a function for calculating an integrated value by weighting absolute values ​​of positional errors; ∫t|e|dt a function for calculating an integrated value of a 2n-th power (n is a natural number) of absolute values ​​of positional errors; and ∫e 2n dt (n is a natural number) a function for calculating the maximum absolute value of position errors Max {|e|}.

[0047] If the evaluation function value f (PD(S')) of the position error in the servo control device 300, which is operated on the basis of the speed feedforward processing unit 308, which is related to the state information S' that is corrected by the action information A, and which has been corrected, is greater than the evaluation function value f (PD(S)) of the position error in the servo control device 300, which is operated on the basis of the speed feedforward processing unit 308, which is related to the state information S before it is corrected by the action information A, and is in the state before the correction, the reward output unit 1021 sets the value of a reward to a negative value.

[0048] Conversely, if the evaluation function value f (PD(S')) of the position error is less than the evaluation function value f (PD(S)) of the position error, the reward output unit 1021 sets the value of a reward to a positive value. If the evaluation function value f (PD(S')) of the position error is equal to the evaluation function value f (PD(S)) of the position error, the reward output unit 1021 sets the value of a reward to zero.

[0049] The negative value can be increased proportionally if the evaluation function value f(PD(S')) of the position error in state S' after action A is performed is greater than the evaluation function value f(PD(S)) of the position error in the preceding state S. In other words, the negative value is preferentially increased by the degree by which the position error value is increased. Conversely, the positive value can be increased proportionally if the evaluation function value f(PD(S')) of the position error in state S' after action A is less than the evaluation function value f(PD(S)) of the position error in the preceding state S. In other words, the positive value is preferentially increased by the degree by which the position error value is decreased.

[0050] The value function update unit 1022 performs Q-learning based on state S, action A, state S' when action A is applied to state S, and the value of a reward, calculated as described above, to update the value function Q stored in the value function memory unit 104. Updating the value function Q can be performed through online learning, batch learning, or mini-batch learning. Online learning is a learning method in which a specific action A is applied to the current state S, and in this way, the value function Q is immediately updated each time state S changes to the new state S'.Batch learning is a learning method in which a specific action A is applied to the current state S, repeatedly changing state S to the new state S' so that data is collected for learning. All collected data is then used to update the value function Q. Mini-batch learning, on the other hand, is an intermediate learning method between online learning and batch learning, in which the value function Q is updated each time a certain amount of data is stored for learning.

[0051] The action information generation unit 1023 selects action A in the Q-Learning process for the current state S. In the Q-Learning process, the action information generation unit 1023 generates the action information A to perform a procedure (corresponding to action A in Q-Learning) to correct the individual coefficients a. i and b j in the speed feedforward processing unit 308 of the servo control device 300, and outputs the generated action information A to the action information output unit 103. More precisely, the action information generation unit 1023 incrementally adds or subtracts, for example (for example, by about 0.01), the individual coefficients a i and b j in the speed control processing unit 308, which are included in action A, to or from the individual coefficients in the speed control processing unit, which are included in state S.

[0052] If the action information generation unit 1023 increases or decreases the individual coefficients a i and b j When applied in the speed feedforward processing unit 308, the state is changed to state S', and a plus reward (a positive reward) is returned. The action information generation unit 1023 can then initiate the subsequent action A', for example, to incrementally perform an addition or a subtraction on each coefficient a. i and b j Select in the speed feedforward processing unit 308 in the same way as in the previous action, so that the value of the position error is reduced more.

[0053] If, on the other hand, a negative reward is returned, the action information generation unit 1023 can initiate the subsequent action A', for example, to incrementally perform a subtraction or addition on the individual coefficient a. i and b j in the speed feedforward processing unit, select in a manner opposite to the previous action so that the position error is smaller than the previous value.

[0054] The action information generation unit 1023 can select action A' by a known procedure such as a greedy procedure to select the action A' whose value Q (S, A) is highest in the value of action A currently being estimated, or by an ε-greedy procedure to randomly select action A' with a low probability ε and otherwise select the action A' whose value Q (S, A) is highest.

[0055] The action information output unit 103 is a unit that outputs the action information A and the evaluation function value, which is output by the learning unit 102, to the output device 200. As described previously, the servo control device 300, through the output device 200, slightly corrects the actual state S based on the action information; that is, the currently set individual coefficients a i and b j in the speed feedforward processing unit 308, which are currently set to change the state of the subsequent state S' (that is, the corrected individual coefficients in the speed feedforward processing unit 308).

[0056] The value function storage unit 104 is a storage device that stores the value function Q. The value function Q can be stored, for example, for each state S or each action A in a table (hereinafter referred to as the action value table). The value function Q stored in the value function storage unit 104 is updated by the value function update unit 1022. The value function Q stored in the value function storage unit 104 can be shared with another machine learning device 100. When the value function Q is shared by multiple machine learning devices 100, reinforcement learning can be distributed across the respective machine learning devices 100, thereby improving the efficiency of reinforcement learning.

[0057] The optimization action information output unit 105 generates action information A (hereinafter referred to as "optimization action information") based on the value function Q, which is updated as a result of the value function update unit 1022 performing Q-learning. This action information causes the velocity feedforward processing unit 308 to perform an operation to maximize the value Q(S, A). More precisely, the optimization action information output unit 105 acquires the value function Q, which is stored in the value function memory unit 104. As described above, the value function Q is updated as a result of the value function update unit 1022 performing Q-learning. Subsequently, the optimization action information output unit 105 generates the action information based on the value function Q and outputs the generated action information to the output device 200.The optimization action information described above, like the action information output by action information output unit 103 in the Q-Learning process, contains information for correcting the individual coefficients a. i and b j in the speed feedforward processing unit 308 and the evaluation function value.

[0058] As described above, the device 100 is used for machine learning according to the present embodiment, and in this way it is possible to simplify the adjustment of the parameters in the speed feedforward processing unit 308 of the servo control device 300.

[0059] When calculating the reward value, the reward output unit 1021 can add another element besides the position error. For example, in addition to the position error, which serves as the output of the subtractor unit 301, the reward output unit 1021 can add at least one element from a position-forward-controlled velocity command, which serves as the output of the adder unit 303, a difference between the position-forward-controlled velocity command and a velocity feedback, which serves as the output of the subtractor unit 304, a velocity-forward-controlled torque command, which serves as the output of the adder unit 306, and the like, to calculate the reward value.

[0060] In the embodiment discussed above, learning on the optimization of the coefficients in the speed feedforward processing unit at the time of the linear process, in which the directions of rotation of the servo motors for the x-axis and the y-axis directions are not changed, is described in the device 100 for machine learning. However, the present embodiment is not limited to learning at the time of the linear process and can also be applied to learning on a nonlinear process.For example, if learning is performed on the optimization of the coefficients in the velocity feedforward processing unit for game correction, a difference between the position command value and the detection position output by an integrator 108 is extracted as a position error between position A2 and position A4 and between position B2 and position B4 in the processing mode described above. This position error is used as detection information, and in this way a reward is given, with the result that reinforcement learning can be carried out.Between positions A2 and A4, and between positions B2 and B4, the directions of rotation of the servo motors are reversed for the y-axis or x-axis direction, respectively. This nonlinear process creates a loop, allowing the machine learning device 100 to learn the coefficients in the transfer function of the feedforward processing unit at the time of the nonlinear process. The servo control device 300 and the machine learning device 100 have been described above. The output device 200 is described below. <Ausgabevorrichtung 200>

[0061] Fig. Figure 6 is a block diagram illustrating an example of the design of the output device 200 included in the control device 10 according to the first embodiment of the present invention. As shown in Fig. As shown in Figure 6, the output device 200 includes an information acquisition unit 201, an information output unit 202, a drawing unit 203, an operating unit 204, a control unit 205, a storage unit 206, an information acquisition unit 207, an information output unit 208, a display unit 209 and an operating unit 210.

[0062] The information acquisition unit 201 serves as an information acquisition unit that acquires the parameters and the evaluation function value from the machine learning device 100. The control unit 205 and the display unit 209 serve as output units that display a relationship between the parameters (for example, the coefficients a) using a scatter plot or the like. i and b j in the speed control processing unit) or values ​​(for example, the center frequency, the bandwidth fw, and the damping coefficient (attenuation)) calculated from the parameters, and the evaluation function value. A liquid crystal display device, a printer, or the like can be used as the display unit 209 of the output unit. The output includes storage in the storage unit 206, and in such a case, the output unit is the control unit 205 and the storage unit 206.

[0063] The output device 200 has an output function for outputting a relationship between the parameters (learning parameters) in the device 100 for machine learning that are being learned or have been learned by machine, and the evaluation function value, or a relationship between the values ​​calculated from the learning parameters and the evaluation function value.The output device 200 also has a transmission function for passing information (such as control commands like the position command and the speed command, the position error and the coefficients in the speed feedforward processing unit) from the servo control device 300 to the machine learning device 100 and information (for example, the correction information of the coefficients in the speed feedforward processing unit) from the machine learning device 100 to the servo control device 310 and an adaptation function for performing the control (for example, an instruction to start the learning program to the machine learning device and an instruction to change the search area) on the operation of the machine learning device 100.The transmission of information and the control of operations are carried out by the information acquisition units 201 and 207 and the information output units 202 and 208.

[0064] A case in which the output device 200 outputs the relationship between the values ​​calculated from the machine-learned parameters and the evaluation function value is first described with reference to Fig. 7 described. Fig. Figure 7 is a diagram that provides an example of the screen display when a scatter plot or similar representation of the relationship between the damping center frequency and the weighting function value calculated from parameters related to state S is shown in display unit 209 to reflect the progress of the machine learning process as it is performed. As in Fig. As shown in Figure 7, the screen P of display unit 209 contains columns P1, P2, and P3. In column P1, display unit 209 shows, for example, selection elements for an axis selection, parameter check, program check / editing, program start, machine learning, and the final determination. In column P2, display unit 209 shows, for example, an adaptation goal such as speed feedforward, a status (state) such as data collection, the number of trials (the total number of trials to date compared to the specified number of trials (hereinafter also referred to as the "maximum number of trials") until the machine learning is complete, and a button to select the interruption of the learning process.In column P3, for example, the display unit 209 shows a scatter plot that represents the relationship between the damping center frequency fc and the weighting function value, which are the values ​​calculated from the coefficients in the transfer function of the speed feedforward processing unit.

[0065] If the user, such as the operator, uses the control unit 204, such as a mouse or keyboard, to access the "machine learning" option in column P1 on the [document / section] Fig. When the screen P shown in Figure 7 is selected in the display unit 209, such as a liquid crystal display device, the control unit 205 executes an instruction to output the coefficients a through the information output unit 202 of the machine learning device 100. i and b j in connection with the state S in conjunction with the number of trials, the evaluation function value f (PD(S)), information about the adaptation target (learning objective) of machine learning, the number of trials, information that includes the maximum number of trials, and so on.

[0066] If the information acquisition unit 201 receives the coefficients a from the machine learning device 100 i and b j In connection with the state S in conjunction with the number of trials, the evaluation function value f (PD(S)), information about the adaptation target (learning objective) of the machine learning, the number of trials, information including the maximum number of trials, and the like, the control unit 205 stores the received information in the memory unit 206 and transmits the control to the operating unit 220. In the memory unit 206, the coefficients a i and b j and the valuation function value f (PS(S)), which contains the coefficient a i and b j They correspond to each other and are stored together.

[0067] The operating unit 220 calculates from the parameters that are machine-learned in the device 100 for machine learning, in particular the parameters (for example, the coefficients a described above). i and b j In connection with state S) at the time of reinforcement learning or after reinforcement learning, the damping center frequency fc of the velocity feedforward processing unit. The damping center frequency fc is a value (a second physical quantity) that is derived from a i and b j The operating unit 220 can calculate the bandwidth fw and the attenuation coefficient R in addition to the attenuation center frequency fc. The following discussion describes a method for calculating the attenuation center frequency fc, the bandwidth fw, and the attenuation coefficient R.

[0068] The procedure in which the operating unit 220 calculates the damping center frequency fc, the bandwidth fw and the damping coefficient R is described below using a case as an example in which the speed feedforward processing unit 308 is driven by the motor reversing characteristic (the transfer function is Js). 2 ) and the notch filter is specified. If the speed feedforward processing unit 308 is driven by the motor reversing characteristic (the transfer function is Js) 2 ) and the notch filter is specified, the transfer function F(s) represented by expression 1 is an expression model specified by the right-hand side of expression 3, and it is given as the right-hand side of expression 3 by using the center angular frequency ω, the normalized bandwidth ζ, and the attenuation coefficient R. The attenuation center frequency fc, the bandwidth fw, and the attenuation coefficient (attenuation) R are derived from the coefficients a i and b j To determine the center angular frequency ω, a normalized bandwidth ζ and the damping coefficient R are determined from expression 3, and the center frequency fc and the bandwidth fw are further determined from ω = 2πfc and ζ = fw / fc. b2s2+b3s3+b4s4a0+a1s+a2s2=Js2⋅ω2+2Rζωs+s2ω2+2ζωs+s2

[0069] Since expression 3 states that a0 = ω 2 , b 4 = J, a1 = 2ζω, b3 = 2JζRω and (b 3 / a1) = R · J and ω = 2πfc and ζ = fw / fc, the center frequency fc, the bandwidth fw and the damping coefficient R can be determined by expression 4. fc=a02π,fw=a14π,R=b3a1⋅b4

[0070] Although the above description uses the case as an example where the speed control processing unit 308 is described by the mathematical formula model of the motor reversing characteristic (the transfer function is Js) 2 The present embodiment is not subject to any special limitation to such a case, and even if the transfer function of the speed feedforward processing unit 308 is in the form of a general formula, as given in Expression 1, the attenuation center frequency fc, the bandwidth fw, and the attenuation coefficient R can be determined if the transfer function exhibits a gain valley. In general, regardless of the order of the filter, it is equally possible to determine one or more of the attenuation center frequencies fc, bandwidth fw, and attenuation coefficients R that are attenuated. Software is known that can analyze the frequency response from the transfer function, and the following, for example, can be used: https:l / jp.mathworks.com / help / signal / ug / frequency-renponse.html; https: / / jp.mathworks.com / help / signal / ref / freqz.html; https: / / docs.scipy.org / doc / scipy-0.19.1 / reference / generated / scipy.signal.freqz.html; https: / / wiki.octave.org / Control_package; and the like. The attenuation center frequency fc, the bandwidth fw and the attenuation coefficient R can be determined from the frequency response.

[0071] The operating unit 220 calculates the attenuation center frequency fc and then transmits the control signal to the control unit 205. The transfer function on the right-hand side of expression 3 is transformed into the transfer function of the speed feedforward processing unit 308, which is specified by the center frequency fc, the bandwidth fw, and the attenuation coefficient R. The parameters of the center frequency fc, the bandwidth fw, and the attenuation coefficient R are machine-learned in the machine learning device 100, with the result that the output device 200 can detect the center frequency fc, the bandwidth fw, and the attenuation coefficient R. In this case, the detected center frequency fc, the bandwidth fw, and the attenuation coefficient R serve as the initial physical quantities.

[0072] The control unit 205 stores the damping center frequency fc in the storage unit 206. The coefficients a are stored in the storage unit 206. i and b j and the valuation function value f (PD(S)), which contains the coefficient a i and b j The values ​​correspond to each other and are stored, and the control unit 205 also stores the damping center frequency fc, which is based on the coefficients a i and b j is calculated such that it is assigned to the evaluation function value f (PD(S)). The control unit 205 further determines the damping center frequency fc, where the evaluation function value has a local minimum value, stores this in the memory unit 206 and transmits the control to the drawing unit 203. If the output device 200 does not determine the physical quantities such as the damping center frequency fc and the information that defines the relationship between the velocity feedforward coefficients a i and b j Since the operating unit 210 outputs the physical quantities such as the damping center frequency fc from the velocity feedforward coefficients a, and these serve as learning parameters and specify the evaluation function value, it is not necessary to use the operating unit 210 to derive the physical quantities such as the damping center frequency fc from the velocity feedforward coefficients a. i and b j To calculate the attenuation center frequency fc, it is not necessary to determine it, and the control unit 205 transmits the control to the drawing unit 203. The drawing unit 203 generates an attenuation center frequency weighting function value scatter plot of the weighting function value f (PD(S)), which is stored such that it contains the coefficient a i and b j is assigned with regard to the damping center frequency fc, which is based on the coefficients a i and b j The system calculates and adds values ​​(here 250 Hz and 400 Hz) of the attenuation center frequency fc, where the weighting function value specifies a local minimum value, to the scatter plot to generate the image information of the attenuation center frequency-weighting function value scatter plot, and transmits the control to the control unit 205. The control unit 205 displays the attenuation center frequency-weighting function value scatter plot in column P3 on the screen shown in the diagram. Fig. The control unit 205 displays "Speed ​​feedforward" in the adaptation target element of column P2 on the screen shown in the 7. Fig. The screen shown in Figure 7, for example, displays information indicating that the speed feedforward processing unit is the target for adjustment, and if the number of trials does not reach the maximum number of trials, "Data is being collected" is displayed in the status element of column P2. Furthermore, the control unit 205 displays a ratio of the number of trials to the maximum number of trials in the number of trials element in column P2. If the information relating to the relationship between the speed feedforward coefficients a i and b j and the evaluation function value, the drawing unit 203, for example, generates a scatter plot that represents the relationship between the velocity feedforward coefficient a0 (parameter associated with the damping center frequency fc) and the evaluation function value, and the control unit 205 displays the scatter plot in column P3 on the screen in Fig. 7 shown on screen P.

[0073] The in Fig. The screen P shown in Figure 7 is an example, and the present invention is not limited to this screen. Information other than the elements illustrated above can be displayed. The display of information from some of the elements illustrated above can be omitted. Although, in the description above, the control unit 205 stores information received from the machine learning device 100 in the storage unit 206 and displays, for example, information about the attenuation center frequency weighting function value scatter plot in real time in the display unit 209, there is no limitation to this design. Design examples where a display is not generated in real time include the following. Design example 1: If the user, such as the operator, provides a display instruction, the information shown in Figure 209 is displayed in the screen. Fig. The information shown in section 7 is displayed. Design example 2: When the total number of attempts (after the start of learning) reaches a predetermined number of attempts, the information shown in Fig. 7. Information shown is displayed. Design example 3: If machine learning is interrupted or completed, the information shown in Fig. The information shown is displayed in 7 images.

[0074] Even in the design examples 1 to 3 described above, as in the real-time display operation described above, when the information acquisition unit 201 from the machine learning device 100 stores the coefficients a i and b j In connection with state S, the control unit 205 receives information about the adaptation goal (learning objective) of the machine learning, the number of trials, information including the maximum number of trials, and the like, storing the received information in the storage unit 206. Subsequently, the control unit 205 transmits the control to the operating unit 210 and the drawing unit 203. In design example 1, the user provides a display instruction; in design example 2, the total number of trials reaches a predetermined number; or in design example 3, the machine learning is interrupted or completed.

[0075] The drawing unit 203 can be used instead of the attenuation center frequency weighting function value scatter plot, as shown in Fig. Figure 8 shows how to create a figure in which the rating function value is not represented as rating points but as a characteristic curve diagram of a rating curve, and the control unit 205 can display the information in Fig. Figure 8 shown in column P3 on the in Fig. Display screen P as shown in 7.

[0076] Although the above discussion describes the example in which the attenuation center frequency weighting function value scatter plot or the characteristic curve plot of the weighting curve is displayed in column P3 on the screen P of display unit 209, a frequency response plot representing the frequency-gain characteristic of the speed feedforward processing unit 308 can be added to the scatter plot or the characteristic curve plot. For example, drawing unit 203, with operating unit 210, determines the frequency response of the speed feedforward processing unit 308 from the transfer function, which includes the center angular frequency ω, the normalized bandwidth ζ, and the attenuation coefficient R on the right-hand side of expression 3, generating a Fig. The frequency-gain characteristic curve shown in Figure 9 transmits the control to the control unit 205. The frequency response of the speed feedforward processing unit 308 can be determined from the transfer function on the right side of expression 3 using the known software described above, which can analyze a frequency response from a transfer function. The control unit 205 shows the frequency response in column P3 of the diagram shown in Figure 9. Fig. The screen shown in Figure 7 displays the frequency-gain characteristic curve diagram (which serves as the frequency characteristic) and the damping center frequency weighting function value scatter plot, or the characteristic curve diagram of the weighting curve. In this way, the user, such as the operator, can simultaneously understand the frequency-gain characteristic curve of the 308 speed feedforward processing unit. Fig. Figure 9 shows that the damping center frequency is 400 Hz.

[0077] In the embodiment discussed above, the example is described in which the scatter plot or characteristic curve diagram of the evaluation curve, which represents the relationship between the attenuation center frequency or the learning parameters and the evaluation function value, is displayed in column P3 on the screen P of display unit 209. However, the physical quantity that specifies the relationship with the evaluation function value is not limited to the attenuation center frequency, and the bandwidth fw or the attenuation coefficient R can be used instead. The bandwidth fw or the attenuation coefficient R can be added to the attenuation center frequency, and in this case, a diagram displayed in column P3 on the screen P can be a three-dimensional diagram (a 3D graph).A damping center frequency weighting function value characteristic curve diagram, in which the bandwidth fw or the damping coefficient R is changed and in which a plurality of curves indicating the relationship between the weighting function value and the damping center frequency are provided, can be found in column P3 on the [table / document]. Fig. The screen P shown in Figure 7 will be displayed. These examples are described below as Examples 1 to 3. It is understood that in individual examples below, the attenuation center frequency, the bandwidth fw, and the attenuation coefficient R can be converted into any one of the coefficients a. i and b j can be changed in the transfer function of the speed control processing unit. <Beispiel 1>

[0078] The present example 1 is an example in which a three-dimensional diagram (3D graph) in which a filter attenuation rate (a damping coefficient (attenuation)) is added to the attenuation center frequency and the weighting function value is displayed in column P3 on the screen P of display unit 209. Fig. Figure 10 is a three-dimensional diagram that illustrates the relationship between the attenuation center frequency, the weighting function value, and the filter attenuation rate. Fig. 10. The filter attenuation rate can be replaced by a filter band (bandwidth). As shown in a curve that indicates the frequency-gain characteristic, from Fig. As shown in Figure 11, the filter attenuation rate indicates the depth of the valley in the curve. This is similar to a curve representing the frequency-gain characteristic. Fig. Figure 12 shows that the filter band indicates the width of the curve's valley. The user can understand the influence of the filter attenuation rate and the attenuation center frequency on the weighting function value. <Beispiel 2>

[0079] Example 2 is an example in which characteristic curve diagrams representing the curves of the damping center frequency and the weighting function value when the filter damping rate (the damping coefficient (damping)) is changed to three fixed values ​​are displayed in column P on screen P3 of display unit 209. Fig. Figure 13 are characteristic curve diagrams that depict the curves of the attenuation center frequency and the weighting function value when the filter attenuation rate (attenuation coefficient (attenuation)) is changed to predefined values ​​(0%, 50%, and 100%). The user can understand the influence of the filter attenuation rate on the characteristic curves of the attenuation center frequency and the weighting function value. <Beispiel 3>

[0080] Example 3 is an example in which a three-dimensional diagram (3D graph) representing a relationship between the attenuation center frequency, the weighting function value and the filter attenuation rate (the attenuation coefficient (attenuation)) is displayed in column P3 on screen P. Fig. Figure 14 is a three-dimensional diagram that illustrates a further detailed relationship between the attenuation center frequency, the weighting function value, and the filter attenuation rate. Fig. 14. The filter attenuation rate can be replaced by the filter band (bandwidth).

[0081] The output function of output device 200 has been described above. The forwarding function and the adaptation function of output device 200 are described below with reference to Fig. 15 and Fig. 16 described. Fig. Figure 15 is a flowchart that illustrates the operation of the control device from the start of machine learning until the completion of machine learning, with a focus on the output device. If, in step S31, the operator uses the operating unit 204, such as a mouse or keyboard, to initiate the "Program Start" in column P1 on screen P of the output device 200, the program will then be activated. Fig. When the display unit 209 shown in Figure 7 is selected, the control unit 205 issues an instruction to start the program via the information output unit 202 to the machine learning device 100. Subsequently, the output device 200 issues a program start instruction message to the servo control device 300 to report the output of the program start instruction for learning to the machine learning device 100. In step S32, the output device 200 provides an instruction to start the learning processing program to a higher-level device, which outputs the learning processing program to the servo control device 300. Step S32 can be performed before or concurrently with step S31. When the higher-level device receives the instruction to start the learning processing program, it generates the position command and outputs it to the servo control device 300.When the machine learning device 100 receives the instruction to start the program, the machine learning device 100 begins machine learning in step S21.

[0082] In step S11, the servo control device 300 controls the servo motor 400 to obtain the parameter information (the coefficients a). i and b j ) in the speed feedforward processing unit 308 and output the information, which includes the position command and the position error, to the output device 200. Subsequently, the output device 200 outputs the parameters, the position command and the position error to the machine learning device 100.

[0083] The machine learning device 100 provides the output device 200 with the evaluation function value related to state S in conjunction with the number of trials used in a reward output unit 2021 while the machine learning process is performed in step S21, the maximum number of trials, the number of trials, and the correction information (serving as parameter correction information) of the coefficients a. i and b j in the transfer function of the speed feedforward processing unit 308. If in step S33 the “machine learning” in column P1 on the in Fig. When screen P shown in 7 is selected, the output device 200 uses the output function described above to generate a graph based on the correction information of coefficients in the transfer function of the speed feedforward processing unit 308, which is machine-learned, and the evaluation function value output by the machine learning device 100. This graph illustrates the relationship between the physical quantities (such as the center frequency fc), which are easily understood by the user, such as the operator, and the evaluation function value, and displays it in column P3 on screen P shown in the Fig. The display unit 209 shown in Figure 7 is used. In step S33, or before or after step S33, the output device 200 feeds the correction information of the coefficients in the transfer function to the speed feedforward processing unit 308 of the servo control device 310. Steps S11, S21, and S33 are repeated until the machine learning is complete.

[0084] Although the case described here is in which information relating to a figure representing a relationship between the physical quantities (such as the center frequency fc) of the coefficients in the transfer function of the speed feedforward processing unit 308 in connection with the machine-learned parameters and the evaluation function is output to the display unit 209 in real time, in the cases of Examples 1 to 3, which have already been described as examples of the case in which a display is not generated in real time, the information relating to the figure representing the relationship between the physical quantities (such as the center frequency fc) of the coefficients in the transfer function of the speed feedforward processing unit 308 and the evaluation function can be output to the display unit 209.

[0085] In step S34, output device 200 determines whether the number of trials has reached the maximum. If so, in step S35, output device 200 sends a termination instruction to machine learning device 100. If the maximum number of trials has not been reached, the process returns to step S33. In step S35, output device 200 sends a termination instruction to machine learning device 100. When machine learning device 100 receives the termination instruction in step S22, it completes the machine learning process.

[0086] The forwarding function of output device 200 has been described above. The following describes the adaptation function of output device 200. If the user, such as the operator, selects column P3 on screen P of display unit 209, the following applies: Fig. When viewing the output device 200 shown in Figure 7, either during or after machine learning, the user might want to provide an instruction to change the orders m and n of the coefficients in the speed feedforward processing unit 308 to the servo control device 300, or an instruction to change or select the search range for the device 100 for machine learning. For example, if the user wants to change the evaluation function value in the scatter plot in column P3 on the device shown in Figure 7, the user might want to provide an instruction to change the orders m and n of the coefficients in the speed feedforward processing unit 308 to the servo control device 300, or an instruction to change or select the search range for the device 100 for machine learning. Fig. Viewed from screen P shown in Figure 7, the user recognizes that, since the evaluation function value is low at 250 Hz and 400 Hz in the machine tool, it is very likely that a machine resonance occurs at these frequencies. In such a case, the user might want to specify the orders m and n of the coefficients a in expression 1. i and b j change or change the search range of the coefficients a i and b j change. During or after machine learning, the output device 200 instructs the servo control device 300 to change the orders m and n of the coefficients a. i and b j to adjust in the speed feedforward processing unit, or instructs the machine learning device 100 to perform relearning.

[0087] Fig. 16 is a flowchart that illustrates the operation of the output device after an instruction to complete the machine learning. If the user in step S35 of Fig. 15, after the output device 200 provides the completion instruction to the machine learning device 100, the evaluation function value in the scatter plot in column P3 on the in Fig. When the user sees the screen P shown in section 7, they recognize that it is very likely that the machine resonance occurs at frequencies of 250 Hz and 400 Hz, and therefore select "Change" on screen P. Fig. 7 with the operating unit 204, such as a mouse or a keyboard. The control unit 205 displays P within the screen. Fig. 7 for example the formula of the transfer function given in expression 3, and an input column for the orders m and n of the coefficients a i and b j The user recognizes from the formula of the transfer function, which is given on the right-hand side of expression 3, that the transfer function is formed with one filter, and to form two filters, the user changes the order m of the coefficient a. i in the transfer function given on the right-hand side of expression 3, from “2” to “4” and the order n of the coefficient b j from “4” to “6”. In step S36 of Fig. 16. The control unit 205 determines whether the order is changed or the search range is changed, and if the control unit 205 determines that the order is changed as a result of the user changing the order, the control unit 205 displays the transfer function of expression 5 on screen P. Fig. 7, and in step S37 the control unit 205 issues a correction instruction to the servo control device 300, which specifies the correction parameters (the change values ​​of the coefficients a). i and b j ) in the speed control processing unit 308 and includes the orders m and n. The right-hand side of the m expression 5 serves as a mathematical formula model. The change values ​​of the coefficients a i and b j can be determined based on a coefficient, where the evaluation function value stored in memory unit 206 is a local minimum value. The servo control device 310 uses the modified coefficients a in step S11. i and b j , to drive the machine tool, and gives the modified coefficients a i and b j and the position error to the output device 200. b2s2+b3s3+b4s4+b5s5+b6s6a0+a1s+a2s2+a3s3+a4s4=Js2⋅ω12+2Rζω1s+s2ω12+2ζω1s+s2⋅ω22+2Rζω2s+s2ω22+2ζω2s+s2

[0088] In step S38, output device 200 instructs machine learning device 100 to reset the number of trials to "0". Step S38 can be performed concurrently with step S37 or prior to step S37. After step S38, output device 200 returns to step S31. Subsequently, machine learning is performed again based on steps S11, S21, and S31 through S33.

[0089] In this way, the user observes the characteristic curves of the damping center frequency and the weighting function value from the scatter plot in column P3 on screen P. Fig. 7, changes the orders m and n of the coefficients a i and b j as needed to perform machine learning, and can thereby determine the coefficients a i and b j Adjust in the speed control processing unit 308.

[0090] If, on the other hand, the user selects the P on the screen from Fig. When the "Relearn" button (shown in section 7) is selected, the control unit 205 displays the input column for the center frequency fc within screen P. Fig. 7. The user enters, for example, 250 Hz and 400 Hz into the input column. If the user enters 250 Hz and 400 Hz into the input column as the center frequency fc, the control unit 205 determines this in step S36 of Fig. 16. The relearning process begins, and in step S39, the machine learning device 100 is instructed to change or select the search range around 250 Hz and 400 Hz. Subsequently, in step S40, the output device 200 instructs the machine learning device 100 to reset the number of trials to "0". Step S40 can be performed concurrently with step S39 or before step S39. After step S40, the output device 200 returns to step S31. When the machine learning device 100 receives the instruction to change or select the search range and the instruction to start the program, it performs relearning around 250 Hz and 400 Hz in step S21. Here, the search range is changed from a wide range to a narrow range, or it is selected to be a range around 250 Hz and 400 Hz.For example, the search range is changed from 100 Hz to 1000 Hz to 200 Hz to 500 Hz, or selected to be 200 Hz to 300 Hz or 400 Hz to 500 Hz. The output device 200 displays a based on the changed coefficients. i and b j and the evaluation function value supplied by the machine learning device 100, a damping center frequency evaluation function value scatter plot in column P3 on screen P of the in Fig. 7 shown display unit 209 and lists the changed coefficients a i and b j the servo control device 300. In this way, the machine learning is performed again based on steps S11, S21 and S31 to S33.

[0091] In this way, the user observes the characteristic curves of the damping center frequency and the weighting function value from the scatter plot, which is displayed in column P3 on screen P. Fig. 7 is displayed, and changes or selects the machine learning search scope as needed to cause the machine learning device 100 to perform the machine learning, and can thereby determine the coefficients a i and b j in the speed control processing unit 308. Although the output device and the control device of the first embodiment have been described above, the output devices and the control device of a second and a third embodiment are described below. (Second embodiment)

[0092] In the first embodiment, the output device 200 is connected to the servo control device 300 and the machine learning device 100 and performs the transfer of information between the machine learning device 100 and the servo control device 300 and the control of the operations of the servo control device 300 and the machine learning device 100. The present embodiment describes a case in which the output device is connected only to the machine learning device. Fig. Figure 17 is a block diagram illustrating an example of the design of a control device according to the second embodiment of the present invention. The control device 10A includes the machine learning device 100, an output device 200A, the servo control device 300, and the servo motor 400. In comparison with the one in Fig. The output device 200 shown in Figure 6 does not include the information acquisition unit 217 or the information output unit 218.

[0093] Since the output device 200A is not connected to the servo control device 300, the output device 200A does not perform the information exchange between the machine learning device 100 and the servo control device 300, nor does it transmit and receive information with the servo control device 300. Specifically, the output device 200A does not execute the instruction to start the learning program in step S31, the output of physical parameter values ​​in step S33, or the instruction to relearn in step S35, as described in Fig. However, the other processes (for example, steps S32 and S34) shown in 15 carry out the other processes (for example, steps S32 and S34) that are shown in Fig. Figure 15 shows that the output device 200A is not connected to the servo control device 300. This reduces the operation of the output device 200A and thus simplifies its design. (Third embodiment)

[0094] Although in the first embodiment the output device 200 is connected to the servo control device 300 and the machine learning device 100, in the present embodiment a case is described in which an adaptation device is connected to the device 100 and the servo control device 300 and in which the output device is connected to the adaptation device. Fig. Figure 18 is a block diagram illustrating an example of the design of a control device according to the third embodiment of the present invention. The control device 10B includes the machine learning device 100, an output device 200A, the servo control device 300, and the adaptation device 500. Although the in Fig. The output device 200A shown in Figure 18 has the same design as the one in Figure 18. Fig. In the output device 200A shown in Figure 17, the information acquisition unit 211 and the information output unit 212 are not connected to the machine learning device 100, but to the adaptation device 700. The adaptation device 500 is designed such that the drawing unit 203, the operating unit 204, the display unit 209, and the operating unit 2100 are located in the output device 200. Fig. 6 have been omitted.

[0095] Although the in Fig. 18 Output device 200A as shown in Fig. 17. The output device 200A of the second embodiment shown in Figure 17 not only provides the instruction to start the learning program in step S31, the output of physical quantities of parameters in step S33, and the fine-tuning instruction for parameters in step S34, which in Fig. In addition to the actions shown in Figure 15, which also execute the instruction for relearning in step S35, these operations are performed by the adaptation device 700. The adaptation device 500 relays information between the machine learning device 100 and the servo control device 300. The adaptation device 500 relays the instruction to start the learning program and the like to the machine learning device 100, which is carried out by the output device 200A, and outputs the start instructions to the machine learning device 100. In this way, compared to the first embodiment, the function of the output device 200 is divided between the output device 200A and the adaptation device 500, thus reducing the operation of the output device 200A and simplifying its design.

[0096] Although the embodiments according to the present invention have been described above, the servo control device and individual constituent units included in the machine learning device and the output device can be implemented by hardware, software, or a combination thereof. The servo control method, which is carried out by the interaction of the individual constituent units included in the servo control device, can also be implemented by hardware, software, or a combination thereof. Here, "implemented by software" means an implementation where a computer reads and executes a program.

[0097] Programs can be stored on various types of non-transitory, computer-readable storage media and transferred to a computer. Non-transitory, computer-readable storage media include various types of physical storage media. Examples of non-transitory, computer-readable storage media include magnetic recording media (for example, a floppy disk and a hard disk drive) and magneto-optical storage media (for example, a magneto-optical disk), a CD-ROM (Read Only Memory), a CD-R, a CD-R / W, and semiconductor memory (for example, a mask ROM, a PROM (Programmable ROM), an EPROM (Erasable PROM), a Flash ROM, and RAM (Random Access Memory)).

[0098] The embodiment described above is a preferred embodiment and an example of the present invention. However, the scope of the present invention is not limited to this embodiment and example; rather, the present invention can be embodied in various modifications without deviating from the essential content of the present invention. <Varianten, bei denen eine Ausgabevorrichtung in einer Servosteuervorrichtung oder einer Vorrichtung für maschinelles Lernen beinhaltet ist>

[0099] The embodiments discussed above describe the first and second embodiments, in which the machine learning device 100, the output device 200 or 200A, and the servo control device 300 are configured as a control device 10, and the third embodiment, in which the output device 200 is divided into the output device 200A and the adaptation device 500 and provided in the control device. Although in these embodiments the machine learning device 100, the output device 200 or 200A, the servo control device 300, and the adaptation device 500 are configured with separate devices, one of these devices can be integrally configured with another device. For example, part or all of the function of the output device 200 or 200A can be implemented with the machine learning device 100 or the servo control device 300.The output device 200 or 200A can be provided outside the control device, which is equipped with the machine learning device 100 and the servo control device 300. <Spielraum bei der Systemgestaltung>

[0100] Fig. Figure 19 is a block diagram illustrating a control device according to a further embodiment of the present invention. As in Fig. As shown in Figure 19, the control device 10C includes n machine learning devices 100-1 to 100-n, output devices 200-1 to 200-n, n servo control devices 300-1 to 300-n, servo motors 400-1 to 400-n, and a network 600. n is any natural number. Each of the n machine learning devices 100-1 to 100-n corresponds to the one shown in Figure 19. Fig. Device 100 for machine learning is shown in Figure 5. The output devices 200-1 to 200-n correspond to those shown in Figure 5. Fig. 6 shown output device 200 or the one in Fig. Output device 200A shown in Figure 17. Each of the n servo control devices 300-1 to 300-n corresponds to the one shown in Figure 17. Fig. 2 servo control device 300 shown. The output device 200A and the adjustment device 500, which are in Fig. Figure 18 shows output devices 200-1 to 200-n.

[0101] Here, the output device 200-1 and the servo control device 300-1 are configured as a one-to-one pair and are connected in such a way that they are able to exchange data with each other. The output devices 200-2 to 200-n and the servo control devices 300-2 to 300-n are also connected in the same way as the output device 200-1 and the servo control device 300-1. Although in Fig. If 19 n pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n are connected via the network 600, then in the n pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n, the output device and the servo control device in each pair can be directly connected via a connection interface. For example, in the n pairs of output devices 200-1 to 200-n and servo control devices 300-1 to 300-n, a plurality of pairs can be set up in the same manufacturing facility, or they can each be set up in different manufacturing facilities.

[0102] Network 600, for example, refers to a LAN (Local Area Network) set up within a factory, the internet, a public telephone network, or a combination thereof. No specific data transmission method is subject to any particular restrictions within Network 600, regardless of whether a wired or wireless connection is used, etc.

[0103] Although in the control device described above Fig. 19. If the output devices 200-1 to 200-n and the servo control devices 300-1 to 300-n are connected as one-to-one pairs so that they are able to exchange data with each other, a configuration can be chosen, for example, in which one output device 200-1 is connected to a plurality of servo control devices 300-1 to 300-m (m < n or m = n) via the network 600 so that they are able to exchange data with each other, and in which a machine learning device connected to the one output device 200-1 performs the machine learning on the individual servo control devices 300-1 to 300-m. In this case, a distributed processing system can be chosen in which, if necessary, the respective functions of the machine learning device 100-1 are distributed across a plurality of servers.The machine learning functions of device 100-1 can be implemented by using a virtual server or similar function in a cloud. If a plurality of machine learning devices 100-1 to 100-n exist, each corresponding to a plurality of servo control devices 300-1 to 300-n with the same type name, specification, or series, the machine learning devices 100-1 to 100-n can be designed to share the learning results across them. This allows for the creation of another optimal model. Explanation of reference symbols 10, 10A, 10B, 10C Control device 100 Device for machine learning 200 output device 211 Information Acquisition Unit 212 Information output unit 213 character units 214 Control unit 215 Control unit 216 storage units 217 Information acquisition unit 218 Information output unit 219 Display unit 220 operating units 300 servo control device 400 servo motor 500 Adjustment device 600 network< / nichtpatentdokument>

Claims

[1] Output device (200, 200A) comprising: an information acquisition unit (201) that is provided by a machine learning device (100) which performs machine learning on a servo control device (300) for controlling a servo motor (400) that drives an axis of a machine tool, robot or industrial machine, (i) a parameter or a first physical quantity of a component of the servo control device (300) that is or has been machine learned, and (ii) a valuation function value, and an output unit (205 and 209, 205 and 206) that outputs information indicating a relationship between (i) the detected parameter, the first physical quantity or a second physical quantity determined from the parameter, and (ii) the evaluation function value, wherein the output unit (205 and 209) includes a display unit (209) which displays on a screen a graph based on the information specifying the relationship between (i) the parameter, the first physical quantity or the second physical quantity and (ii) the evaluation function value, and wherein the parameter is a coefficient in a transfer function of the component of the servo control device (300) and an instruction to change the order of the coefficient is provided based on the information from the servo control device (300). [2] Output device (200, 200A) according to claim 1, wherein the machine learning device (100) is provided with an instruction to change or select the parameter of the component of the servo control device (300) or a search area of ​​the machine learning of the first physical quantity based on the information specifying the relationship between (i) the detected parameter, the first physical quantity or the second physical quantity determined from the parameter and (ii) the evaluation function value. [3] Output device (200, 200A) according to one of claims 1 to 2, wherein the parameter of the component of the servo control device (300) includes a parameter of a mathematical formula model or a filter. [4] Output device (200, 200A) according to claim 3, wherein the mathematical formula model or the filter is included in a speed feedforward processing unit (308) or a position feedforward processing unit (309) and the parameter includes a coefficient in a transfer function of the filter. [5] Control device (10) comprising: the output device (200, 200A) according to any one of claims 1 to 4; the servo control device (300) that controls the servo motor (400) that drives the axis of the machine tool, robot or industrial machine; and the machine learning device (100) that performs the machine learning on the servo control device (300). [6] Control device (10) according to claim 5, wherein the output device (200, 200A) is included in the servo control device (300) or in the machine learning device (100). [7] Method for outputting an evaluation function value of an output device (200, 200A) used in machine learning of a machine learning device (100) that performs machine learning on a servo control device (300) for controlling a servo motor (400) driving an axis of a machine tool, robot or industrial machine, wherein the method comprises: Capture (i) a parameter or first physical quantity of a component of the servo control device (300) that is or has been machine learned, and (ii) the evaluation function value from the machine learning device (100), and Outputting information that specifies a relationship between (i) the detected parameter, the first physical quantity or a second physical quantity derived from the parameter, and (ii) the evaluation function value, Displays on a screen of a display unit (209) of the output device (200, 200A) of a graph based on the information that specifies the relationship between (i) the parameter, the first physical quantity or the second physical quantity and (ii) the evaluation function value, wherein the parameter is a coefficient in a transfer function of the component of the servo control device (300) and Providing, based on information from the servo control device (300), an instruction to change the order of the coefficient.

Citation Information

Patent Citations

  • adjustment device and adjustment method

    DE102018205015A1

  • Output device, control device and output method for a valuation function value

    DE102019209104A1

  • Output device, control device and method for outputting a learning parameter

    DE102019216081A1

  • PREDICTIVE REPEAT CONTROL METHOD AND APPARATUS FOR SERVO MOTOR

    DE69220561T2

  • Signal converter and signal conversion method using the same

    JP1999031139A