Electric motor control device and electric motor control method
By optimizing the timing and parallel execution of the motor control device, the problem of long adjustment time for motor control commands in the existing technology has been solved, and more efficient automatic adjustment has been achieved.
Patent Information
- Application Number
- CN201980100397.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-19
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2039-09-19
AI Technical Summary
Existing technologies struggle to shorten the time required for automatic adjustment when repeatedly performing initialization, evaluation, and learning operations to adjust motor control commands, and the adjustment work relies on the operator's knowledge and experience.
The system employs an electric motor control unit, which includes a drive control unit, a learning unit, and an adjustment management unit. By optimizing process timing and performing initialization, evaluation, and learning actions in parallel, the automatic adjustment time is shortened.
It reduces the time required for automatic adjustment when repeatedly performing initialization, evaluation, and learning operations, and reduces reliance on operator experience.
Smart Images

Figure CN114514481B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a motor control device that automatically adjusts control commands for controlling an electric motor. Background Technology
[0002] In electronic component mounting machines, semiconductor manufacturing equipment, and the like, positioning control is used to move the mounting head and other machines a target distance by driving an electric motor. In positioning control, parameters that define the trajectory of the position contained in the command signal used to drive the electric motor, as well as parameters of the control system, are adjusted and set in order to shorten positioning time and improve equipment productivity.
[0003] Adjusting these parameters sometimes requires trial and error, which takes time and labor. Furthermore, there are issues such as the time required for adjustment and the dependence of the results on the operator's knowledge and experience. As a technique to address these problems, a method for automating parameter adjustment operations has been proposed.
[0004] The control parameter adjustment device described in Patent Document 1 includes a model updating unit that updates the control object model using data from when the controlled object operates. Furthermore, it includes a first search unit that searches for control parameters within a first range and repeatedly performs simulations using the updated control object model to extract candidate optimal values. Additionally, it includes a second search unit that repeatedly operates the controlled object within a second range, narrower than the first range, to obtain an operation result.
[0005] The machine learning device described in Patent Document 2 includes a state observation unit that observes the state variables of a motor that is driven and controlled by a motor control device. Furthermore, it includes a learning unit that learns conditions associated with correction amounts used to correct commands from the motor control device, based on a training data set consisting of state variables.
[0006] Patent Document 1: Japanese Patent Application Publication No. 2017-102619
[0007] Patent Document 2: Japanese Patent Application Publication No. 2017-102613 Summary of the Invention
[0008] The devices described in Patent Documents 1 and 2 both automate the parameter adjustment process by repeatedly and alternately performing evaluation operations to obtain sensor values when the motor is driven, and then using the sensor values obtained through the evaluation operations for calculation processing. Here, the calculation processing includes simulation, learning, etc. However, when adjustments are performed by repeatedly performing evaluation operations and calculation processing based on the motor drive, an initialization operation is sometimes required. This initialization operation refers to setting the motor, etc., to the state before the evaluation operation began, i.e., the initial state. Furthermore, in such cases, it is difficult to shorten the time required for automatic adjustment when automatically adjusting the control commands for controlling the motor by repeatedly performing initialization operations, evaluation operations, and learning actions.
[0009] The present invention was made in view of the above circumstances, and its object is to provide a motor control device that can shorten the time required for automatic adjustment when performing automatic adjustment of control commands for controlling the motor by repeatedly performing initialization operation, evaluation operation and learning actions.
[0010] The electric motor control device of the present invention comprises: a drive control unit that drives an electric motor based on control commands to cause a controlled object consisting of the electric motor and a mechanical load mechanically connected to the electric motor to operate, and executes an initialization operation that sets the controlled object to an initial state and an evaluation operation that starts from the initial state; a learning unit that learns by associating control commands for the evaluation operation with state sensor signals obtained by detecting the state of the controlled object during the evaluation operation, and determines control commands for the evaluation operation to be executed after the evaluation operation in which the state sensor signals are obtained based on the learning results; and an adjustment management unit that determines the timing of executing a second step, which is any one of the learning operation, initialization operation, and evaluation operation, based on the timing of executing a first step, which is the operation of the learning unit, i.e., the learning operation, the initialization operation, and the evaluation operation.
[0011] The effects of the invention
[0012] According to the present invention, a motor control device can be provided such that, when automatically adjusting the control commands for controlling the motor by repeatedly performing initialization operation, evaluation operation, and learning actions, the time required for automatic adjustment can be shortened. Attached Figure Description
[0013] Figure 1 This is a block diagram illustrating an example of the structure of the motor control device in Embodiment 1.
[0014] Figure 2 This is a diagram illustrating an example of the timing of the operation of the motor control device in Embodiment 1.
[0015] Figure 3 This is a flowchart illustrating an example of the actions of the adjustment management department in Implementation Method 1.
[0016] Figure 4 This is a diagram illustrating an example of the instruction pattern in Implementation Method 1.
[0017] Figure 5 This is a block diagram illustrating an example of the structure of the learning unit in Implementation Method 1.
[0018] Figure 6 This is a diagram illustrating an example of the time response to the deviation in Implementation 1.
[0019] Figure 7 This is a diagram illustrating a structural example of the processing circuit of the motor control device in Embodiment 1, which consists of a processor and a memory.
[0020] Figure 8 This is a diagram illustrating a structural example of the processing circuit of the motor control device in Embodiment 1, which is constructed using dedicated hardware.
[0021] Figure 9 This is a block diagram illustrating an example of the structure of the motor control device in Embodiment 2.
[0022] Figure 10 This is a diagram illustrating an example of the timing of the operation of the motor control device in Embodiment 2.
[0023] Figure 11 This is a flowchart illustrating an example of the operation of the adjustment management department in Implementation Method 2.
[0024] Figure 12 This is a block diagram illustrating an example of the structure of the motor control device in Embodiment 3.
[0025] Figure 13 This is a diagram illustrating an example of the timing of the operation of the motor control device in Embodiment 3.
[0026] Figure 14 This is a block diagram illustrating an example of the structure of the motor control device in Embodiment 4.
[0027] Figure 15 This is a diagram illustrating an example of the timing of the operation of the motor control device in Embodiment 4.
[0028] Figure 16 This is a flowchart illustrating an example of the actions of the adjustment management department in Implementation Method 4. Detailed Implementation
[0029] The embodiments will now be described in detail with reference to the accompanying drawings. Furthermore, the embodiments described below are merely illustrative. Additionally, the various embodiments can be appropriately combined and implemented.
[0030] Implementation Method 1
[0031] Figure 1 This is a block diagram illustrating an example of the structure of the motor control device 1000 in Embodiment 1. The motor control device 1000 includes: a drive control unit 4 that drives the motor 1 in accordance with a command signal 103; and a command generation unit 2 that acquires command parameters 104 and generates the command signal 103. Furthermore, the motor control device 1000 includes a learning unit 7 that acquires a learning start signal 106 and a status sensor signal 101, and determines a learning completion signal 107 and the command parameters 104. Moreover, the motor control device 1000 includes an adjustment management unit 9 that acquires the learning completion signal 107 and determines a learning start signal 106 and a command start signal 105.
[0032] The electric motor 1 generates torque, thrust, etc., by means of drive power E output from the drive control unit 4. Examples of electric motor 1 include rotary servo motors, linear motors, and stepper motors. The mechanical load 3 is mechanically connected to the electric motor 1 and is driven by the electric motor 1. The electric motor 1 and the mechanical load 3 are referred to as the controlled object 2000. The mechanical load 3 can be a device that operates by means of torque, thrust, etc., generated by the electric motor 1. The mechanical load 3 can also be a device that performs positioning control. Examples of mechanical load 3 include electronic component assembly machines and semiconductor manufacturing equipment.
[0033] The drive control unit 4 supplies drive power E to the motor 1 based on the command signal 103, thereby driving the motor 1 to follow the command signal 103 and causing the controlled object 2000 to operate, performing evaluation operation and initialization operation. Here, the command signal 103 can also be at least one of the position, speed, acceleration, current, torque, or thrust of the motor 1. Initialization operation is operation that sets the controlled object 2000 to an initial state. Evaluation operation is operation that starts from the initial state, and the state sensor signal 101 obtained during the evaluation operation is used for the learning operation described later. As the drive control unit 4, a structure that makes the position of the motor 1 follow the command signal 103 can be appropriately adopted. For example, it can also be configured as a feedback control system, which calculates the torque or current of the motor 1 based on PID control in a way that minimizes the difference between the detected position of the motor 1 and the command signal 103. Alternatively, the drive control unit 4 may be a 2-DOF control system, which adds feedforward control to the feedback control that drives the motor 1 in such a way that the position of the detected mechanical load 3 follows the command signal 103.
[0034] The instruction generation unit 2 generates an instruction signal 103 based on instruction parameters 104. Furthermore, the instruction generation unit 2 generates the instruction signal 103 in accordance with the timing indicated by the instruction start signal 105. Moreover, the motor 1 starts operating at the timing indicated by the instruction signal 103 generated by the instruction generation unit 2. As described above, the motor 1 starts operating in accordance with the timing indicated by the instruction start signal 105. That is, the motor 1 starts operating according to the instruction start signal 105. Here, the evaluation operation or initialization operation is referred to as operation. The initialization operation and evaluation operation are performed by following the instruction signal 103 of each operation, and the instruction signal 103 of the initialization operation and evaluation operation is generated based on the instruction parameters 104 used by each operation. An example of the operation of the instruction generation unit 2 will be used later. Figure 4 To narrate.
[0035] The state sensor 5 detects the state quantity of at least one of the motor 1 or the mechanical load 3, i.e., the state quantity of the controlled object 2000, and outputs the result as a state sensor signal 101. Examples of state quantities include the position, speed, acceleration, current, torque, and thrust of the motor 1. Similarly, examples of state quantities include the position, speed, and acceleration of the mechanical load 3. Examples of the state sensor 5 include encoders, laser displacement gauges, gyroscopes, accelerometers, current sensors, and force sensors. Figure 1 The state sensor 5 is described as an encoder that detects the position of the motor 1 as a state quantity.
[0036] The learning unit 7 learns by associating the instruction parameters 104 used in the evaluation operation with the state sensor signals 101 obtained by detecting the state of the controlled object 2000 during the evaluation operation. Furthermore, it determines the instruction parameters 104 used in the evaluation operation to be executed after the state sensor signals 101 are acquired. The actions of the learning unit 7 from the start of this learning until the instruction parameters 104 are determined are called learning actions. The learning unit 7 begins learning according to the learning start signal 106. Here, the learning start signal 106 is a signal indicating the start time of the learning action, determined by the adjustment management unit 9 described later. The learning unit 7 further determines the learning completion signal 107. The learning completion signal 107 indicates the time point at which the instruction parameters 104 are determined, i.e., the completion time point of the learning action. Detailed operations of the learning unit 7 will be described using... Figure 5 and Figure 6 The details will be provided later.
[0037] The adjustment management department 9 determines the value of the instruction start signal 105, which indicates the start time of the evaluation operation, based on the learning completion signal 107, thereby determining the start time of the evaluation operation based on the completion time of the learning action. Additionally, in Figure 2 In the example operation, the adjustment management unit 9 determines the learning start signal 106, indicating the start time of the learning action, and the instruction start signal 105, indicating the start time of the initialization operation, based on the completion time of the evaluation operation. Furthermore, as described later, the adjustment management unit 9 can detect the elapsed time from the start time of the evaluation operation and detect the completion time of the evaluation operation. In other words, the adjustment management unit 9 determines the start time of the learning action and the initialization operation based on the completion time of the evaluation operation.
[0038] Figure 2 This is a diagram illustrating an example of the timing of the operation of the motor control device 1000 in Embodiment 1. Figure 2 (a) to Figure 2 (e) The horizontal axis represents time. Figure 2 (a) to Figure 2 (e) represents the learning action, action processing (initialization operation and evaluation operation), learning start signal 106, learning completion signal 107 and instruction start signal 105.
[0039] The relationship between the values of the instruction start signal 105, the learning start signal 106, and the learning completion signal 107 and the content indicated by each signal is explained. Figure 2In this process, when the value of the instruction start signal 105 becomes 1, the motor 1 starts running. Additionally, when the value of the learning start signal 106 becomes 1, the learning unit 7 begins learning the action. Furthermore, the learning unit 7 sets the value of the learning completion signal 107, representing the time when the learning action is completed, to 1. Moreover, the values of the instruction start signal 105, learning start signal 106, and learning completion signal 107 can also be reset to 0 after becoming 1 until the next action is indicated. These signals only need to represent the start and completion times of the action, and are not limited to the above methods.
[0040] The evaluation operation, initialization operation, and learning action are referred to as processes. A cycle in which each process, including at least one initialization operation, evaluation operation, and learning action, is executed periodically and repeatedly is called a learning cycle. Figure 2 Each learning cycle includes one initialization run, one evaluation run, and one learning action. The instruction parameter 104 can also be updated on a per-learning-cycle basis. The motor control device 1000 advances its learning by repeatedly executing the learning cycle. The adjustment action that repeatedly executes the learning cycle to search for the instruction parameter 104 that provides the optimal action for the controlled object 2000 will be referred to as automatic adjustment.
[0041] Figure 3 This is a flowchart illustrating an example of the operation of the adjustment management unit 9 in Implementation Method 1. (See attached diagram.) Figure 2 and Figure 3 The operation of the motor control device 1000 is illustrated below. If automatic adjustment is initiated, in step S101, the adjustment management unit 9 sets the value of the learning start signal 106 at time TL111 to 1, thus determining the start time of the learning action L11. The learning unit 7 starts learning action L11 at time TL111 according to the learning start signal 106. Furthermore, as with learning action L11, if the learning unit 7 starts learning action after automatic adjustment has begun but before obtaining the status sensor signal 101 for evaluation operation, the learning unit 7 can also randomly determine the command parameter 104. Alternatively, it can be determined based on a preset setting. In the case of random determination, the action value function Q (described later) can be initialized with a random number, and action a can be randomly determined. t That is, instruction parameter 104.
[0042] In step S102, the adjustment management unit 9 sets the value of the instruction start signal 105 at time TL111 to 1, thus determining the start time of the initialization operation IN11. The motor 1 begins the initialization operation IN11 at time TL111 according to the instruction start signal 105. The initialization operation IN11 is executed in parallel with the learning operation L11. Hereinafter, parallel execution means that at least a portion of the two processes are executed repeatedly in time. Furthermore, the time required for the initialization operation IN11 is shorter than the time required for the learning operation L11. Therefore, the adjustment management unit 9 can also set the start time of the initialization operation IN11 to be later than the start time of the learning operation L11, within a range where the waiting time is not extended (i.e., the completion of the initialization operation IN11 is not later than the completion of the learning operation L11). The motor 1 completes the initialization operation IN11 at time TL112 and enters a standby state after the initialization operation IN11 is completed. Furthermore, the motor 1 in the standby state can be controlled to be within a predetermined position range, or it can be stopped. Additionally, the power supply can be stopped. Next, the learning unit 7 sets the value of the learning completion signal 107 at time TL113, the time point at which the learning action is completed, to 1.
[0043] In step S103, the adjustment management unit 9 detects the time point when the value of the learning completion signal 107 becomes 1, and detects time TL113 as the completion time point of the learning action L11. Furthermore, regarding the action in step S103, the adjustment management unit 9 only needs to detect the completion time point of the learning action; for example, it can detect the time point when the learning unit 7 outputs the instruction parameter 104. In step S104, based on the completion time point of the learning action, i.e., time TL113, the adjustment management unit 9 determines the value of the instruction start signal 105 at time TL113 to be 1, and determines the start time point of the evaluation operation EV11 (first evaluation operation). The motor 1 starts the evaluation operation EV11 at time TL113 according to the instruction start signal 105. If the evaluation operation EV11 is completed at time TL114, the motor 1 enters a standby state.
[0044] In step S105, the adjustment management unit 9 detects the elapsed time since the start of the evaluation operation EV11, and detects time TL121 as the completion time of the evaluation operation EV11. Here, the aforementioned predetermined time is set to be the same as or longer than the estimated time required for the evaluation operation EV11. Furthermore, in this embodiment, it should be noted that the time detected by the adjustment management unit 9 as the completion time of the evaluation operation EV11 is different from the time when the evaluation operation EV11 ends and the motor 1 stops. In step S106, the adjustment management unit 9 performs a determination on whether to continue automatic adjustment. If it is determined that automatic adjustment should continue, the process proceeds to step S107; if it is determined that automatic adjustment should not continue, the process proceeds to step S108.
[0045] Regarding the method for determining step S106, for example, it could be that if the number of learning cycles performed in automatic adjustment is less than a predetermined number, automatic adjustment continues; if the number is the same as the predetermined number, it is determined that automatic adjustment will not continue. Alternatively, it could be determined that automatic adjustment will not continue if the state sensor signal 101 obtained from the evaluation operation immediately preceding step S106 meets a predetermined benchmark; otherwise, it is determined that automatic adjustment will continue. The benchmark for this state sensor signal 101 could be, for example, a value used later. Figure 6 The convergence time of the described positioning action is less than or equal to a predetermined time, as a benchmark.
[0046] In step S106 executed at time TL121, the adjustment management unit 9 determines that automatic adjustment should continue and proceeds to step S107. In step S107, the adjustment management unit 9, based on the completion time of the evaluation operation EV11, i.e., time TL121, sets the values of the learning start signal 106 and the command start signal 105 at time TL121 to 1. Through this action, the start times of the learning operation L12 (first learning operation) and the initialization operation IN12 (first initialization operation) are determined respectively. The learning unit 7 and the motor 1 each begin the learning operation L12 and the initialization operation IN12 at time TL121 according to the learning start signal 106 and the command start signal 105. The period from time TL111 to time TL121 is designated as the learning cycle CYC11.
[0047] Then, steps S103 to S107 are repeated until step S106 determines that automatic adjustment should not continue. Furthermore, in step S103 of the learning loop CYC12, the adjustment management unit 9 detects time TL123 as the completion time of learning action L12. Moreover, in step S104 of the learning loop CYC12, based on the detected completion time of learning action L12, the adjustment management unit 9 determines the start time of evaluation operation EV12 (second evaluation operation) as time TL123.
[0048] The adjustment management unit 9 executes step S106 of the learning loop CYC1X at time TL1X1. Then, it determines that automatic adjustment should not continue and proceeds to step S108. In step S108, the adjustment management unit 9 determines the value of the learning start signal at time TL1X1 to be greater than 1 and instructs the learning unit 7 to end the process T1. The instruction for end process T1 only needs to be a time that can instruct the learning unit 7 to start the end process. For example, the value of the learning start signal 106 indicating the end process time can be determined to be a value other than 0 and 1, or other signals can be output to the learning unit 7 at the time indicating the end process time. The learning unit 7 detects the start time of end process T1 and executes end process T1.
[0049] In the end processing T1, the learning unit 7 can also determine the optimal instruction parameter 104, i.e., the optimal instruction parameter 104, based on the learning actions repeatedly executed in automatic adjustment. As an example of evaluation operation, the end processing T1 is shown when performing positioning that moves the controlled object 2000 to a target distance. First, evaluation operations using the instruction parameters 104 from all learning cycles are selected where the difference between the position of the motor 1 and the target moving distance (i.e., the deviation) falls within a predetermined allowable range and does not exceed it. Furthermore, the instruction parameters 104 used in these evaluation operations are set as candidates for the optimal instruction parameter 104. Moreover, among the candidates for instruction parameters 104, it is also possible to further select the instruction parameter 104 that executes an evaluation operation that keeps the deviation within the allowable range for the shortest period from the start of the evaluation operation, and determine it as the optimal instruction parameter 104. The aforementioned deviation will be used later. Figure 4 To narrate.
[0050] Furthermore, the learning unit 7 can also determine the optimal instruction parameter 104 from the instruction parameters 104 that have not yet been used in the evaluation operation. For example, from the instruction parameters 104 used in the evaluation operation of all learning cycles, the instruction parameter 104 that performs the action that makes the deviation fall within the allowable range within a predetermined time is selected. Moreover, the average value of the selected instruction parameters 104 can also be determined as the optimal instruction parameter 104. Figure 2 At time TL1Y1, if the learning unit 7 has completed the termination process T1, the termination is automatically adjusted. Alternatively, the termination process T1 can be omitted. For example, the instruction parameter 104 used to evaluate the operation of EV1X can be determined as the optimal instruction parameter 104.
[0051] The first and second processes can be set as any one of evaluation operation, initialization operation, or learning operation. The adjustment management unit 9 can also determine the timing of the second process based on the timing of the first process. In addition, the timing of the first and second processes can be set to the start or end time of each process, or to a time point offset from the start or end time by a predetermined time. By determining the timing of the second process based on the timing of the first process, the interval between the two processes can be shortened, thereby reducing the waiting time until the motor 1 or the learning unit 7 begins its process.
[0052] right Figure 2 Describe the relationship between the various processes in the action example. Figure 2 In the example operation, the next evaluation operation is executed using the instruction parameter 104 determined in the learning operation, and the next learning operation is executed using the state sensor signal 101 obtained through the evaluation operation. Therefore, the learning operation and the evaluation operation are not executed in parallel. Furthermore, since the evaluation operation and the initialization operation are executed through a single control object 2000, the evaluation operation and the initialization operation are not executed in parallel. On the other hand, since the initialization operation and the learning operation do not interfere with each other, they can be executed in parallel. Moreover, in Figure 2 In the example shown, the time required to learn the action is longer than the time required to initialize the operation.
[0053] exist Figure 2 In the example operation, the adjustment management unit 9 determines the learning start signal 106, indicating the start time of the learning operation, and the instruction start signal 105, indicating the start time of the initialization operation, based on the completion time of the evaluation operation. Furthermore, the learning operation L12 and the initialization operation IN12 begin at the completion time of the evaluation operation EV11 detected by the adjustment management unit 9, and the evaluation operation EV12 begins at the completion time of the learning operation L12. This embodiment is not limited to this operation.
[0054] For example, evaluation operation EV11 (first evaluation operation), which is one of the evaluation operations, can be performed. Learning action L12 is executed using the state sensor signal 101 acquired during evaluation operation EV11, and initialization operation IN12 is performed in parallel with learning action L12. Furthermore, based on the instruction parameters 104 (control instructions) determined by learning action L12, the next evaluation operation after evaluation operation EV11, namely evaluation operation EV12 (second evaluation operation), can be performed from the initial state set by initialization operation IN12. By executing each process as described above, initialization operation IN12 and learning action L11 can be performed in parallel, adjusting the timing between processes and shortening waiting time. The motor control device 1000 or motor control method can also be provided in this manner.
[0055] Additionally, for example, the adjustment management unit 9 can detect the completion time of evaluation operation EV11, and based on the detected completion time of evaluation operation EV11, determine the start time of learning action L12 and the start time of initialization operation IN12, adjusting the timing between processes to shorten waiting time. Furthermore, for example, the adjustment management unit 9 can also determine the start time of the one requiring more time in learning action L12 and initialization operation IN12 to be simultaneous with or earlier than the other's start time, shortening waiting time. Additionally, the adjustment management unit 9 can also detect the completion time of the one completing simultaneously or later in learning action L12 or initialization operation IN12, and based on the detected completion time, determine the start time of evaluation operation EV12, shortening waiting time. In the above-described action examples, when determining the start time of the next process based on the completion time of the previous process, it is preferable to set the interval between the completion time of the previous process and the start time of the next process to be as short as feasible, and more preferably to set them to be simultaneous or approximately simultaneous.
[0056] Furthermore, the adjustment management unit 9 detects the completion time of the learning action L11 by detecting the time elapsed since the start time of the learning action L11, but this embodiment is not limited to this method. For example, sometimes two processes, namely the first process and the second process, are executed. During the period from the completion of the first process to the start of the second process, an intermediate process including at least one of initialization operation, evaluation operation, and learning action is executed. In such cases, the adjustment management unit 9 can also estimate the time required for the intermediate process in advance and determine the start time of the second process as a time point later than the time point after the estimated time required for the intermediate process from the completion time of the first process. By doing so, the start time of the second process can be adjusted with the estimated time required for the intermediate process as the target, shortening the waiting time and thereby reducing the time required for automatic adjustment. Alternatively, it can also be used as follows: Figure 2 As illustrated in the example, by learning the completion signal 107, the adjustment management department 9 can more accurately detect the completion time of the learned action and accurately determine the start time of the next process. Furthermore, it can also shorten waiting time.
[0057] Next, the operation of the instruction generation unit 2 in generating instruction signal 103 based on instruction parameter 104 will be illustrated. Figure 4 This diagram illustrates an example of the command pattern in Implementation Method 1. Here, the command pattern represents the command values of the motor 1 in a timing sequence. The command value for this pattern is any one of the position, speed, acceleration, or jerk of the motor 1. The command value can also be the same as the value of the command signal 103. Furthermore, in... Figure 4 In the action example, the instruction pattern is the one that represents instruction signal 103 in sequence.
[0058] During evaluation operation, command parameter 104, together with operating conditions, defines the command pattern. In other words, if command parameter 104 and operating conditions are specified, the command pattern is uniquely determined. Here, operating conditions are constraints on the operation of motor 1 during evaluation operation, and are constant during repeated evaluation operations in automatic adjustment. On the other hand, command parameter 104 can be updated in automatic adjustment on a per-learning-cycle basis. Figure 1In the motor control device 1000, the instruction generation unit 2 generates an instruction signal 103 based on the instruction parameter 104. At this time, as a result, the drive control unit 4 drives the motor 1 based on the instruction parameter 104. Moreover, the drive control unit 4 can also drive the motor 1 based on the instruction pattern. As described above, if the instruction signal 103, the instruction parameter 104, or the instruction pattern is set as an instruction for controlling the motor 1, that is, a control instruction, the drive control unit 4 drives the motor 1 based on the control instruction.
[0059] Figure 4 (a) to Figure 4 (d), the horizontal axis is time. In Figure 4 (a) to Figure 4 (d), the vertical axis of each shows the position, speed, acceleration, and jerk of the motor 1, which are the instruction signal 103. Here, the speed, acceleration, and jerk are the first derivative, second derivative, and third derivative of the position of the motor 1, respectively. The intersection point of the horizontal axis and the vertical axis is the time 0 on the horizontal axis, which is the instruction start time point for starting the evaluation of the operation. Regarding Figure 4 the operating conditions of the operation example, the target moving distance is set to D. That is, the position of the motor 1 is 0 at the start time point 0 of the evaluation operation, and the position of the motor 1 is D at the time t = T1 + T2 + T3 + T4 + T5 + T6 + T7, which is the end time point.
[0060] Figure 4 The instruction pattern is sequentially divided into the first interval to the seventh interval from the instruction start time point, that is, time 0, to the end time point. Let n be a natural number from 1 to 7, and the time length of the nth interval is set as the nth time length Tn. In Figure 4 the operation example, the seven parameters from the first time length T1 to the seventh time length T7 are set as the instruction parameter 104. The magnitudes of the acceleration in the second interval and the sixth interval are set as Aa and Ad, respectively, and they are constant within the interval. It should be noted that the magnitude of the acceleration Aa and the magnitude of the acceleration Ad are dependent variables of the instruction parameter 104 and have no setting freedom.
[0061] The instruction signal 103 at the time t (0 ≤ t < T1) in the first interval can be calculated as follows. The acceleration A1, speed V1, and position P1 obtained by integrating the jerk, acceleration A1, and speed V1 during the period from time 0 in the first interval to the time t within the first interval with respect to time. Moreover, in the first interval, the acceleration increases at a certain rate and reaches the magnitude of the acceleration Aa at time T1. Therefore, the jerk in the first interval is the value obtained by dividing the magnitude of the acceleration Aa by T1. As described above, the acceleration A1, speed V1, and position P1 can be calculated as in equations (1) to (3), respectively.
[0062] [Mathematical formula 1]
[0063]
[0064] [Mathematical formula 2]
[0065] V1(t)=∫0 t A1(τ)dτ···(2)
[0066] [Mathematical formula 3]
[0067] P1(t)=∫0 t V1(τ)dτ···(3) <000021३>In addition, similar to the first interval, the command signal 103 at time t (T1≤t<T1+T2) in the second interval, i.e., the acceleration A2, velocity V2, and position Pz, can be calculated as in equations (4) to (6).
[0069] [Mathematical formula 4]
[0070] A2(t)=Aa···(4)
[0071] [Mathematical formula 5]
[0072]
[0073] [Mathematical formula 6]
[0074]
[0075] In addition, similar to the first interval, the command signal 103 at time t (T1+T2≤t<T1+T2+T3) in the third interval, i.e., the acceleration A3, velocity V3, and position P3, can be calculated as in equations (7) to (9).
[0076] [Mathematical formula 7]
[0077] <00००234><00002५5>
[0078] [Mathematical formula 8]
[0079] [[ID=५6]]
[0080] [Mathematical formula 9]
[0081] <00१०245>
[0082] In addition, similar to the first interval, the command signal 103 at time t (T1+T2+T3≤t<T1+T2+T3+T4) in the fourth interval, i.e., the acceleration A4, velocity V4, and position P4, can be calculated as in equations (10) to (12). It should be noted that there may be some inaccuracies in the original text, such as "Pz" in item 23 which may be a miswriting. It is recommended to check and correct the original text for more accurate translation.
[0083] [Mathematical formula 10]
[0084] A4(t)=0···(10)
[0085] [Mathematical formula 11]
[0086]
[0087] [Mathematical formula 12]
[0088]
[0089] In addition, similar to the first interval, the command signal 103 at time t (T1 + T2 + T3 + T4 ≤ t < T1 + T2 + T3 + T4 + T5) in the fifth interval, that is, the acceleration A5, the velocity V5, and the position P5 can be calculated as in equations (13) to (15).
[0090] [Mathematical formula 13]
[0091]
[0092] [Mathematical formula 14]
[0093]
[0094] [Mathematical formula 15]
[0095]
[0096] In addition, similar to the first interval, the command signal 103 at time t (T1 + T2 + T3 + T4 + T5 ≤ t < T1 + T2 + T3 + T4 + T5 + T6) in the sixth interval, that is, the acceleration A6, the velocity V6, and the position P6 can be calculated as in equations (16) to (18).
[0097] [Mathematical formula 16]
[0098] A6(t)=-Ad···(16)
[0099] [Mathematical formula 17]
[0100]
[0101] [Mathematical formula 18]
[0102]
[0103] In addition, similar to the first interval, the command signal 103 at time t (T1+T2+T3+T4+T5+T6≤t≤T1+T2+T3+T4+T5+T6+T7) in the seventh interval can be calculated as in equations (19) to (21), namely, acceleration A7, velocity V7 and position P7.
[0104] [Mathematical Expression 19]
[0105]
[0106] [Mathematical Expression 20]
[0107]
[0108] [Mathematical Expression 21]
[0109]
[0110] Furthermore, at the final time point t = T1 + T2 + T3 + T4 + T5 + T6 + T7, the velocity V7 is equal to 0, and the position P7 is equal to the target's distance D. Therefore, at the final time point, equations (22) and (23) hold. The magnitude of the acceleration Aa in the second interval and the magnitude of the acceleration Ad in the sixth interval can be determined by equations (22) and (23).
[0111] [Mathematical Expression 22]
[0112] V7=0···(22)
[0113] [Mathematical Expression 23]
[0114] P7=D···(23)
[0115] The above is an example of the operation of the instruction generation unit 2, which generates the instruction signal 103 based on the instruction parameter 104 and the operating conditions. Here, in the first interval, the third interval, the fifth interval, and the seventh interval, the jerk is a non-zero constant value. That is, the first time length T1, the third time length T3, the fifth time length T5, and the seventh time length T7 specify the time when the jerk becomes a non-zero constant value. Here, a non-zero constant value means a constant value greater than 0 or a constant value less than 0. In addition, in these intervals, the magnitude of the jerk can also be set as the instruction parameter 104 instead of the time length Tn. For example, when the magnitude of the jerk in the first interval is defined as J1, the first time length T1 and the jerk J1 have the relationship shown in equation (24).
[0116] [Mathematical Expression 24]
[0117]
[0118] Defining the duration of the interval during which the accelerometer becomes a non-zero constant value as command parameter 104 is equivalent to defining the magnitude of the accelerometer during the interval during which the accelerometer becomes a non-zero constant value as command parameter 104. As in the example above, the command pattern can be determined simply by combining command parameter 104 with the operating conditions. As in the example given here, there may be multiple options for selecting command parameter 104 even under the same operating conditions. Moreover, the method for selecting command parameter 104 is not limited to the method described in this embodiment.
[0119] Explanation of Section 7 of the Study Department. Figure 5 This is a block diagram illustrating an example of the structure of the learning unit 7 in Embodiment 1. The learning unit 7 includes a reward calculation unit 71, a value function update unit 72, an intent determination unit 73, a learning completion signal determination unit 74, a command parameter determination unit 75, and an evaluation sensor signal determination unit 76. The reward calculation unit 71 calculates the reward r for the command parameter 104 used in the evaluation operation based on the evaluation sensor signal 102. The value function update unit 72 updates the action value function corresponding to the reward r. The intent determination unit 73 uses the action value function updated by the value function update unit 72 to determine a candidate evaluation parameter 108 that will become the command parameter 104 used in the evaluation operation. The command parameter determination unit 75 determines the command parameter 104 used in the evaluation operation based on the candidate evaluation parameter 108. The evaluation sensor signal determination unit 76 determines the evaluation sensor signal 102 based on the state sensor signal 101 during the evaluation operation. Furthermore, the intent determination unit 73 may replace the candidate evaluation parameter 108 in determining the command parameter 104. Moreover, the command parameter determination unit 75 may be omitted from the learning unit 7.
[0120] Furthermore, the learning unit 7 can also learn the command signal 103 or the command pattern instead of the command parameter 104, thus enabling the learning unit 7 to learn control commands. In this case, the learning unit 7 has a control command determination unit instead of the command parameter determination unit 75. The control command determination unit determines the control command used for evaluating operation based on the evaluation candidate parameter 108. Moreover, while the command pattern and command signal 103 each individually specify the operation of the motor 1, the command parameter 104 specifies the operation of the motor 1 through a combination of the command parameter 104 and the operating conditions. Therefore, compared to the case where the learning unit 7 learns the command pattern or command signal 103, the amount of data required for the learning unit 7 to learn the command parameter 104 is reduced, thus decreasing the computational load and computation time of the learning unit 7. Therefore, the learning operation can be performed efficiently when learning the command parameter 104.
[0121] The evaluation sensor signal determination unit 76 can also derive the evaluation sensor signal 102 by performing calculations such as extraction, transformation, correction, and filtering on the state sensor signal 101. For example, the evaluation sensor signal 102 can be obtained by extracting the state sensor signal 101 from the state sensor signal 101 over time during the evaluation operation. Here, the state sensor signal 101 from the start to the end of the evaluation operation can be extracted, and the state sensor signal 101 from the end of the evaluation operation to the elapsed time after a predetermined period can also be extracted to evaluate the impact of oscillations immediately after the completion of the evaluation operation. In addition, it can be configured to apply correction to the obtained state sensor signal 101 and remove the offset value when determining the evaluation sensor signal 102. Alternatively, it can be configured to pass the state sensor signal 101 through a low-pass filter to remove noise. These signal processing methods can also improve the accuracy of the learned action. In addition, the reward calculation unit 71 can be configured to calculate the reward r based on the state sensor signal 101, thus omitting the evaluation sensor signal determination unit 76.
[0122] The learning unit 7 can perform learning using various learning algorithms. In this embodiment, the application of reinforcement learning will be explained as an example. Reinforcement learning involves an agent (acting entity) observing the current state within a certain environment and deciding on the action to be taken. The agent selects actions and receives rewards from the environment. Furthermore, it learns the strategy that yields the highest reward through a series of actions. Representative methods of reinforcement learning include Q-learning and TD-learning. For example, in the case of Q-learning, the usual update formula for the action value function Q(s, a) is expressed by equation (25). The update formula can also be expressed using an action value table.
[0123] [Mathematical Expression 25]
[0124]
[0125] In equation (25), s t a represents the environment at time t. t This represents the action at time t. Through action a... t Environmental changes are s t+1 r t+1 Let γ represent the reward obtained through changes in the environment, α represent the discount rate, and α represent the learning coefficient. Furthermore, the discount rate γ is set to a range greater than 0 and less than or equal to 1 (0 < γ ≤ 1), and the learning coefficient α is set to a range greater than 0 and less than or equal to 1 (0 < α ≤ 1). When Q-learning is applied, action a...t The decision is made for instruction parameter 104, but in reality, the action that determines the evaluation of candidate parameter 108 is sometimes also action a. t Environments t It consists of operating conditions, the initial position of motor 1, etc.
[0126] use Figure 6 An example of the operation of the remuneration calculation unit 71 is given. Figure 6 This is a diagram illustrating an example of the time response to the deviation in Implementation 1. Figure 6 The deviation is the difference between the target movement distance and the position of motor 1 when motor 1 is activated during operation. Figure 6 The horizontal axis represents time, and the vertical axis represents deviation. The intersection of the vertical and horizontal axes represents a state of zero deviation on the vertical axis and a point on the horizontal axis that marks the start of the evaluation operation (0). Figure 6 In this context, IMP is the limit of the allowable range of deviation, representing the magnitude of the error in the allowable motion accuracy under mechanical load 3.
[0127] Figure 6 The deviation of (a) falls within the allowable range from the start of the evaluation operation to time Tst1, and then oscillates and converges within the allowable range. Figure 6 The deviation in (b) falls within the allowable range from the start of the evaluation operation up to time Tst2, then temporarily exceeds the allowable range. Then, it falls back within the allowable range. Figure 6 The deviation in (c) falls within the allowable range from the start of the evaluation operation to time Tst3, and then oscillates and converges within the allowable range. Here, between time Tst1, time Tst2, and time Tst3, there exists a relationship where the value of time Tst2 is smaller than the value of time Tst3, and the value of time Tst3 is smaller than the value of time Tst1 (Tst1>Tst3>Tst2). Figure 6 (c) deviation and Figure 6 (a) and Figure 6 Compared to the deviation in (b), it converges more quickly.
[0128] By changing the method by which the reward calculation unit 71 calculates the reward r, the characteristics of the optimal instruction parameter 104 obtained as a result of learning can be selected. For example, in order to learn the instruction parameter 104 that enables the deviation to converge quickly, the reward calculation unit 71 can assign a large reward r if the time from the start of the action to the deviation falling into the allowable range is less than or equal to a predetermined time. Alternatively, the shorter the time from the start of the action to the deviation falling into the allowable range, the larger the reward r can be assigned. Furthermore, the reward calculation unit 71 can also calculate the reward r as the reciprocal of the time from the start of the evaluation operation to the deviation falling into the allowable range. Additionally, it can also be as follows... Figure 3 As shown in (b), when the deviation falls within the allowable range but then exceeds it, a small reward r is assigned, and the instruction parameter 104 that prevents the mechanical load 3 from oscillating is learned. The above is... Figure 6 Explanation of the operation example of the reward calculation unit 71 shown.
[0129] If the reward r is calculated, the value function update unit 72 updates the action value function Q accordingly. The intention determination unit 73 selects the action a that maximizes the updated action value function Q. t That is, the instruction parameter 104 with the largest value in the updated action value function Q is determined as the evaluation candidate parameter 108.
[0130] In addition, Figure 1 The description of the electric motor control device 1000 shows the case where reinforcement learning is performed on the learning algorithm used by the learning unit 7, but the learning algorithm in this embodiment is not limited to reinforcement learning. Teacher-assisted learning, teacherless learning, and semi-teacher-assisted learning algorithms can also be applied. Furthermore, deep learning, which learns by extracting the features themselves, can also be used as the learning algorithm. Additionally, machine learning can be performed using other methods, such as neural networks, genetic programming, functional logic programming, support vector machines, and Bayesian optimization.
[0131] Figure 7This diagram illustrates a structural example of the processing circuit of the motor control device 1000 in Embodiment 1, which is composed of a processor 10001 and a memory 10002. When the processing circuit is composed of a processor 10001 and a memory 10002, each function of the processing circuit of the motor control device 1000 is implemented by software, firmware, or a combination of software and firmware. The software or firmware is described as a program and stored in the memory 10002. In the processing circuit, each function is implemented by reading from the processor 10001 and executing the program stored in the memory 10002. That is, the processing circuit has a memory 10002, which stores programs that, as a result, cause the processing of the motor control device 1000 to be executed. Furthermore, these programs can be said to cause a computer to execute the processes and methods of the motor control device 1000.
[0132] Here, processor 10001 can also be a CPU (Central Processing Unit), processing device, arithmetic device, microprocessor, microcomputer, or DSP (Digital Signal Processor), etc. Memory 10002 can also be, for example, non-volatile or volatile semiconductor memory such as RAM (Random Access Memory), ROM (Read Only Memory), flash memory, EPROM (Erasable Programmable ROM), or EEPROM (Electrically EPROM). Additionally, memory 10002 can also be a hard disk, floppy disk, optical disk, high-density disk, mini-disk, or DVD (Digital Versatile Disc), etc.
[0133] Figure 8 This diagram illustrates a structural example of the processing circuit of the motor control device 1000 in Embodiment 1, which is constructed using dedicated hardware. In the case where the processing circuit is constructed using dedicated hardware, Figure 8The processing circuit 10003 shown can be, for example, a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. The functions of the motor control device 1000 can be implemented individually by the processing circuit 10003, or multiple functions can be implemented centrally by the processing circuit 10003. Furthermore, the motor control device 1000 and the controlled object 2000 can also be connected via a network. Additionally, the motor control device 1000 can also reside on a cloud server.
[0134] Alternatively, multiple control objects identical to control object 2000 can be set up to execute the evaluation process implemented by multiple control objects in parallel, enabling efficient learning. For example, in Figure 2 During the evaluation operation EV11, evaluation operations based on multiple controlled objects are performed in parallel, acquiring a set of data including multiple command parameters and evaluation sensor signals. Then, during the learning operation L12, the action value function Q is updated multiple times using the data acquired during evaluation operation EV11, determining multiple command parameters. Furthermore, during evaluation operation EV12, evaluation operations based on multiple controlled objects are performed using the multiple command parameters determined during learning operation L12. By performing the learning loop in this way, multiple evaluation operations can be performed in parallel. Additionally, the learning unit can use the method described later in Embodiment 4 for actions that determine multiple command parameters. Furthermore, during the repeated learning loop, some or all of the multiple controlled objects can be changed, and the number of controlled objects constituting the multiple controlled objects can be increased or decreased.
[0135] Alternatively, the motor control device 1000, which has learned using data obtained from the controlled object 2000, can be connected to other controlled objects, and further learning can be performed using data obtained from other controlled objects. Alternatively, the motor control device can be configured using a learning-completed learner equipped with the learning results of this embodiment. The aforementioned learning-completed learner can also be implemented using a learning-completed program that uses the action value function Q updated through learning to determine the instruction parameter 104. Alternatively, the aforementioned learning-completed learner can be implemented using learning-completed data that stores the results of adjusting the instruction parameter 104. According to the motor control device using the learning-completed learner, a motor control device capable of utilizing the learning results can be provided in a short time. Furthermore, the method described in this embodiment allows for the automatic adjustment of the instruction parameter 104 of the motor control device, and also allows for the manufacture of the motor control device. Moreover, the automatic adjustment in this embodiment only requires automating at least a portion of the adjustment operation; human operation or human intervention is not excluded.
[0136] As described above, the motor control device 1000 of this embodiment includes a drive control unit 4, a learning unit 7, and an adjustment management unit 9. The drive control unit 4 drives the motor 1 based on command parameters 104 (control commands), causing the controlled object 2000, consisting of the motor 1 and a mechanical load 3 mechanically connected to the motor 1, to operate. Furthermore, it performs an initialization operation to set the controlled object 2000 to an initial state and an evaluation operation from the initial state. The learning unit 7 learns by associating the command parameters 104 (control commands) used in the evaluation operation with the state sensor signal 101 obtained by detecting the state of the controlled object 2000 during the evaluation operation. Based on the learning results, it determines the command parameters 104 (control commands) used in the evaluation operation to be performed after the evaluation operation that has acquired the state sensor signal 101. The adjustment management unit 9 determines the timing for performing any one of the initialization operation, evaluation operation, and learning operation (i.e., the first step) based on the timing of the first step. Therefore, the execution timing of the first and second processes can be adjusted to shorten the waiting time and efficiently execute the adjustment of instruction parameter 104 (control instruction).
[0137] Furthermore, the motor control method of this embodiment drives the motor 1 based on the command parameter 104 (control command), causing the controlled object 2000, which consists of the motor 1 and the mechanical load 3 mechanically connected to the motor 1, to operate. It also performs an initialization operation to set the controlled object 2000 to an initial state and an evaluation operation from the initial state. Additionally, a learning operation is performed, that is, the command parameter 104 used for the evaluation operation and the state sensor signal 101 obtained by detecting the state of the controlled object 2000 during the evaluation operation are learned in association, and the command parameter 104 used for the evaluation operation to be performed after the evaluation operation in which the state sensor signal 101 is obtained is determined based on the learning result. Here, the learning operation is the operation from the start of learning until the command parameter 104 is determined. Furthermore, the timing of performing any one of the learning operation, the initialization operation, and the evaluation operation (i.e., the first step) is determined based on the timing of the first step. In this way, a motor control method capable of performing automatic adjustments can be provided efficiently.
[0138] Alternatively, the timing for executing the second step can be set to be simultaneous with or later than the timing for executing the first step. This allows the timing of the first step to be used in determining the timing of the second step, reliably shortening the interval between steps. Furthermore, if the time required for the first step changes, the timing of the second step can be adjusted accordingly. Here, it is preferable that the interval between the completion time of the first step and the start time of the second step be as short as feasible; it is even more preferable to set the completion time of the first step and the start time of the second step to be simultaneous or approximately simultaneous.
[0139] As described above, according to this embodiment, a motor control device can be provided such that, when automatically adjusting the control commands for controlling the motor by repeatedly performing initialization operation, evaluation operation, and learning operations, the time required for automatic adjustment can be shortened.
[0140] Implementation Method 2
[0141] Figure 9 This is a block diagram illustrating an example of the structure of the motor control device 1000a in Embodiment 2. Figure 9 (a) shows an example of the overall structure of the motor control device 1000a. Figure 9 (b) shows a structural example of the learning unit 7a. The motor control device 1000a replaces the embodiment 1. Figure 1 The learning unit 7 of the motor control device 1000 shown has a learning unit 7a instead of a learning unit 7a. Figure 1The adjustment management department 9 has an adjustment management department 9a instead of an adjustment management department 9a. The structure of the learning department 7a is derived from the structure of the learning department 7, omitting the learning completion signal determination department 74. Moreover, Figure 9 The adjustment management department 9a detects the completion time points of evaluation operation and initialization operation based on the status sensor signal 101. Furthermore, Figure 9 The Adjustment Management Department 9a uses the initial operation completion time when determining the start time of the evaluation operation. Figure 9 In the description of the electric motor control device 1000a shown, regarding the... Figure 1 Structural elements that are identical or corresponding are labeled with the same number.
[0142] Figure 10 This is a diagram illustrating an example of the timing of the operation of the motor control device 1000a in Embodiment 2. Figure 10 (a) to Figure 10 (d) The horizontal axis represents time. Figure 10 (a) to Figure 10 (d) The vertical axes represent the learning action, action processing (initialization operation and evaluation operation), learning start signal 106, and instruction start signal 105. The relationship between the values of instruction start signal 105 and learning start signal 106 and the content indicated by each signal is similar to that in embodiment 1. Figure 2 The same as described above.
[0143] exist Figure 10 In the example action, the time required for initialization is longer than the time required for learning the action. Furthermore, initialization is completed after the learning action. Therefore, the start time of the evaluation operation is determined based on the completion time of the initialization operation, not the completion time of the learning action. Moreover, the completion times of both the initialization and evaluation operations are detected based on the state sensor signal 101. At these points... Figure 2 The examples of actions are different.
[0144] Figure 11 This is a flowchart illustrating an example of the operation of the adjustment management unit 9a in Embodiment 2. (See attached diagram.) Figure 10 and Figure 11 The operation of the motor control device 1000a is illustrated below. If automatic adjustment is initiated, in step S201, the adjustment management unit 9a sets the value of the command start signal 105 at time TL211 to 1, and sets the start time of initialization operation IN21 to time TL211. Motor 1 begins initialization operation IN21 at time TL211 according to the command start signal 105. Then, initialization operation IN21 is completed at time TL213.
[0145] In step S202, the adjustment management unit 9a sets the value of the learning start signal 106 at time TL211 to 1, and sets the start time of learning action L21 to time TL211. The learning unit 7a starts learning action L21 at time TL211 according to the learning start signal 106. Then, learning action L21 is completed at time TL212. Figure 2 Similar to learning action L11, in learning action L21, the learning unit 7a can also determine the instruction parameter 104 based on a preset setting or randomly. Initialization operation IN21 is executed in parallel with learning action L21. Since the time required for initialization operation IN21 is longer than the time required for learning action L21, time TL213 becomes a later time point than time TL212. Figure 2 Similarly, the start time of the learning action L21 can be later than the start time of the initialization operation IN21, without prolonging the waiting time.
[0146] In step S203, the adjustment management unit 9a detects time TL213 as the completion time of initialization operation IN21 based on the status sensor signal 101. In step S204, the adjustment management unit 9a determines the value of the command start signal 105 at time TL213 to be 1 based on the detected completion time of initialization operation IN21, thus determining the start time of evaluation operation EV21 (first evaluation operation). Motor 1 starts evaluation operation EV21 at time TL213 according to the command start signal 105. Then, evaluation operation EV21 is completed at time TL221.
[0147] In step S205, the adjustment management unit 9a, based on the status sensor signal 101, detects time TL221 as the completion time point for evaluating operation EV21. Then, in step S206, it... Figure 3 Similar to step S106, a decision is made on whether to continue automatic adjustment. In step S206 executed at time TL221, the adjustment management unit 9a determines that automatic adjustment should continue and proceeds to step S207. The period from time TL211 to time TL221 is set as the learning loop CYC21.
[0148] In step S207, the adjustment management unit 9a sets the values of the command start signal 105 and the learning start signal 106 at time TL221 to 1 based on the completion time of the evaluation operation EV21. Then, through this action, time TL221 is determined as the start time of the initialization operation IN22 (first initialization operation) and the learning operation L22 (first learning operation). The motor 1 and the learning unit 7a each start the initialization operation IN22 and the learning operation L22 according to the command start signal 105 and the learning start signal 106, respectively. The initialization operation IN22 and the learning operation L22 are executed in parallel.
[0149] Then, steps S203 to S207 are repeated until the adjustment management unit 9a determines in step S206 that automatic adjustment will not continue. Then, in step S204 of the learning loop CYC22, the adjustment management unit 9a sets the value of the command start signal 105 at time TL223 to 1 based on the completion time point of the initialization operation IN22, i.e., TL223. Then, through this action, time TL223 is determined as the start time point of the evaluation operation EV22 (the second evaluation operation). Motor 1 begins the evaluation operation EV22 at time TL223 according to the command start signal 105.
[0150] In step S205 of the final learning cycle, namely learning cycle CYC2X, the adjustment management unit 9a detects time TL2X2 as the completion time point for evaluating the operation of EV2X. Then, in step S206, it is determined that automatic adjustment will not continue, and proceeds to step S208. In step S208, the adjustment management unit 9a and... Figure 3 Similar to step S108, the learning unit 7a is instructed to end processing T2. The learning unit 7a and... Figure 2 The same as the termination process T1 is executed in the same way as the termination process T2. Furthermore, similar to Embodiment 1, this embodiment allows multiple control objects identical to the controlled object 2000 to perform evaluation operations in parallel, efficiently performing automatic adjustments. Additionally, a motor control device can be constructed using a learning-completed learner equipped with the learning results obtained from this embodiment. Furthermore, through the learning in this embodiment, automatic adjustments of control commands for controlling the motor can be performed, and the manufacture of a motor control device can also be carried out.
[0151] Furthermore, when the adjustment management unit 9a detects the completion of operation in step S203 or S205, the completion of operation can also be detected by detecting whether the difference between the status sensor signal 101 indicating the position of the motor 1 and the target movement distance, i.e., the deviation, is less than or equal to a predetermined reference value. In addition to the deviation being less than or equal to the reference value, operation can also be determined to be complete when the deviation does not exceed the reference value within a predetermined time period. Furthermore, the adjustment management unit 9a is not limited to the status sensor signal 101; signals obtained by detecting the state of the controlled object 2000 can also be used to detect the completion time point of operation. Moreover, the command signal 103 can also be used to detect the completion time point of operation.
[0152] According to this embodiment, a motor control device can be provided such that, when automatically adjusting the control commands for controlling the motor by repeatedly performing initialization operation, evaluation operation, and learning actions, the time required for automatic adjustment can be shortened.
[0153] Evaluation operation EV21 (first evaluation operation), which is one of the evaluation operations, can also be executed, and learning operation L22 (first learning operation) can be performed using the state sensor signal 101 obtained during evaluation operation EV21. Furthermore, initialization operation IN22 (first initialization operation) can be executed in parallel with learning operation L22. Starting from the initial state set by initialization operation IN22, the next evaluation operation after evaluation operation EV21, namely evaluation operation EV22 (second evaluation operation), is executed based on the command parameter 104 (control command) determined by learning operation L22. Through this operation, learning operation L22 and initialization operation IN22 can be executed in parallel, shortening the time required for automatic adjustment. Therefore, a motor control device 1000a or motor control method capable of efficiently performing automatic adjustment can also be provided.
[0154] Furthermore, the adjustment management unit 9a can also detect the completion time of the evaluation operation EV21, and based on the detected completion time, determine the start time of the learning action L22 and the start time of the initialization operation IN22, thereby shortening the waiting time between processes. Additionally, the adjustment management unit 9a can also determine the start time of the one requiring more time between the learning action L22 and the initialization operation IN22 to be simultaneous with or earlier than the other's start time, further shortening the waiting time between processes. Furthermore, the adjustment management unit 9a can also detect the completion time of the one completing simultaneously or later between the initialization operation IN22 and the learning action L22, and based on the detected completion time, determine the start time of the evaluation operation EV22, further shortening the waiting time between processes. Moreover, if two consecutively executed processes are designated as a preceding and a following process, it is preferable to keep the interval between the completion time of the preceding process and the start time of the following process as short as feasible; setting them to be simultaneous or approximately simultaneous is even more preferable. Furthermore, the drive control unit 4 can drive the motor 1 by following the command values for controlling the motor 1, namely the command values for position, speed, acceleration, current, torque, or thrust, i.e., the command signal 103. It can detect the completion time of the evaluation operation or initialization operation by detecting the signal or command signal 103 obtained from detecting the state of the controlled object 2000, and can detect the completion time of the operation with high precision. Moreover, if the time required for operation changes, the start time of the next process can be accurately determined, and this can be used to shorten the time required for automatic adjustment. As described above, a motor control device 1000a or a motor control method capable of efficiently performing automatic adjustment can also be provided.
[0155] Implementation Method 3
[0156] Figure 12 This is a block diagram illustrating an example of the structure of the motor control device 1000b in Embodiment 3. Figure 12 (a) shows an example of the overall structure of the motor control device 1000b. Figure 12 (b) shows a structural example of the learning unit 7b. The structure of the motor control device 1000b, except that it replaces the learning unit 7a and has a learning unit 7b, is similar to that of Embodiment 2. Figure 9 The motor control device 1000a shown is the same. This embodiment... Figure 12 The structural elements shown are similar to those in Implementation 2. Figure 9 Structural elements that are identical or corresponding to those shown are labeled with the same number.
[0157] In addition to the Study Department 7b Figure 9In addition to the structural elements of the learning unit 7a in (b), there is a learning limit time determination unit 77. The learning limit time determination unit 77 calculates the estimated time required for initialization operation as the estimated time required for initialization operation. Moreover, based on the estimated time required for initialization operation, the upper limit of the learning time, i.e., the time for the learning unit 7b to perform the learning action, is determined as the learning limit time TLIM1. The learning limit time determination unit 77 can also determine the learning limit time TLIM1 to be the same as or shorter than the estimated time required for initialization operation. Moreover, the learning unit 7b can also perform the learning action within a period that is the same as or shorter than the learning limit time TLIM1. By performing the learning action in this way, the learning action can be completed before the initialization operation is completed. Here, the learning unit 7b can also obtain the estimated time required for initialization operation from an external source. In addition, the learning unit 7b can also calculate the measured value of the time required for initialization operation based on the status sensor signal 101, the command signal 103, etc., and use the measured value to estimate or update the estimated time required for initialization operation.
[0158] The learning time limit determination unit 77 can further predetermine the basic learning time TSL1. The basic learning time TSL1 is the lower limit of the learning time, and the learning unit 7b can also perform the learning action for the same or longer time as the basic learning time TSL1. For example, the basic learning time TSL1 can be set as the minimum time for determining the instruction parameter 104, or it can be set as the minimum time for determining the instruction parameter 104 with the desired accuracy. The learning time limit determination unit 77 can also further determine the additional learning time TAD1 based on the basic learning time TSL1 and the learning time limit TLIM1, such that the sum of the basic learning time TSL1 and the additional learning time TAD1 does not exceed the learning time limit TLIM1. This condition is expressed by equation (26). In addition, the learning time limit TLIM1 is set to be longer than the basic learning time TSL1.
[0159] [Mathematical Expression 26]
[0160] TSL1+TAD1<TLIM1···(26)
[0161] The learning unit 7b performs learning during the basic learning time TSL1. Furthermore, it can also perform learning actions during the additional learning time TAD1 to improve the accuracy of the instruction parameter 104. The learning unit 7b can utilize the basic learning time TSL1 to perform learning for a predetermined lower limit. Alternatively, it can set only the learning limit time TLIM1 without setting the basic learning time TSL1 and the additional learning time TAD1. Additionally, the learning limit time determination unit 77 can store the estimated initialization operation time, the learning limit time TLIM1, the basic learning time TSL1, and the additional learning time TAD1 in a storage device.
[0162] Next, the relationship between learning time and the accuracy of instruction parameters determined in the learning action will be described. For example, in the case of using Q-learning as a learning algorithm, the intention determination unit 73 will increase the value of the action value function Q in the action a. t 108 is selected as the candidate evaluation parameter. When making this selection, if the action value function Q is a continuous function, the intention determination unit 73 sometimes performs iterative calculations. Thus, when iterative calculations are performed during the learning action, the intention determination unit 73 can improve calculation accuracy by ensuring a longer calculation time and increasing the number of calculation steps. As described above, the effects of this embodiment are more significant when the learning action includes iterative calculations. Furthermore, examples of iterative calculations include methods that numerically calculate the gradient, such as the fastest descent method or Newton's method, and methods that use probabilistic factors, such as the Monte Carlo method.
[0163] Figure 13 This is a diagram illustrating an example of the timing of the operation of the motor control device 1000b in Embodiment 3. Figure 13 (a) to Figure 13 (d) The horizontal axis represents time. Figure 13 (a) to Figure 13 (d) represents the learning action, action processing (initialization operation and evaluation operation), learning start signal 106 and instruction start signal 105. Figure 13 The relationship between the values of the instruction start signal 105 and the learning start signal 106 and the timing of the actions indicated by each signal, and the relationship between these values and the timing of the actions indicated by each signal in Embodiment 1. Figure 2 The same as described above. Figure 13 The operation of the motor control device 1000b shown, in addition to the operation of the learning unit 7b, is related to... Figure 10 Same. Figure 13 In the middle, to and Figure 10 Same or corresponding operation, learning, learning cycle, time, etc. annotations and Figure 10 Same label. Also... Figure 13The flowchart of the operation of the adjustment management department 9a in the action example and the implementation method 2. Figure 11 Same. (Refer to...) Figure 11 and Figure 13 An example of the operation of the motor control device 1000b will be explained.
[0164] exist Figure 13 In the example operation, the learning limit time determination unit 77 calculates the estimated time required for initialization operation based on the measured value of the time required for initialization operation IN21. Furthermore, it determines the learning limit time TLIM1 to be the same as or shorter than the estimated time required for initialization operation. Moreover, the learning limit time determination unit 77 determines the basic learning time TSL1 as the lower limit of the learning time, and sets the difference between the learning limit time TLIM1 and the basic learning time TSL1 as the additional learning time TAD1.
[0165] exist Figure 13 In the example of the action, since only the action of the learning unit 7b is similar to that of implementation method 2... Figure 10 Since the learning loop CYC22 is different, the operation of the learning unit 7b will be explained using the example of the learning loop CYC22. The learning unit 7b starts learning action L22 (the first learning action) at time TL221 according to the learning start signal 106 determined in step S202 of the learning loop CYC22. Here, the learning unit 7b executes sub-learning action L221 and sub-learning action L222 as learning action L222. The length of sub-learning action L221 is the basic learning time TSL1. Moreover, the length of sub-learning action L222 is the additional learning time TAD1. Moreover, the learning unit 7b completes learning action L22 at time TL222, which is the point in time from time TL221 when the basic learning time TSL1 and the additional learning time TAD1 have elapsed. Here, the value of time TL222 is equal to the sum of the three values: the value of time TL221, the basic learning time TSL1, and the additional learning time TAD1, and the relationship of equation (27) holds.
[0166] [Mathematical Expression 27]
[0167] TL222=TL221+TSL1+TAD1···(27)
[0168] exist Figure 13In the example operation, the start time of the initialization operation and the start time of the learning operation are simultaneous. However, if the time required for the initialization operation is longer than the time required for the learning operation, the learning operation can start later than the initialization operation. The learning limit time determination unit 77 can also determine the learning limit time TLIM1 by ensuring that the time elapsed from the start time of the initialization operation IN22, which is the estimated time required for the initialization operation, is later than the time elapsed from the start time of the learning operation L22 (the first learning operation), which is the learning limit time TLIM1. Furthermore, the learning unit 7b can also execute the learning operation L22 within a period that is the same as or shorter than the learning limit time TLIM1. In this way, even if the start time of the learning operation L22 is later than the start time of the initialization operation IN22, the learning operation L22 can be completed before the initialization operation IN22 is completed. In this case, it is not necessary to wait for the completion of the learning operation L22, and the evaluation of operation EV22 can begin immediately after the initialization operation IN22 is completed. Therefore, no increase in delay time caused by waiting for the completion of the learning operation L22 will occur. Therefore, the time required for automatic adjustment can be shortened. Consequently, a motor control device 1000b or motor control method capable of efficiently performing automatic adjustment can also be provided.
[0169] In addition to the learning limit time TLIM1, the learning limit time determination unit 77 can also determine the lower limit of the learning time, namely the basic learning time TSL1. Furthermore, the learning unit 7b can also perform the learning action L22 during a period that is the same as or longer than the basic learning time TSL1, and the same as or shorter than the learning limit time TLIM1. If the learning action is performed in this way, the learning time, which is predetermined as the lower limit, can be ensured using the learning limit time TLIM1. Moreover, for example, if the basic learning time TSL1 is set as the minimum time required to obtain the instruction parameter 104, the instruction parameter 104 can be calculated with higher probability on a per-learning-cycle basis. As described above, a motor control device 1000b or motor control method capable of efficiently performing automatic adjustments can also be provided.
[0170] According to this embodiment, a motor control device can be provided such that, when automatically adjusting the instruction parameters 104 (control instructions) for controlling the motor 1 by repeatedly performing initialization operation, evaluation operation, and learning actions, the time required for automatic adjustment can be shortened.
[0171] Implementation Method 4
[0172] Figure 14This is a block diagram illustrating an example of the structure of the motor control device 1000c in Embodiment 4. Figure 14 (a) shows an example of the overall structure of the motor control device 1000c. Figure 14 (b) represents the structure example of the learning section 7c. Figure 14 The motor control device 1000c shown is a replacement Figure 1 The motor control device 1000 of Embodiment 1 shown has a learning unit 7c instead of a learning unit 7, and an adjustment management unit 9b instead of an adjustment management unit 9. Furthermore, besides... Figure 1 In addition to the structural elements of the motor control device 1000, it has a learning time estimation unit 10. Figure 14 In the description of the electric motor control device 1000c shown, regarding the embodiment 1... Figure 1 or Figure 5 Identical or corresponding structural elements are labeled with the same number.
[0173] In this implementation, various learning algorithms can be applied, but an example is given using reinforcement learning implemented by Q-learning. Figure 14 The learning section 7c shown is replaced Figure 5 The learning unit 7 of the embodiment 1 shown has an intention determination unit 73a. Figure 5 In one learning operation, the learning unit 7c acquires one set of command parameters 104 for evaluating operation and the status sensor signals 101 during evaluation operation, and executes a decision on the command parameters 104. On the other hand, the learning unit 7c acquires multiple sets of the above in one learning cycle. Then, the reward calculation unit 71 and the value function update unit 72 calculate the reward r and update the action value function Q based on the calculated reward r for each acquired set. As a result, the learning unit 7c performs the calculation of reward r and the update of action value function Q multiple times in one learning cycle.
[0174] The intention determination unit 73a determines multiple evaluation candidate parameters 108 based on the action value function Q that has been updated multiple times and the multiple data sets used for the update. Then, the instruction parameter determination unit 75 determines the instruction parameters 104 used for the evaluation operation after the learning action being performed based on the determined evaluation candidate parameters 108.
[0175] The operation of the intention determination unit 73a will be explained. The intention determination unit 73a obtains the action value function Q(s) of equation (25) updated by the value function update unit 72. t a t Then, for multiple actions a t That is, multiple instruction parameters 104 contained in multiple data sets are used to calculate the value of the corresponding action value function Q. Here, when action a is selected...t When (instruction parameter 104) is given, assign a value function Q(s) to a certain action. t a t Given the value of ), action a t (Instruction parameter 104) and action value function Q(s) t a t The values of the action value functions Q correspond to each other. Furthermore, from the calculated values of multiple action value functions Q, a predetermined number of action value function Q values are selected in descending order. Then, the instruction parameter 104 corresponding to the selected action value function Q values is determined as the evaluation candidate parameter 108. This is an example of the operation of the intention determination unit 73a. Furthermore, the number of instruction parameters 104 determined by the instruction parameter determination unit 75 can also be the same as the number of evaluation runs performed in the next learning cycle after the currently executed learning action.
[0176] Next, the learning time estimation unit 10 will be described. The learning time estimation unit 10 calculates an estimated learning time based on an estimated value of the learning time for the performed learning action and outputs an estimated learning time signal 109 representing the estimated learning time. Furthermore, the learning time estimation unit 10 can also acquire the learning start signal 106 and the learning completion signal 107 for the performed learning action, and obtain a measured value of the learning time based on the difference between the learning start time and the learning completion time. Moreover, it can also calculate an estimated learning time based on the measured value of the learning time for the performed learning action. Additionally, the learning time estimation unit 10 can acquire the estimated learning time from external input, and can also update the estimated learning time based on the measured value of the learning time.
[0177] Next, the adjustment management unit 9b will be explained. The adjustment management unit 9b determines the learning start signal 106 based on the learning completion signal 107, and thus determines the start time of the next learning action based on the completion time of the learning action. Furthermore, the adjustment management unit 9b pre-determines the time required for initialization operation (initialization operation time) and the time required for evaluation operation (evaluation operation time). Moreover, by detecting the passage of the initialization operation time and evaluation operation time from the start time of the initialization operation and evaluation operation, the completion time of the initialization operation and evaluation operation is detected respectively. Furthermore, based on the detected completion time of the initialization operation and evaluation operation, the start time of the next evaluation operation and initialization operation to be executed is determined respectively. Here, the adjustment management unit 9b can also, as in the adjustment management unit 9a of embodiment 2, accurately detect the completion time of the initialization operation and evaluation operation based on the signal or command signal 103 obtained by detecting the state of the controlled object 2000. Here, the operation of the motor 1, consisting of the initialization operation and the evaluation operation starting from the initial state set by the initialization operation, is called the evaluation operation cycle. The Adjustment Management Department (9b) determines whether the evaluation cycle has been completed at each evaluation operation's completion time point. Hereinafter, the completion time point of the evaluation operation will sometimes be referred to as the judgment time point.
[0178] Figure 15 This is a diagram illustrating an example of the timing of the operation of the motor control device 1000c in Embodiment 4. Figure 15 (a) to Figure 15 (e) The horizontal axis represents time. Figure 15 (a) to Figure 15 (e) represents the learning action, action processing (initialization operation and evaluation operation), learning start signal 106, learning completion signal 107, and instruction start signal 105. The relationship between the values of the learning start signal 106, learning completion signal 107, and instruction start signal 105 and the timing of the learning action or operation indicated by each signal is consistent with that in Implementation Method 1. Figure 2 The same as described above. Figure 16 This is a flowchart illustrating an example of the operation of the adjustment management unit 9b in embodiment 4. Figure 15 In one learning cycle, one learning action is performed, and two evaluation operation cycles are performed in parallel with the learning action. However, the number of evaluation operation cycles performed in parallel with the learning action can be greater than or equal to three.
[0179] use Figure 15 and Figure 16The operation of the motor control device 1000c is illustrated below. If automatic adjustment is initiated, in step S401, the adjustment management unit 9b sets the value of the learning start signal 106 at time TL411 to 1, thus setting time TL411 as the start time of learning action L41 (the third learning action). The learning unit 7c begins learning action L41 at time TL411 according to the learning start signal 106. In step S402, the adjustment management unit 9b sets the value of the command start signal 105 at time TL411 to 1 based on the start time of learning action L41, thus setting time TL411 as the start time of initialization operation IN41. The motor 1 begins initialization operation IN41 at time TL411 according to the command start signal 105. Then, the motor 1 completes initialization operation IN41 at time TL412 and enters standby mode after completing initialization operation IN41. Here, in step S402, the adjustment management unit 9b determines the start time of the first evaluation operation cycle ECYC1 (first evaluation operation cycle) by determining the start time of the initialization operation IN41.
[0180] In step S403, the adjustment management unit 9b detects the time elapsed since time TL411 for the initialization operation, and determines time TL413 as the completion time of the initialization operation IN41. In step S404, based on the detected completion time of the initialization operation IN41, the adjustment management unit 9b sets the value of the command start signal 105 at time TL413 to 1, and determines time TL413 as the start time of the evaluation operation EV41. The motor 1 begins the evaluation operation EV41 at time TL413 according to the command start signal 105. Then, the motor 1 completes the evaluation operation EV41 at time TL414, and enters a standby state after the evaluation operation EV41 is completed.
[0181] In step S405, the adjustment management unit 9b detects whether the time required for the evaluation operation has elapsed since time TL413, and sets time TL415 as the completion time of evaluation operation EV41. In step S406, the adjustment management unit 9b determines whether the currently executing evaluation operation cycle has been completed. If the evaluation operation cycle is not completed, the process proceeds to step S407; if the evaluation operation cycle is completed, the process proceeds to step S408.
[0182] An example of the judgment in step S406 is given. The adjustment management unit 9b pre-determines an estimated value for the time required for one evaluation operation cycle, i.e., the estimated evaluation operation cycle time. At the judgment time point, the adjustment management unit 9b obtains the estimated learning time signal 109 and calculates the time point from the start time of the learning action L41 that has elapsed since the estimated learning time, i.e., the estimated learning time elapsed time point. Furthermore, if the time from the completion time of the evaluation operation (i.e., the judgment time point) to the estimated learning time elapsed time point is shorter than the estimated evaluation operation cycle time, the adjustment management unit 9b determines that the evaluation operation cycle ECYC1 has been completed. Conversely, if the time from the aforementioned judgment time point to the estimated learning time elapsed time point is longer than or equal to the estimated evaluation operation cycle time, the adjustment management unit 9b determines that the evaluation operation cycle ECYC1 has not yet been completed. In other words, if the adjustment management unit 9b cannot execute one evaluation operation cycle within the remaining time until the estimated learning time elapsed time point, it determines that the evaluation operation cycle ECYC1 has been completed. Furthermore, if one evaluation cycle can be performed within the remaining time, it is determined that the evaluation cycle ECYC1 has not yet been completed. The above is an example of the judgment in step S406.
[0183] In step S406 at time TL415, the adjustment management unit 9b determines that the evaluation operation cycle ECYC1 has not yet been completed and proceeds to step S407. In step S407, based on the completion time of evaluation operation EV41, the adjustment management unit 9b sets the value of the command start signal 105 at time TL415 to 1, and sets time TL415 as the start time of initialization operation IN42. Motor 1 begins initialization operation IN42 at time TL415 according to the command start signal 105. Afterwards, the adjustment management unit 9b repeatedly executes steps S403 to S407 until it determines in step S406 that the evaluation operation cycle ECYC1 has been completed.
[0184] At the judgment point TL421, the adjustment management unit 9b executes the judgment in step S406, determining that the evaluation operation cycle ECYC1 has been completed, and proceeds to step S408. In step S408, the adjustment management unit 9b, based on the learning completion signal 107, detects time TL421 as the completion point of learning action L41. Next, in step S409, the adjustment management unit 9b and the implementation method 1... Figure 3 Similar to step S106, a judgment is made on whether to continue automatic adjustment. If the judgment is to continue automatic adjustment, the process proceeds to step S410; if the judgment is not to continue automatic adjustment, the process proceeds to step S411. In step S409 at time TL421, the adjustment management unit 9b determines that automatic adjustment should continue.
[0185] Here, the learning cycle CYC41 is the period from time TL411 to time TL421. Furthermore, the evaluation operation cycle ECYC1 begins from a state where no learning action has been executed even once. Therefore, evaluation operations EV41 and EV42 can be executed using preset command parameters 104 or randomly determined command parameters 104. Additionally, in learning action L41, similar to learning action L11 in Embodiment 1, the command parameters 104 can be randomly determined or determined based on preset parameters.
[0186] In step S410, the adjustment management unit 9b sets the value of the learning start signal 106 at time TL421 to 1 based on the completion time of learning action L41, and sets time TL421 as the start time of learning action L42 (the fourth learning action). The learning unit 7c starts learning action L42 at time TL421 according to the learning start signal 106. Learning action L42 is executed based on the instruction parameter 104 used in the evaluation operation cycle ECYC1 and the status sensor signal 101 obtained in the evaluation operation cycle ECYC1. Afterwards, the adjustment management unit 9b repeatedly executes steps S402 to S410 until it is determined in step S409 that automatic adjustment will not continue. Here, the evaluation operation cycle ECYC2 (the second evaluation operation cycle) is executed using the instruction parameter 104 determined in learning action L41. In addition, in step S402, by determining time TL421 as the start time of initialization operation IN43, the adjustment management unit 9b determines time TL421 as the start time of evaluation operation cycle ECYC2.
[0187] In the judgment of step S409 of TL4X3 during the learning cycle CYC4Z, the adjustment management unit 9b determines that automatic adjustment will not continue and proceeds to step S411. In step S411, the adjustment management unit 9b and the implementation method 1... Figure 3 Step S108 similarly indicates the end of processing T4. Furthermore, the learning unit 7c is the same as in Embodiment 1. Figure 2 Similarly to the termination process T1, termination process T4 is executed.
[0188] Furthermore, in this embodiment, similar to Embodiment 1, multiple control objects identical to the controlled object 2000 can be executed in parallel for evaluation operations, efficiently performing automatic adjustments. For example, if in Figure 15During the learning action L41, multiple controlled objects are executed evaluation operation cycles in parallel. This allows for the acquisition of more sets of state sensor signals 101 and command parameters 104 within a single evaluation operation cycle, thus enabling efficient learning. Furthermore, a learning-completed learner equipped with the learning results of this embodiment can be used to construct a motor control device. Additionally, by executing the learning of this embodiment, automatic adjustment of control commands for controlling the motor and the manufacture of the motor control device can be performed. Moreover, a motor control method capable of performing automatic adjustment can be efficiently provided.
[0189] Alternatively, learning action L41 (the third learning action), which is one of the learning actions, can be executed multiple times in parallel with learning action L41. Furthermore, using the state sensor signal 101 obtained during evaluation operation cycle ECYC1, the next learning action after learning action L41, namely learning action L42 (the fourth learning action), can be executed. Moreover, using the command parameter 104 (control command) determined by learning action L41, the next evaluation operation cycle after evaluation operation cycle ECYC1, namely evaluation operation cycle ECYC2 (the second evaluation operation cycle), can be executed multiple times in parallel with learning action L42. Through such actions, the evaluation operation cycle can be executed multiple times during one learning action, efficiently obtaining the set of command parameter 104 and evaluation sensor signal 102, thus shortening the time required for automatic adjustment. Therefore, a motor control device 1000c or motor control method capable of efficiently performing automatic adjustment can be provided.
[0190] Furthermore, the adjustment management unit 9b can determine the start time of learning action L42 based on the completion time of learning action L41, and determine the start times of evaluation operation cycle ECYC1 and evaluation operation cycle ECYC2 based on the start times of learning action L41 and learning action L42, respectively. Through this operation, the relationship between the timing periods of the two learning actions can be adjusted, as can the relationship between the execution timing of the learning action and the execution timing of the evaluation operation cycle. Moreover, waiting time can be shortened in this way. Therefore, a motor control device 1000c or motor control method capable of efficiently performing automatic adjustments can be provided.
[0191] Furthermore, the motor control device 1000c also includes a learning time estimation unit 10, which estimates the time required for learning action L41 as the estimated learning time. Moreover, the adjustment management unit 9b can also pre-determine the estimated value of the time required to execute the evaluation operation cycle as the estimated time required for the evaluation operation cycle. Furthermore, at the completion time of the evaluation operation cycle ECYC1 (i.e., the judgment time point), if the difference between the estimated learning time and the time elapsed from the start time of learning action L21 to the judgment time point is the same as or longer than the estimated time required for the evaluation operation cycle, the adjustment management unit 9b determines that the evaluation operation cycle ECYC1 should continue to be executed; if the difference is shorter than the estimated time required for the evaluation operation cycle, the evaluation operation cycle ECYC1 should not continue to be executed. Through such operations, the number of evaluation operation cycles can be increased within the range where the evaluation operation cycle can be completed by the completion time of the learning time. Furthermore, when changes occur in the estimated learning time, the estimated time required for the evaluation cycle, etc., the number of times the evaluation cycle is executed can be adjusted accordingly, thus enabling efficient automatic adjustment. Therefore, a motor control device 1000c or a motor control method capable of efficiently performing automatic adjustment can also be provided.
[0192] In addition, Figure 15 In the example operation, the adjustment management unit 9b determines the completion time of the initialization operation IN41 based on the start time of the initialization operation IN41 and the time required for the initialization operation. This embodiment is not limited to this operation. For example, sometimes, during the period from the completion of the first operation as a process to the start of the second operation as a process, an intermediate process, including any one of initialization operation, evaluation operation, or learning operation, is performed. In such cases, the adjustment management unit 9b may also estimate the time required to perform the intermediate process in advance and determine the start time of the second process as a time point later than the time point after the estimated time required to perform the intermediate process from the completion time of the first process. By adjusting the start time of the second process with the estimated time required for the intermediate process as the target, the waiting time is shortened, thereby reducing the time required for automatic adjustment. Therefore, a motor control device 1000c or motor control method capable of efficiently performing automatic adjustment can also be provided.
[0193] As described above, according to this embodiment, a motor control device can be provided such that, when automatically adjusting the control commands for controlling the motor by repeatedly performing initialization operation, evaluation operation, and learning operations, the time required for automatic adjustment can be shortened.
[0194] Explanation of the label
[0195] 1. Electric motor; 2. Command generation unit; 3. Mechanical load; 4. Drive control unit; 7. Learning units (7a, 7b, 7c); 9. Adjustment and management units (9a, 9b); 10. Learning time estimation unit; 77. Learning limit time determination unit; 101. Status sensor signal; 103. Command signal; 1000, 1000a, 1000b, 1000c. Electric motor control device; 2000. Controlled object; ECYC1, ECYC2. Evaluation operation cycle; EV11, EV12, EV21, EV22, EV41, EV42, EV43, EV44. Evaluation operation; IN12, IN22, IN41, IN42, IN43, IN44. Initialization operation; L12, L22, L23, L41, L42. Learning action; TLIM1. Learning limit time; TSL1. Basic learning time.
Claims
1. An electric motor control device, comprising: The drive control unit drives the motor based on control commands, causing the controlled object, which consists of the motor and a mechanical load mechanically connected to the motor, to operate, and performs initialization operation to set the controlled object to an initial state and evaluation operation starting from the initial state. The learning unit learns in association the control commands used for the evaluation operation and the state sensor signals obtained by detecting the state of the controlled object during the evaluation operation, and determines the control commands for the evaluation operation to be executed after the evaluation operation has obtained the state sensor signals based on the learning results. as well as The management department adjusts the timing of the execution of the second step, which is any one of the learning action, the initialization operation, or the evaluation operation, based on the timing of the first step, which is the action of the learning department.
2. The motor control device according to claim 1, characterized in that, The first evaluation operation, which is one of the evaluation operations, is executed. The first learning action, which is the learning action, is executed using the state sensor signal obtained during the first evaluation operation. The first initialization operation, which is performed in parallel with the first learning action, is the first initialization operation. Starting from the initial state set by the first initialization operation, the next evaluation operation, namely the second evaluation operation, is executed based on the control command determined by the first learning action after the first evaluation operation.
3. The motor control device according to claim 2, characterized in that, The adjustment management department detects the completion time of the first evaluation operation, and based on the detected completion time of the first evaluation operation, determines the start time of the first learning action and the start time of the first initialization operation.
4. The motor control device according to claim 2, characterized in that, The adjustment management department determines the start time of the first learning action and the first initialization operation, which requires a longer time, to be at the same time as or earlier than the start time of the other.
5. The motor control device according to claim 3, characterized in that, The adjustment management department determines the start time of the first learning action and the first initialization operation, which requires a longer time, to be at the same time as or earlier than the start time of the other.
6. The motor control device according to any one of claims 2 to 5, characterized in that, The adjustment management department detects the completion time of either the first learning action or the first initialization operation, which is completed simultaneously or later, and determines the start time of the second evaluation operation based on the detected completion time.
7. The motor control device according to any one of claims 2 to 5, characterized in that, The time required for the first initialization operation is longer than the time required for the first learning action. The motor control device has a learning limit time determination unit. This learning limit time determination unit determines the time for executing the learning action, i.e., the upper limit of the learning time, i.e. the learning limit time, in such a way that the estimated time required for initial operation has elapsed since the start time of the first initial operation is later than the time point after which the learning limit time has elapsed since the start time of the first learning action. The learning unit executes the first learning action during a period that is the same as or shorter than the learning limit time.
8. The motor control device according to claim 7, characterized in that, The learning time limit determination unit further determines a time shorter than the learning time limit, namely a basic learning time, which is the lower limit of the learning time. The learning unit performs the first learning action during a period that is the same as or longer than the basic learning time.
9. The motor control device according to claim 1, characterized in that, Perform the third learning action, which is one of the learning actions. In parallel with the third learning action, the first evaluation operation loop, which is one of the evaluation operation loops consisting of the initialization operation and the evaluation operation, is executed multiple times. The fourth learning action, which follows the third learning action, is performed using the state sensor signal obtained during the first evaluation cycle. Using the control command determined by the third learning action, the next evaluation cycle, namely the second evaluation cycle, is executed multiple times in parallel with the fourth learning action, following the first evaluation cycle.
10. The motor control device according to claim 9, characterized in that, The adjustment management department determines the start time of the fourth learning action based on the completion time of the third learning action, and determines the start time of the first evaluation cycle and the second evaluation cycle based on the start times of the third and fourth learning actions, respectively.
11. The motor control device according to claim 9 or 10, characterized in that, It also includes a learning time estimation unit, which estimates the time required for the third learning action as the estimated learning time. The adjustment management department pre-determines the estimated time required to execute the evaluation cycle as the estimated time required for the evaluation cycle. The adjustment management department determines whether to continue executing the first evaluation cycle at the time point when the first evaluation cycle is completed (i.e., the judgment time point). If the difference between the estimated learning time and the time elapsed from the start time of the third learning action to the judgment time point is the same as or longer than the estimated time required for the evaluation cycle, the adjustment management department determines whether to continue executing the first evaluation cycle. If the difference is shorter than the estimated time required for the evaluation cycle, the adjustment management department determines whether to stop executing the first evaluation cycle.
12. The motor control device according to any one of claims 1 to 5, characterized in that, During the period from the completion of the first step to the commencement of the second step, intermediate steps are performed, including at least one of the initialization operation, the evaluation operation, or the learning action. The adjustment management department pre-estimates the time required to execute the intermediate process and determines the start time of the second process to be a time point later than the time point after the estimated time required to execute the intermediate process has elapsed since the completion time of the first process.
13. The motor control device according to any one of claims 1 to 5, characterized in that, The drive control unit drives the motor in a manner that follows command signals, which are command values for controlling the motor, such as position, speed, acceleration, current, torque, or thrust. The adjustment management department detects the timing of the evaluation operation or the initialization operation based on the detection results obtained by detecting the state of the controlled object or the instruction signal.
14. A method for controlling an electric motor, wherein, Based on control commands, the motor is driven to cause the controlled object, consisting of the motor and a mechanical load mechanically connected to the motor, to operate, performing initialization operation to set the controlled object to an initial state and evaluation operation starting from the initial state. The learning action involves associating the control commands used for the evaluation operation with the state sensor signals obtained by detecting the state of the controlled object during the evaluation operation, and learning based on the learning results to determine the control commands used for the evaluation operation to be executed after the evaluation operation obtains the state sensor signals. Based on the timing of the execution of the first step, which is any one of the learning action, the initialization operation, and the evaluation operation, the timing of the execution of the second step, which is any one of the learning action, the initialization operation, and the evaluation operation, is determined.
Citation Information
Patent Citations
Machine learning device and method for optimizing smoothness of feeding of feed shaft of machine and motor control device having machine learning device
JP2017102613A
Control parameter adjustment device, control parameter adjustment method, and control parameter adjustment program
JP2017102619A
Packaging tact, component mounter for reducing power consumption, and machine learning device
JP2017033979A
Control device, recording medium, and control system
US20180374001A1