Learning control device, learning control method, and learning control program
The learning control device addresses the issue of reduced control performance due to deviations in the learning control start time by calculating and correcting for these deviations, ensuring consistent performance across trials.
Patent Information
- Application Number
- JP2021205941
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-20
AI Technical Summary
In learning control systems, the deviation in the learning control start time due to discrete operation timing leads to a reduction in control performance, as the same learning controller is used across trials without considering this deviation.
A learning control device that includes an update unit, a calculation unit, and a correction unit. The calculation unit calculates the deviation of the learning control start time based on the controlled object's state, and the correction unit corrects the updated correction control input to cancel out this deviation, ensuring consistent performance across learning trials.
The proposed solution effectively suppresses the reduction in control performance caused by deviations in the learning control start time, maintaining consistent control performance across multiple learning trials.
Smart Images

Figure 0007693530000008 
Figure 0007693530000009 
Figure 0007693530000010
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to a learning control device, a learning control method, and a learning control program.
Background Art
[0002] As a digital control device, a learning control device is known that repeatedly controls a control target according to a correction control input that is a learning value stored in a memory, and sequentially updates the learning value of the memory from the tracking error between the target value and the output value of the control target, improving the control performance each time it repeats.
[0003] As learning control devices, for example, learning control using look-ahead and learning control using a zero-phase filter at the time of updating the memory are disclosed (see, for example, Patent Document 1 and Patent Document 2).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0005] In learning control, it is common to continuously perform learning control from the start to the end of the operation of the controlled object. However, as the operation time becomes longer, it is necessary to prepare a larger-capacity memory. Therefore, it is conceivable to reduce the amount of memory used by starting the learning control when it is determined that the state of the controlled object satisfies the learning start condition and starting the learning control from the middle of the operation of the controlled object. However, in the prior art, the same learning controller is always used, and since the operation timing of the learning control device that performs digital control is discrete, a deviation occurs in the learning control start time between learning trials. For this reason, in the prior art, a reduction in control performance due to the deviation of the learning control start time may occur.
[0006] The problem to be solved by the present invention is to provide a learning control device, a learning control method, and a learning control program that can suppress a reduction in control performance due to a deviation in the learning control start time.
Means for Solving the Problem
[0007] The learning control device according to the embodiment includes an update unit, a calculation unit, and a correction unit. The update unit updates the correction control input used at the time of a learning trial according to the tracking error. The calculation unit calculates the deviation of the learning control start time according to the state of the controlled object at the start of learning control. The correction unit corrects the updated correction control input using the deviation so that the value cancels out the deviation. The time when the state of the controlled object satisfies the learning start condition and when the learning control actually starts learning control start time and of the deviation.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7A
Figure 7B
Figure 8
Figure 9A
Figure 9B
Figure 10
Embodiments for Carrying Out the Invention
[0009] With reference to the attached drawings below, the learning control device, learning control method, and learning control program according to this embodiment will be described in detail.
[0010] FIG. 1 is a schematic diagram of an example of the learning control device 10 according to this embodiment.
[0011] The learning control device 10 is a digital control device that repeatedly controls the control target 50, sequentially updates the learning value, and performs learning control to improve the control performance every time it repeats.
[0012] The learning control device 10 performs a state control trial at a predetermined sampling period that is a fixed interval. By repeating the state control trial at each sampling period by the learning control device 10, a learning trial, which is one learning control, is completed. Therefore, one learning trial includes a plurality of state control trials. One repetition of the above-described repetitive control corresponds to one learning trial.
[0013] The controlled object 50 is the object to be controlled by the learning control device 10. The controlled object 50 is the object of state control by the learning control device 10, such as a disk head drive of an HDD (Hard Disk Drive), a semiconductor manufacturing device, and a robot. The state of the controlled object 50 is, for example, the position on the disk or the position of the robot. Note that the state of the controlled object 50 is not limited to the position. For example, the state of the controlled object 50 may be a position, a speed, an acceleration, and a combination of two or more of these. In the present embodiment, the state of the controlled object 50 will be described by taking the form representing the position of the controlled object 50 as an example.
[0014] The learning control device 10 includes a learning control unit 20, a calculation unit 22, a feedback control unit 24, a first addition unit 26, an error calculation unit 28, and the controlled object 50.
[0015] The controlled object 50 operates in response to an input control signal sequentially received from the first addition unit 26 for each state control trial, and sequentially outputs a control amount y representing the state that is the operation result. As described above, in the present embodiment, the state of the controlled object 50 will be described by taking the form representing the position of the controlled object 50 as an example. For this reason, in the present embodiment, the controlled object 50 sequentially outputs a control amount y representing the position of the controlled object 50 as a result of operating in response to the received input control signal. Note that the control amount y of the controlled object 50 may be configured to be detected by a detection device such as a known sensor provided outside the controlled object 50.
[0016] The error calculation unit 28 calculates the error between the control amount y of the controlled object 50 and the target value r of the controlled object 50, and outputs it to the learning control unit 20 and the feedback control unit 24. The error calculation unit 28 sequentially receives the control amount y output from the controlled object 50 for each state control trial, calculates the error with the target value r each time the control amount y is received, and outputs it to the learning control unit 20 and the feedback control unit 24.
[0017] The feedback control unit 24 generates a feedback signal for causing the state of the controlled object 50 to follow the target value r using the error received from the error calculation unit 28, and outputs it to the first addition unit 26.
[0018] The first adder 26 outputs an input control signal obtained by adding the feedback signal received from the feedback control unit 24 and the correction control input received from the learning control unit 20 to the control target 50.
[0019] The correction control input is a learned value learned by the learning control unit 20 for each state control trial.
[0020] The learning control unit 20 includes an update unit 30 and a correction unit 40.
[0021] The update unit 30 updates the correction control input according to the tracking error. Specifically, the update unit 30 updates the correction control input to be used in the next learning trial according to the tracking error observed in the current learning trial.
[0022] In this embodiment, "current" and "next" represent one and the other of two consecutive learning trials in time series.
[0023] In this embodiment, "the current learning trial" means the latest learning trial, and "the next learning trial" means the next learning trial after the current one for the purpose of explanation.
[0024] In this embodiment, the update unit 30 includes a memory 32, a gain multiplication unit 34, and a third adder 36.
[0025] The memory 32 is a memory for storing the correction control input every sampling step i. The sampling step i represents the step of the state control trial for each sampling period by the learning control device 10. The correction control input at the sampling step i stored in the memory 32 is a learned value updated by the operation of the control target 50 until the previous learning trial.
[0026] The gain multiplication unit 34 multiplies the tracking error observed during the current learning trial by the gain g. The tracking error represents the error between the current state and the target state of the control target 50. In the present embodiment, the gain multiplication unit 34 uses the error between the target value r received from the error calculation unit 28 and the control amount y as the tracking error. Note that the gain multiplication unit 34 is not limited to the form of receiving the tracking error from the error calculation unit 28. For example, the gain multiplication unit 34 may obtain the tracking error from other functional units or the like mounted on the learning control device 10 and use it for multiplication by the gain g.
[0027] The third addition unit 36 stores, in the memory 32, as the corrected control input for the sampling step i, the addition result obtained by adding the multiplication result of multiplying the tracking error observed during the current learning trial by the gain g and the corrected control input for the sampling step i stored in the memory 32. Therefore, the corrected control input for the sampling step i stored in the memory 32 is sequentially updated for each learning trial according to the newly observed tracking error.
[0028] Here, in learning control, it is common to continuously perform learning control from the start to the end of the operation of the control target 50.
[0029] On the other hand, the learning control device 10 of the present embodiment starts learning control when it determines that the state of the control target 50 satisfies the learning start condition for each learning trial. That is, the learning control device 10 of the present embodiment starts learning control from the middle of the operation, which is the timing in the middle from the start to the end of the operation of the control target 50. By starting learning control from the middle of the operation of the control target 50, the learning control device 10 of the present embodiment can reduce the usage amount of the memory 32.
[0030] FIG. 2A, FIG. 2B, and FIG. 2C are explanatory diagrams of an example of learning control.
[0031] In FIGS. 2A and 2B, the horizontal axis represents time and the vertical axis represents position. The time indicated by the horizontal axis in FIG. 2A is the elapsed time since the start of the operation of the controlled object 50. The position indicated by the vertical axis is an example of a state that is the operation result of the controlled object 50. FIG. 2A shows, as an example, a diagram 60 (diagram 60a, diagram 60b, diagram 60c) showing the relationship between time and position due to the repetition of the state control trial in each of the three learning trials.
[0032] The learning control is performed by repeating the state control trial every sampling period T. For this reason, as shown in FIG. 2A, in the learning control device 10, the time ts which is the first sampling timing after the learning start condition is satisfied becomes the learning control start time. That is, there is a deviation between the time when the learning start condition is satisfied and the learning control start time when the learning control is actually started.
[0033] FIG. 2B shows the conversion of each of the plurality of learning trials represented by the diagram 60 (diagram 60a, diagram 60b, diagram 60c) shown in FIG. 2A into a diagram 62 showing the transition of the position of the controlled object 50 when it is assumed that the learning start condition is satisfied at the same time. FIG. 2B shows a plot Pc representing the learning control start time represented by the diagram 60a, a plot Pb representing the learning control start time represented by the diagram 60b, and a plot Pa representing the learning control start time represented by the diagram 60c.
[0034] As shown in FIG. 2B, in the operation of the controlled object 50 due to the repetition of the state control trial in each of the plurality of learning trials, when it is assumed that the learning start condition is satisfied at the same time, there is a deviation in the learning control start time between the plurality of learning trials. That is, due to the variation in the entire operation until the learning start condition is satisfied in each of the plurality of learning trials, the learning control cannot be started at the same time every time, and the learning control start time deviates from the time when the learning start condition is satisfied by a maximum width of one sampling period T.
[0035] However, in the prior art, the same learning controller was used in each learning trial, and the same control signal that did not consider the above deviation was output to the control target 50 between learning trials.
[0036] FIG. 2C is an explanatory diagram of an example of conventional learning control. In FIG. 2C, the diagram 70 represents the transition of the control signal output to the control target 50 for learning control. In FIG. 2C, the diagram 72 represents the transition of the state of the operation of the control target 50 by the control according to the control signal represented by the diagram 70. In FIG. 2C, the transitions of the states of the control target 50 in each of the three learning trials are shown as diagrams 72a, 72b, and 72c.
[0037] As shown in FIG. 2C, in the prior art, for the diagram 70 which is the transition of one type of control signal, a plurality of different types of state transitions were obtained between learning trials. That is, in the prior art, a deviation also occurred between the operation of the control target 50 and the output of the learning control, and the effect of the learning control was reduced. That is, in the prior art, a reduction in control performance may occur due to the deviation of the learning control start time.
[0038] Returning to FIG. 1 and continuing the explanation. Therefore, the learning control device 10 of the present embodiment includes a calculation unit 22 and a correction unit 40.
[0039] The calculation unit 22 calculates the deviation of the learning control start time, which is the time at the start of learning control, according to the state of the control target 50 at the start of learning control.
[0040] The calculation unit 22 acquires the control amount y of the control target 50 at the start of learning control as the state x0 of the control target 50 at the start of learning control. As described above, since the first sampling timing after satisfying the learning start condition is the learning control start time, the state x0 of the control target 50 at the learning control start time does not match the learning control start condition.
[0041] The calculation unit 22 calculates the deviation Δt0 of the learning control start time according to the acquired state x0.
[0042] The deviation Δt0 of the learning control start time represents the deviation of the learning control start time between a plurality of learning trials. Further, the deviation Δt0 of the learning control start time may represent the deviation between the time when the learning start condition is satisfied and the learning control start time.
[0043] As described above, the learning control start time is deviated by a maximum width of one sampling period T from the time when the learning start condition is satisfied. Therefore, in the present embodiment, the calculation unit 22 calculates the deviation between the learning control start time and the reference timing within the period of the sampling period T including the learning control start time as the deviation Δt0 of the learning control start time. Any timing within the period of the sampling period T may be predetermined as the reference timing. For example, the reference timing may be the central timing within the period of the sampling period T. In the present embodiment, the case where the reference timing is the central timing within the period of the sampling period T will be described as an example.
[0044] FIG. 3 is an explanatory diagram of an example of the calculation of the deviation Δt0 of the learning control start time. The calculation unit 22 calculates the deviation Δt0 of the learning control start time by the following formula (1) using the state x0 of the control target 50 at the start of learning control.
[0045]
Equation
[0046] In formula (1), Δt0 is the deviation Δt0 of the learning control start time. T is the sampling period T. x max and x min are parameters. x max is the maximum value of the state x0 when learning control is started so that the deviation Δt0 is in the range of -T / 2 or more and T / 2 or less in a certain learning trial. x min is the minimum value of the state x0 when learning control is started so that the deviation Δt0 is in the range of -T / 2 or more and T / 2 or less.
[0047] In this case, the state x0 at the start of learning control is x maxIn this case, the calculated deviation Δt0 becomes the maximum value T / 2. When the state x0 at the start of learning control is x min In this case, the calculated deviation Δt0 becomes the minimum value -T / 2.
[0048] Returning to FIG. 1, the description will be continued. The calculation unit 22 outputs the calculated deviation Δt0 of the learning control start time to the correction unit 40. The correction unit 40 stores the deviation Δt0 of the learning control start time received from the calculation unit 22. Note that the calculation unit 22 calculates the deviation Δt0 of the learning control start time from the state x0 of the control target 50 at the start of learning control for each learning trial, and outputs it to the correction unit 40. Each time the correction unit 40 receives a new deviation Δt0 of the learning control start time from the calculation unit 22, the stored deviation Δt0 of the learning control start time is updated to the newly received deviation Δt0 of the learning control start time. For this reason, the calculation unit 22 stores the newly calculated deviation Δt0 of the learning control start time used in each learning trial.
[0049] The correction unit 40 corrects the corrected control input updated by the update unit 30 using the deviation Δt0 so that it becomes a value that cancels out the deviation Δt0. In other words, the correction unit 40 corrects the corrected control input to be used at the next learning trial updated by the update unit 30 using the deviation Δt0 received from the calculation unit 22.
[0050] FIG. 4 is a schematic diagram showing an example of the configuration of the correction unit 40.
[0051] The correction unit 40 includes an HPF (high-pass filter) 40A, an LPF (low-pass filter) 40B, a linear interpolation unit 40C, and a second addition unit 40F.
[0052] The HPF 40A and the LPF 40B are filters for dividing the updated corrected control input into a high-frequency component and a low-frequency component. In other words, the HPF 40A and the LPF 40B are filters for dividing the corrected control input at the updated sampling step i into a high-frequency component and a low-frequency component.
[0053] HPF40A extracts the high-frequency components included in the corrected control input updated by the update unit 30 and outputs them to the second adder 40F.
[0054] LPF40B extracts the low-frequency components included in the corrected control input updated by the update unit 30 and outputs them to the linear interpolation unit 40C.
[0055] The linear interpolation unit 40C is a filter that performs linear interpolation to shift the output of the learning control.
[0056] Returning to FIG. 1, the description continues. The output of the learning control means the signal output from the learning control unit 20 to the controlled object 50 via the first adder 26. Therefore, the output of the learning control means at least one of the corrected control input output from the learning control unit 20 to the first adder 26 and the input control signal output from the first adder 26 to the controlled object 50.
[0057] Returning to FIG. 4, the description continues. The linear interpolation unit 40C is a filter for correcting the value of the corrected control input, which is the output of the learning control, to a value that cancels out the deviation Δt0 at the learning start time in accordance with the deviation of the operation of the controlled object 50 due to the deviation Δt0 of the learning control start time.
[0058] The linear interpolation unit 40C includes a first linear interpolation unit 40D and a second linear interpolation unit 40E.
[0059] The first linear interpolation unit 40D is a filter that performs linear interpolation when the deviation Δt0 is a positive value (+ value). The first linear interpolation unit 40D is a filter that performs linear interpolation on the corrected control input, which is the learning value at the current sampling step i stored in the memory 32, and the corrected control input, which is the learning value at the previous sampling step i.
[0060] The second linear interpolation unit 40E is a filter that performs linear interpolation when the deviation Δt0 is a negative value (- value). The second linear interpolation unit 40E is a filter that performs linear interpolation on the corrected control input that is the learning value of the current sampling step i stored in the memory 32 and the corrected control input that is the learning value of the next sampling step i + 1.
[0061] The filters used when each of the first linear interpolation unit 40D and the second linear interpolation unit 40E performs linear interpolation are 1 when the deviation Δt0 → 0, and (1 + z -1 ) / 2, (1 + z) / 2 when the deviation Δt0 → T / 2, -T / 2. T is the sampling period T. z is a variable in the Z-transform.
[0062] The filter used by the first linear interpolation unit 40D for linear interpolation is represented by Equation (2). Also, the filter used by the second linear interpolation unit 40E for linear interpolation is represented by Equation (3).
[0063]
Equation
[0064]
Equation
[0065] Figure 5 is a diagram showing the relationship of the corrected control input before and after linear interpolation by the linear interpolation unit 40C.
[0066] In Figure 5, the horizontal axis represents time, and the vertical axis represents the value of the corrected control input. In Figure 5, the plot Pa represents the plot of the corrected control input for each time before linear interpolation. The plot Pb represents the plot of the corrected control input for each time after linear interpolation. Figure 5 shows the relationship before and after linear interpolation when the deviation Δt0 at the start time of learning control is -T / 2 of the sampling period T.
[0067] As shown in FIG. 5, by linear interpolation by the linear interpolation unit 40C, the value of the corrected control input at the sampling step i updated by the update unit 30 is corrected and sequentially output to the control target 50 via the first addition unit 26. Therefore, by the linear interpolation by the linear interpolation unit 40C, the entire output of the learning control is pseudo-shifted by 1 / 2 step compared to the case where the entire output of the learning control is not linearly interpolated. That is, by the linear interpolation by the linear interpolation unit 40C, the value of the corrected control input output from the learning control unit 20 toward the first addition unit 26 for each state control trial is corrected to be the value at the timing that cancels the deviation Δt0 of the learning control start time.
[0068] However, if linear interpolation by the linear interpolation unit 40C is performed on all frequency components of the corrected control input received from the update unit 30, the gain of the high-frequency components will decrease. Therefore, the correction unit 40 includes an HPF 40A and an LPF 40B, and divides the corrected control input at the sampling step i updated by the update unit 30 into high-frequency components and low-frequency components. Then, the linear interpolation unit 40C of the correction unit 40 selectively performs linear interpolation on the output from the LPF 40B, which is a low-frequency component, and outputs it to the second addition unit 40F.
[0069] Therefore, the correction unit 40 of the present embodiment can suppress a decrease in the gain of the high-frequency components included in the corrected control input, and output the corrected control input corrected to be the value at the timing that cancels the deviation Δt0 of the learning control start time to the first addition unit 26.
[0070] Returning to FIG. 1, the description will be continued. The first addition unit 26 outputs an input control signal obtained by adding the corrected corrected control input received from the correction unit 40 and the feedback signal received from the feedback control unit 24 to the control target 50.
[0071] Therefore, an input control signal in which the deviation Δt0 of the learning control start time is canceled is input to the control target 50. Thus, in the learning control device 10 of the present embodiment, it is possible to suppress a reduction in the control performance of the control target 50 due to the deviation Δt0 of the learning control start time.
[0072] Next, an example of the information processing flow executed by the learning control device 10 of the present embodiment will be described.
[0073] FIG. 6 is a flowchart showing an example of the information processing flow executed by the learning control device 10 of the present embodiment.
[0074] When the operation of the control target 50 is started, the calculation unit 22 determines whether the state of the control target 50 satisfies the learning start condition (step S100). The calculation unit 22 repeats the negative determination (step S100: No) until an affirmative determination (step S100: Yes) is made in step S100. When an affirmative determination (step S100: Yes) is made in step S100, the process proceeds to step S102.
[0075] In step S101, the learning control unit 20 starts learning control (step S102).
[0076] The calculation unit 22 acquires the state x0 of the control target 50 at the start of learning control, which is the time when learning control is started in step S102 (step S104). Then, the calculation unit 22 calculates the deviation Δt0 of the start time of learning control according to the state x0 acquired in step S104 (step S106).
[0077] The correction unit 40 corrects the corrected control input of the sampling step i updated by the update unit 30 when learning control is started, using the deviation Δt0 calculated in step S106 (step S108).
[0078] The first addition unit 26 outputs an input control signal obtained by adding the corrected control input corrected in step S108 and the feedback signal received from the feedback control unit 24 to the control target 50 (step S112).
[0079] The learning control unit 20 determines whether to end the learning control (step S114). The learning control unit 20 makes the determination in step S114 by determining whether a predetermined learning control end condition is satisfied. If a negative determination is made in step S114 (step S114: No), the process returns to step S108. If an affirmative determination is made in step S114 (step S114: Yes), this routine ends.
[0080] As described above, the learning control device 10 of the present embodiment includes an update unit 30, a calculation unit 22, and a correction unit 40. The update unit 30 updates the correction control input used during a learning trial according to the tracking error. The calculation unit 22 calculates a deviation Δt0 of the learning control start time, which is the time at the start of learning control, according to the state of the control target 50 at the start of learning control. The correction unit 40 corrects the updated correction control input using the deviation Δt0 so that the value becomes one that cancels out the deviation Δt0.
[0081] In the present embodiment, the correction unit 40 corrects the correction control input updated by the update unit 30 so that the value cancels out the deviation Δt0 of the learning control start time. For this reason, an input control signal corresponding to the correction control input in which the deviation Δt0 of the learning control start time is canceled out is input to the control target 50.
[0082] Therefore, in the learning control device 10 of the present embodiment, it is possible to suppress a reduction in control performance due to the deviation Δt0 of the learning control start time.
[0083] Figures 7A to 8 are explanatory diagrams of the effects of the learning control device 10 of the present embodiment. In the description of Figures 7A to 8, a learning control device having the same configuration as the learning control device 10 of the present embodiment shown in Figure 1 was used as a comparative learning control device, which is a conventional learning control device, except that it does not include the correction unit 40 and the calculation unit 22.
[0084] Also, in the description of Figures 7A to 8, the filter of the HPF40A of the learning control device 10 of the present embodiment was set as the filter shown in the following formula (4), and the filter of the LPF40B was set as the filter shown in the following formula (5).
[0085]
Number
[0086]
Number
[0087] The simulation results of the difference between the target position and the actual position of the control target 50 when the control target 50 moves toward the target position are shown in FIGS. 7A and 7B. FIG. 7A shows the simulation results when using the comparative learning device. FIG. 7B shows the simulation results when using the learning control device 10 of the present embodiment. Note that FIGS. 7A and 7B are the results after sufficient learning, and the results of learning trials consisting of multiple state control trials are overwritten and shown.
[0088] In FIG. 7A, which is the simulation result of the comparative learning device that is the prior art, the entire waveform varies due to the deviation in the timing of starting learning control. On the other hand, in FIG. 7B, which is the simulation result of the learning control device 10 of the present embodiment, it was confirmed that such variation is suppressed and the reduction in control performance due to the deviation in the learning control start time is suppressed.
[0089] FIG. 8 is a diagram showing the output change from the correction unit 40 in the learning control device 10 of the present embodiment. FIG. 8 shows the cases where the deviation Δt0 of the learning control start time is approximately 0 and -T / 2. Therefore, it can be confirmed that in the learning control device 10 of the present embodiment, the output of the learning control is also corrected in accordance with the deviation Δt0 of the learning control start time.
[0090] FIGS. 9A and 9B are explanatory diagrams of the effects of the learning control device 10 of the present embodiment by actual machine experiments. In the description of FIGS. 9A to 9B, a learning control device having the same configuration as the learning control device 10 of the present embodiment shown in FIG. 1 was used as the comparative learning control device, which is the conventional learning control device, except that it does not include the correction unit 40 and the calculation unit 22.
[0091] Also, in the description of FIGS. 9A to 9B, the filter of the HPF 40A of the learning control device 10 of the present embodiment is the filter shown in the following formula (6), and the filter of the LPF 40B is the filter shown in the following formula (7).
[0092]
Equation
[0093]
Equation
[0094] The actual machine experiment results of the difference between the target position and the position of the actual control target 50 when the control target 50 operates toward the target position are shown in FIGS. 9A and 9B. FIG. 9A shows the actual machine experiment results when using the comparative learning device. FIG. 9B shows the actual machine experiment results when using the learning control device 10 of the present embodiment. Note that FIGS. 9A and 9B are the results after sufficient learning, and the results of learning trials consisting of multiple state control trials are overwritten and shown.
[0095] Similar to FIGS. 7A and 7B which are the results of the simulation, in FIG. 9A which is the actual machine experiment results of the prior art comparative learning device, the entire waveform varies due to the deviation in the timing of starting the learning control. On the other hand, in FIG. 9B which is the actual machine experiment results of the learning control device 10 of the present embodiment, it was confirmed that such variation is suppressed and the reduction in control performance due to the deviation in the learning control start time is suppressed.
[0096] From the above simulation results and actual machine experiment results, it was confirmed that the learning control device 10 of the present embodiment suppresses the reduction in control performance due to the deviation Δt0 of the learning control start time.
[0097] Next, an example of the hardware configuration of the learning control device 10 of the present embodiment will be described.
[0098] FIG. 10 is a hardware configuration diagram of an example of the learning control device 10 of the present embodiment.
[0099] The learning control device 10 of the present embodiment includes a control device such as a CPU (Central Processing Unit) 90B, a storage device such as a ROM (Read Only Memory) 90C, a RAM (Random Access Memory) 90D, and an HDD (hard disk drive) 90E, an I / F unit 90A that is an interface with various devices, and a bus 90F that connects each unit, and has a hardware configuration using a normal computer.
[0100] In the learning control device 10 of the present embodiment, the CPU 90B reads a program from the ROM 90C and executes it on the RAM 90D, whereby each of the above units is realized on a computer.
[0101] Note that the program for executing each of the above processes executed by the learning control device 10 of the present embodiment may be stored in the HDD 90E. Further, the program for executing each of the above processes executed by the learning control device 10 of the present embodiment may be provided by being pre-installed in the ROM 90C.
[0102] Also, the program for executing the above processes executed by the learning control device 10 of the present embodiment may be stored in a computer-readable storage medium such as a CD-ROM, a CD-R, a memory card, a DVD (Digital Versatile Disc), or a flexible disk (FD) in an installable or executable file format and provided as a computer program product. Further, the program for executing the above processes executed by the learning control device 10 of the present embodiment may be stored on a computer connected to a network such as the Internet and provided by being downloaded via the network. Further, the program for executing the above processes executed by the learning control device 10 of the present embodiment may be provided or distributed via a network such as the Internet.
[0103] Note that the embodiments of the present invention have been described above, but the above embodiments are presented as examples and are not intended to limit the scope of the invention. This novel embodiment can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. This embodiment and its modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and its equivalent scope.
Explanation of Reference Numerals
[0104] 10 Learning control device 22 Calculation unit 26 First addition unit 30 Update unit 40 Correction unit 40A HPF 40B LPF 40C Linear interpolation unit 40F Second addition unit
Claims
1. An update unit that updates a correction control input used during a learning trial according to a tracking error; A calculation unit that calculates a deviation between a time when the state of the control target satisfies a learning start condition and a learning control start time when learning control is actually started, according to the state of the control target at the start of learning control; A correction unit that corrects the updated correction control input using the deviation so that the value cancels out the deviation; A learning control device comprising the above.
2. The correction unit includes: A low-pass filter that extracts a low-frequency component included in the updated correction control input; A high-pass filter that extracts a high-frequency component included in the updated correction control input; A linear interpolation unit that linearly interpolates the output of the low-pass filter using the deviation; A second addition unit that outputs the addition result of the output of the high-pass filter and the output of the linear interpolation unit as the corrected correction control input; The learning control device according to claim 1, comprising the above.
3. A first addition unit that outputs an input control signal obtained by adding a feedback signal for causing the state of the control target to follow a target value and the corrected correction control input to the control target; Comprising the above, The learning control device according to claim 1 or claim 2.
4. A step of updating a correction control input used during a learning trial according to a tracking error; A step of calculating a deviation between a time when the state of the control target satisfies a learning start condition and a learning control start time when learning control is actually started, according to the state of the control target at the start of learning control; A step of correcting the updated correction control input using the deviation so that the value cancels out the deviation; A learning control method including the above.
5. A step of updating a corrected control input used during a learning trial according to a tracking error; A step of calculating a deviation between a time when the state of the controlled object satisfies a learning start condition and a learning control start time when learning control is actually started, according to the state of the controlled object at the start of learning control; A step of correcting the updated corrected control input using the deviation so as to be a value that cancels out the deviation; A learning control program for causing a computer to execute the above.
Citation Information
Patent Citations
Digital servo device using look-ahead learning
JP1997146645A
Disk drive device and head positioning controlling method for the device
JP2001126421A
Disk drive and tracking control method
JP2002093085A
Control device
JP2006118429A
Detection of and responses to time delays in networked control systems
US20170060102A1