Residual Feedback Learning for Control Arrangement Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional controllers face challenges in complex tasks such as robot assembly due to the perception of correction terms as disturbances, leading to counteractive behavior and reduced safety.
Innovation Solution
A method involving reinforcement learning to adapt feedback information for the regulation device, using residual feedback learning (RFL) to correct the feedback signals rather than directly correcting the control actions, thereby enhancing the control strategy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the actuator corrects the control actions output by the regulation device, then the control performance is improved, but the regulation device perceives the corrections as disturbances and counteracts them
Solution Approach 1:
The patent introduces feedback information as an intermediary between the actuator and the regulation device. Instead of directly correcting control actions, the actuator corrects the feedback information that the regulation device uses to generate control actions. This indirect approach allows the regulation device to operate without perceiving disturbances, while still achieving improved control performance through the corrected feedback information.
Solution Approach 2:
The patent inverts the conventional correction approach. Instead of the actuator correcting the regulation device's output (control actions), the actuator corrects the regulation device's input (feedback information). This inversion resolves the contradiction by eliminating the perception of disturbance while maintaining the corrective function.
2Adaptability or versatility
If reinforcement learning is used to train the actuator for generating correction terms, then the adaptability is improved, but the system complexity increases
Solution Approach 1:
The patent employs reinforcement learning with a reward mechanism that provides feedback to the actuator. The reward is generated based on whether the controlled system successfully fulfills the task (e.g., pin insertion success) and cost functions (e.g., energy consumption). This feedback loop enables the actuator to learn optimal correction strategies through trial and error, improving adaptability while keeping the training system manageable through clear reward signals.
Data Source
AI summary
A method for training a control arrangement for a controlled system. The control arrangement includes a regulation device and an actuator that operates according to a control strategy. The method includes the generation of control actions by the regulation device, each control action being generated by detecting measured variables that indicate a state of the controlled system, ascertaining a correction term for the detected measured variables by the actuator according to the control strategy, adapting the detected measured variables using the correction term for the detected measured variables, and generating the control action by supplying the adapted measured variables to the regulation device as the actual value. The method further includes training the control strategy by reinforcement learning for maximizing the gain that is achieved by the generated control actions.


