Motor Control Map Adaptation for Noise-Resistant Learning Convergence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing adaptation systems face challenges in converging learning due to fluctuations in state variables caused by noise in sensor signals, making it difficult to optimize motor control functions effectively.
Innovation Solution
An adaptation system and method that includes a learning routine with first and second trials to adjust command values in a sign-reversing direction, followed by recording changes in a storage device without reflection in the function, and finally using summary statistics to update the function based on stored changes, thereby optimizing motor control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If learning is continuously updated based on state variables during motor driving, then the function optimization progresses, but noise in sensor signals causes fluctuations in state variables that prevent learning convergence
Solution Approach 1:
The system performs preliminary actions by storing multiple candidate changes in the storage device before finalizing the function update. During the learning routine, changes are accumulated and evaluated, with the best change selected and applied only after thorough evaluation, preventing premature convergence due to noisy state variables
Solution Approach 2:
The storage device acts as an intermediary between the learning routine and the function. Instead of directly updating the function with each learned change, the system stores multiple candidate changes and selects the optimal one, mediating the impact of noisy state variables on function optimization
Data Source
AI summary
Processing circuitry of an adaptation system executes a second process when a number of times of execution of a learning routine is greater than or equal to a specified number of times and less than a termination number of times. The second process performs a first trial and a second trial in each execution of the learning routine, and ends the learning routine by recording, in a storage device, a change in one of the first trial and the second trial in which a reward is larger. The processing circuitry executes a third process when the number of times of execution of the learning routine reaches a specified number of times. The third process is a process of calculating summary statistics of multiple changes and reflecting the summary statistics in a control map to complete optimization of the control map.


