Neural Oscillator Robot Gait Control for Faster RL Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for rhythmic motion control of robots, particularly quadruped robots, require extensive professional knowledge and are time-consuming due to the difficulty in designing a reward function for model-free reinforcement learning, which hinders unbiased learning.

Innovation Solution

A control structure composed of a neural oscillator and pattern formation network, combined with a designed action space for joint position increment, accelerates the training process of rhythmic motion reinforcement learning.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If model-free reinforcement learning is used for autonomous learning of motion strategy, then the robot can learn independently without extensive professional knowledge, but the reward function design becomes time-consuming and difficult

Engineering Contradiction:
Improveautonomous learning capabilityVSAvoidreward function design time
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The patent segments the control task into two parts: a neural oscillator that generates rhythmic motion patterns autonomously, and a reinforcement learning component that only needs to learn timing and coordination. This segmentation allows autonomous learning while reducing the complexity of reward function design, as the RL agent only needs to learn when to activate each leg rather than the entire motion strategy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The neural oscillator acts as an intermediary between the high-level motion goal and the low-level actuator control. It generates intermediate rhythmic patterns that guide the reinforcement learning process, making the learning task more manageable and reducing the time needed to design effective reward functions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional control methods with sensory feedback and complex control theory are used, then better motion performance can be obtained, but the design process requires rich professional knowledge and is time-consuming

Engineering Contradiction:
Improvemotion performanceVSAvoidcontrol theory complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical control theory with a biologically-inspired neural oscillator model. Instead of using sophisticated control algorithms requiring deep theoretical knowledge, the system uses a simplified neural network that naturally generates rhythmic patterns, achieving good motion performance with reduced complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The neural oscillator is self-regulating and automatically generates appropriate rhythmic patterns without requiring external intervention or complex feedback processing. The system serves itself by internally generating the timing and coordination signals needed for locomotion, reducing the need for external control complexity.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12372981B2Method and system for rhythmic motion control of robot based on neural oscillator
Publication Date: 2025.07.29 SHANDONG UNIV
  • US12372981B2 patent drawing

AI summary

A method and a system for rhythmic motion control of a robot based on a neural oscillator, including: acquiring a current state of the robot, and a phase and a frequency generated by the neural oscillator; and obtaining a control instruction according to the acquired current state, phase and frequency and a preset reinforcement learning network so as to control the robot. The preset reinforcement learning network includes an action space, a pattern formation network and the neural oscillator. A control structure designed by the present disclosure, which is composed of the neural oscillator and the pattern formation network, can ensure formation of an expected rhythmic motion behavior; and meanwhile, a designed action space for joint position increment can effectively accelerate the training process of rhythmic motion reinforcement learning, and solve a problem that design of the reward function is time-consuming and difficult in learning with existing model-free reinforcement learning.