Bipedal Action Model With Dual-Rate Control for Humanoid Balance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional robotic control systems are limited in their ability to manage the vast number of degrees of freedom in humanoid robots, resulting in rigid, imprecise, and unnatural movements, particularly in unstructured environments, due to their reliance on discrete action outputs and lack of dynamic balance control.
Innovation Solution
A bipedal action model (BAM) architecture with a decoupled dual-system design, comprising a high-level cognitive alpha model and a low-level reactive beta model, where the alpha model processes complex multimodal inputs and generates a task-conditioning latent vector, while the beta model translates this intent into precise, continuous robot actions, enabling whole-body control and fluid motion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional discrete action outputs are used to control humanoid robot, then device complexity is reduced, but manufacturing precision and movement smoothness deteriorate
Solution Approach 1:
The patent replaces the conventional discrete mechanical control system with a neural network-based continuous control system. The neural network generates continuous action outputs that directly control the robot's degrees of freedom, eliminating the need for discrete action binning and resulting in smoother, more precise movements without increasing overall system complexity.
Solution Approach 2:
The patent changes the control output from discrete parameter values to continuous parameter values. By using continuous action outputs from the neural network instead of discrete binned actions, the system achieves smoother transitions and more precise control over the robot's movements, directly addressing the precision issue while maintaining manageable complexity.
2Ease of operation
If discrete binned actions are generated for robot control, then ease of operation is improved, but productivity and task completion accuracy deteriorate
Solution Approach 1:
The patent substitutes the discrete action generation mechanism with a continuous action generation mechanism based on neural networks. This allows the robot to execute tasks more efficiently by generating smooth, continuous motion trajectories that reduce the compounding errors inherent in discrete action sequences, thereby improving task completion speed and accuracy while maintaining operational simplicity through learned policies.
3Device complexity
If conventional control systems are used for humanoid robot, then device complexity is reduced, but reliability and dynamic balance control deteriorate
Solution Approach 1:
The patent segments the control architecture into a neural network policy that directly outputs continuous actions for each degree of freedom. This segmentation allows independent optimization of each joint's control while maintaining overall system coordination, improving dynamic balance and reliability without requiring overly complex centralized control mechanisms.
Solution Approach 2:
The patent implements dynamic control by using continuous action outputs that can adapt in real-time to changing conditions. The neural network generates continuous control signals that smoothly adjust the robot's posture and movements, enabling reliable dynamic balance control that responds fluidly to environmental changes rather than relying on rigid pre-programmed motions.
4Loss of information
If discrete action sequences are used for robot control, then loss of information is reduced, but measurement precision and temporal consistency deteriorate
Solution Approach 1:
The patent changes the action representation from discrete to continuous parameters. This allows for more precise action outputs that better capture the nuanced movements required for accurate control, reducing information loss while maintaining better measurement precision through the continuous nature of the neural network outputs.
Data Source
AI summary
The present disclosure provides a system for generating motor control commands for a humanoid robot, comprising an alpha model with over 1 billion parameters that processes visual observations and language instructions at a first frequency to generate contextual embeddings, and a beta model operating at a higher second frequency. The beta model includes an embodiment-specific state encoder projecting robot state information into a shared embedding space, a diffusion transformer module generating denoised action sequences through iterative flow-matching that cross-attends to the alpha model's contextual embeddings, and an embodiment-specific action decoder converting denoised sequences into motor control commands. The beta model generates action chunks comprising future action sequences over a predetermined time horizon in a single inference step, with the complete system having less than 5 billion parameters.


