A humanoid robot full-body cooperative voice control system and method
Through a closed-loop control system that integrates multimodal intent understanding, full-body motion planning, and dynamic balance safety arbitration, the semantic gap and motion fragmentation problems in voice interaction and control of humanoid robots have been solved. This system enables the direct conversion of natural language commands into full-body coordinated motion, ensuring the safety and stability of task execution.
Patent Information
- Application Number
- CN202511620812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Existing voice interaction and control systems for humanoid robots suffer from semantic gaps, fragmented movements, and safety lags, making it impossible to effectively perform complex full-body coordinated movement tasks.
A closed-loop control system employs a multimodal intent understanding module, a full-body motion planning module, a dynamic balance safety arbitrator, and a microkernel safety scheduler. Combined with a semantic-action mapping knowledge base and an asymmetric load balancing algorithm, it achieves direct conversion of natural language commands into full-body coordinated movements. Stability is ensured through dynamic balance safety arbitration and multi-level safety protection mechanisms.
It enables robots to perform complex tasks safely, stably, and in a human-like manner, improves the naturalness of human-computer interaction and the robustness of the system, reduces the risk of instability, and enhances adaptability and intelligence.
Smart Images

Figure CN121075330B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, specifically to the intelligent control of humanoid robots, and in particular to a humanoid robot's full-body collaborative voice control system and method. Background Technology
[0002] Intelligent voice interaction has become an indispensable function for service robots, especially humanoid robots. However, existing humanoid robots have several interconnected fundamental bottlenecks in their voice interaction and control capabilities, making them unable to handle complex daily tasks.
[0003] First, at the interaction level, most existing systems can only map voice commands to a limited, pre-defined library of simple actions (such as "forward" and "shake hands"). When users issue complex commands requiring full-body coordination, such as "walk around the chair and steadily bring me that glass of water on the table," existing systems cannot deeply understand the complex requirements implied in the command, such as obstacle avoidance, two-handed coordination, fine manipulation ("bring"), and dynamic stability ("steadily"), resulting in a "semantic gap." The core reason for this is the lack of a mechanism that can systematically map the rich semantics of natural language (especially descriptive words) to the robot's high-dimensional, nonlinear motion control space.
[0004] Secondly, at the planning level, traditional solutions are often simply a combination of "voice-controlled robotic arms" or "voice-controlled mobile chassis," rather than planning the humanoid robot as a complete dynamic system for coordinated full-body movement, resulting in a "disjointed movement" problem. This leads to stiff and unnatural robot movements, failing to leverage the advantages of its humanoid form in completing tasks in complex environments.
[0005] Finally, at the control level, humanoid robots must maintain dynamic balance when performing complex actions. However, traditional balance control algorithms (such as feedback control based on zero-torque point (ZMP)) are usually decoupled from higher-level semantic understanding and motion planning modules. This can lead to planned actions that, while semantically consistent, are dynamically infeasible or unstable. The lower-level controller can only provide "post-hoc" remediation after instability has occurred or is about to occur, resulting in a "safety lag" problem and a high risk of loss of control.
[0006] In summary, the shortcomings of existing technologies do not exist in isolation, but rather stem from the fragmentation of the three core components: interaction, planning, and balanced control. This invention aims to solve this fundamental problem from the system architecture level. Summary of the Invention
[0007] The purpose of this invention is to overcome the shortcomings of the existing humanoid robot voice control, which separates interaction, planning and balance control, and to provide a humanoid robot full-body coordinated voice control system and method, so as to realize the direct and intelligent conversion from natural language commands to safe, stable and human-like full-body coordinated movement.
[0008] To achieve the above objectives, the specific solution adopted by the present invention is as follows:
[0009] On one hand, the present invention provides a humanoid robot full-body collaborative voice control system, integrated into the humanoid robot body, comprising:
[0010] The multimodal intent understanding module is used to receive and recognize the user's voice commands, obtain natural language commands, and combine them with the physical attributes of environmental objects and the user's relative pose information obtained by the visual sensor. By querying the semantic-action mapping knowledge base, the natural language commands are parsed into specific task objectives and their implicit whole-body motion constraints.
[0011] The full-body motion planning module, connected to the multimodal intent understanding module, is used to generate an initial joint trajectory sequence containing coordinated movements of the "torso-arm-hand-leg" based on the task objective and full-body motion constraints, estimate the range of dynamic centroid offset caused by executing the sequence, and use an asymmetric load balancing algorithm for motion compensation in tasks involving asymmetric loads on both arms.
[0012] The dynamic balance safety arbitrator, connected to the whole-body motion planning module, is used to perform feasibility verification and smoothing optimization of the initial joint trajectory sequence based on the current posture of the humanoid robot and the estimated dynamic offset range of the center of mass, using a preset stability boundary model, and outputs safe joint control commands.
[0013] The microkernel safety scheduler, connected to the dynamic balance safety arbiter, responds to the safety joint control command and distributes it to the corresponding joint driver at a fixed scheduling cycle. At the same time, it monitors the status of each joint and the overall balance in real time, and immediately executes the predefined safety protection mechanism when it detects actual center of mass shift or the plantar pressure center approaching the stability boundary.
[0014] The multimodal intent understanding module, the whole-body motion planning module, the dynamic balance safety arbitrator, and the microkernel safety scheduler are connected in sequence to form a closed-loop control system from voice interaction and whole-body motion planning to dynamic balance control.
[0015] Furthermore, the semantic-action mapping knowledge base is constructed by parsing massive amounts of human-computer interaction data, and it can map adverbs and adjectives in natural language into specific motion control parameters.
[0016] Furthermore, the core formula of the asymmetric load balancing algorithm is:
[0017]
[0018]
[0019]
[0020] Among them, △COM x This represents the offset of the center of mass in the lateral (X-axis) direction between the actual and desired center of mass positions caused by the load from both arms; m1 and m2 represent the masses of the loads held by the left and right hands, respectively; x1 and x2 represent the lateral position coordinates of the loads held by the left and right hands in the robot's torso coordinate system; COM x0 θ represents the desired position coordinates of the robot's center of mass in the horizontal (X-axis) direction when the robot is unloaded (i.e., without any load); spine K represents the virtual spinal curvature angle calculated to compensate for centroid shift; p K d These represent the proportional gain coefficient and the differential gain coefficient of the spinal curvature controller, respectively. Indicates the backswing compensation angle required by the unloaded (or lightly loaded) arm to provide balancing torque; K comp This represents the compensation coefficient for the arm swing motion. This represents the rate of change of the lateral offset of the centroid with respect to time.
[0021] Furthermore, the expression for the stability boundary model built into the dynamic equilibrium arbitrator is:
[0022]
[0023]
[0024] Where x represents the state vector of the robot system, x=[θ,ω,P] COM V COM ]ᵀ, where θ and ω represent the joint angle vector and joint angular velocity vector of all joints of the robot, respectively, P COM and V COM Let V(x) and Q(x) represent the position vector and velocity vector of the robot's center of mass in three-dimensional space, respectively; V(x) represents the Lyapunov function constructed based on Lyapunov stability theory, used to measure the total energy or generalized distance of the system; P represents a positive definite weight matrix used to configure the weights of the state vector x in the Lyapunov function; Q i θ represents the joint angle deviation weighting coefficient of the i-th joint, 1≤i≤n, where n represents the total number of joints; i θ represents the joint angle of the i-th joint. i-refThe reference angle of the i-th joint is represented by S; S represents the stability margin, where S > 0 indicates stability and S < 0 indicates a risk of instability. This represents the partial derivative of the Lyapunov function with respect to time.
[0025] Furthermore, the security protection mechanism triggered by the microkernel security scheduler includes three levels: Level 1 slows down the movement speed, Level 2 stops the gait and switches to a stationary balance mode, and Level 3 immediately locks all joints and activates mechanical brakes.
[0026] Furthermore, it also includes an interactive context learning module, which records the actual motion trajectory and user feedback after the task is executed, and fine-tunes the parameters of the semantic-action mapping knowledge base based on this to achieve adaptive optimization of system performance.
[0027] On the other hand, the present invention provides a method for control using the above-mentioned humanoid robot full-body collaborative voice control system, comprising the following steps:
[0028] Step S1, Intent Understanding: The multimodal intent understanding module receives and recognizes the user's voice commands to obtain natural language commands. Combined with environmental visual information, the natural language commands are parsed into task objectives containing whole-body motion constraints through the semantic-action mapping knowledge base.
[0029] Step S2, Motion Planning: Based on the task objective and whole-body motion constraints, the whole-body motion planning module generates an initial whole-body joint trajectory sequence and estimates the dynamic shift of the center of mass during the execution of the sequence.
[0030] Step S3, Safety Arbitration: Using a dynamic balance safety arbitrator, based on the dynamic offset of the centroid and the current state of the robot, the initial joint trajectory sequence is verified and optimized using a stability boundary model to generate safe joint control commands; wherein, the stability boundary model is used to calculate the stability margin of the robot;
[0031] Step S4, Scheduling and Execution: The safety joint control instructions are executed through the microkernel safety scheduler, and the robot's posture stability is monitored in real time. When the stability margin determination system approaches the stability boundary, the predefined safety protection mechanism is automatically executed.
[0032] Furthermore, in step S1, for instructions containing delivery semantics, the multimodal intent understanding module automatically loads a socially compliant motion template from the semantic-action mapping knowledge base, the template constraining the robot at least during the delivery process:
[0033] Maintain a smooth motion trajectory and stable speed for the end effector;
[0034] And after confirming that the item has been received, perform a small-scale withdrawal action.
[0035] Furthermore, after step S4, the following steps are also included:
[0036] Step S5, Learning Optimization: Through the interactive context learning module, the control parameters in the semantic-action mapping knowledge base are adaptively adjusted based on the actual motion trajectory and user feedback during task execution.
[0037] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0038] (1) This invention constructs a unified control architecture through the closed-loop collaboration of a multimodal intent understanding module, a full-body motion planning module, a dynamic balance safety arbitrator, and a microkernel safety scheduler. This architecture fundamentally solves the three core problems of "semantic gap," "motion fragmentation," and "safety lag" caused by the fragmentation of each link in the prior art. Specifically, the tight coupling between modules ensures a smooth transition from high-level semantics to low-level execution, enabling the robot to understand and execute complex task instructions that require full-body coordination, while ensuring the safety and reliability of the execution process.
[0039] (2) By constructing a semantic-action mapping knowledge base, this invention systematically maps adverbs (such as "carefully" and "quickly") and adjectives in natural language to specific robot motion control parameters for the first time. This makes the robot's actions no longer mechanical and preset, but can be adjusted according to the context and emotional tone of the instructions, possessing "style" and "emotional tone", significantly improving the naturalness of human-computer interaction and user experience.
[0040] (3) This invention moves stability verification from the traditional low-level feedback control to the high-level action planning layer through a dynamic balance safety arbitrator, achieving "pre-event" safety verification of actions. Combined with the real-time monitoring and multi-level circuit breaker mechanism of the microkernel safety scheduler during task execution, a dual safety barrier of "pre-event verification" and "in-event protection" is formed. This proactive protection mechanism greatly reduces the risk of instability of humanoid robots when performing tasks in complex and dynamic environments, and improves the overall robustness of the system.
[0041] (4) Through the microkernel secure scheduler, a low-level core designed specifically for real-time operation, this invention ensures that the complex intelligent decision-making algorithms at higher levels can be seamlessly integrated with the hard real-time control requirements of the underlying components. This design meets the stringent requirements for system response speed in dynamic environments, ensuring that intelligent planning can be executed in a timely and accurate manner.
[0042] (5) The asymmetric load balancing algorithm introduced in this invention calculates the centroid shift caused by unilateral load online using a formula, and dynamically adjusts the virtual spinal curvature and the swing of the non-loaded arm to compensate for it. This algorithm effectively overcomes the stability and balance challenges of humanoid robots when performing single-arm load-bearing or double-arm asymmetric load tasks (such as lifting or delivering items with one hand), and expands the application scope of robots in practical scenarios such as logistics and housekeeping.
[0043] (6) The stability boundary model based on Lyapunov stability adopted in this invention outputs a clear and quantifiable stability margin (S) through a calculation formula. This index provides a precise arbitration basis for the dynamic equilibrium safety arbitrator, transforming stability verification from traditional qualitative and empirical judgment to precise quantitative analysis. This enables earlier and more accurate identification of instability risks and triggers corresponding optimization or protection strategies, achieving refined and intelligent safety control.
[0044] (7) This invention establishes a continuous optimization mechanism for system performance by introducing an interactive context learning module. This module automatically records actual motion trajectory data and user feedback information after each task execution, and adaptively fine-tunes the parameters in the semantic-action mapping knowledge base based on this. This closed-loop learning mechanism enables the system to continuously optimize its behavior patterns from practical operational experience. It can not only gradually adapt to the personalized preferences and usage habits of specific users, but also adjust control strategies according to environmental characteristics, significantly improving the system's adaptability and intelligence level in different application scenarios, achieving a leap from single task execution to long-term performance evolution.
[0045] (8) The multi-level safety protection mechanism designed in this invention provides hierarchical and progressive safety assurance. Through the three-level protection strategy implemented by the microkernel safety scheduler, the system can take corresponding countermeasures according to the severity of the instability risk: from gentle adjustment to reduce speed, to attitude maintenance to stop movement, and then to emergency braking with complete locking. This hierarchical design ensures that the safety response is not overly sensitive and affects task execution, while providing the highest level of protection in truly dangerous situations, achieving the best balance between safety assurance and task efficiency, and providing full-cycle protection for the safe and reliable operation of the robot in unstructured environments. Attached Figure Description
[0046] Figure 1 This is a block diagram of the overall architecture of the system of the present invention.
[0047] Figure 2 This is a schematic diagram of the semantic-action mapping process in this invention.
[0048] Figure 3 This is a flowchart illustrating the collaborative workflow of whole-body motion planning and dynamic balance safety arbitration in this invention.
[0049] Figure 4 This is a sequence diagram of a specific application scenario of the present invention ("carefully picking up the ball under the table"). Detailed Implementation
[0050] The technical solution of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0051] This invention discloses a humanoid robot full-body collaborative voice control system, such as... Figure 1 As shown, this system adopts a layered architecture design, integrated into the humanoid robot body. It includes a multimodal intent understanding module, a full-body motion planning module, a dynamic balance safety arbitrator, and a microkernel safety scheduler connected sequentially. The multimodal intent understanding module, full-body motion planning module, and dynamic balance safety arbitrator run on the robot's main control computer (e.g., a host computer), responsible for high-level intelligent decision-making. The microkernel safety scheduler, on the other hand, is directly deployed on the processor with the highest real-time requirements (e.g., an embedded MCU or a real-time Linux kernel), communicating directly with the joint drivers via a real-time bus (e.g., EtherCAT) to ensure the hard real-time performance of the underlying control. This architecture achieves effective separation and seamless integration of intelligent decision-making and real-time control. The system architecture is described in detail below.
[0052] (1) Multimodal intent understanding module
[0053] The multimodal intent understanding module serves as the human-computer interaction entry point for the entire system. Its workflow is as follows: it receives voice signals from the microphone and converts them into natural language commands in text form through a speech recognition engine; simultaneously, it acquires the physical attributes of environmental objects (size, material, shape, etc.) and the user's relative pose information from visual sensors. It then performs deep parsing of the natural language commands by querying a semantic-action mapping knowledge base. This knowledge base is not a simple command-action lookup table, but a multi-level mapping network containing verbs, nouns, adverbs, adjectives, and even prepositional phrases, down to motion primitives and their parameterized constraints.
[0054] Semantic mapping examples are as follows (e.g.) Figure 2 (as shown)
[0055] For example, the instruction "Carefully lift this box and hand it to me":
[0056] The verb "to move" is mapped to the basic motion unit of bi-arm load.
[0057] The adverb "carefully" is mapped to the parameters: {"End velocity ratio": 0.6, "Torque detection gain": 1.5}
[0058] The noun "box" (visually identified as "rigid and of a certain weight") is mapped to {"grip strength": "medium", "grip pattern": "palm-wrapped"}.
[0059] The prepositional phrase "hand it to me" (in conjunction with the user's pose) is mapped to {"final pose": "facing the user", "delivery height": "user's hand height"}.
[0060] All mapping results are integrated into a structured task description, containing explicit task objectives and their implicit whole-body motion constraints, and then passed to the next module.
[0061] (2) Full-body motion planning module
[0062] This module receives a structured task description from the intent understanding module and is responsible for generating a whole-body coordinated motion trajectory: using inverse kinematics algorithms and trajectory optimization techniques, it generates a timestamped initial joint trajectory sequence containing coordinated movements of the torso, arms, hands, and legs; based on the robot's complete dynamic model, it performs forward dynamic simulation on the generated trajectory sequence to predict the range of dynamic displacement of the center of mass and changes in plantar reaction force caused by executing the sequence.
[0063] When the task involves asymmetrical loads on both arms, the asymmetrical load balancing algorithm is automatically activated. It calculates the center of mass offset using a formula, thereby determining the virtual spinal curvature angle and the backswing compensation angle of the unloaded arm. Specifically, this is achieved through the following formula: Center of mass offset calculation:
[0064]
[0065] Spinal compensation control:
[0066]
[0067] Arm compensation control:
[0068]
[0069] Among them, △COM x This represents the offset between the actual and desired center of mass positions in the lateral direction due to the load on both arms; m1 and m2 represent the masses of the loads held by the left and right hands, respectively; x1 and x2 represent the lateral coordinates of the loads held by the left and right hands in the robot's torso coordinate system; COM x0 θ represents the desired lateral position coordinates of the robot's center of mass when it is unloaded. spine K represents the virtual spinal curvature angle calculated to compensate for centroid shift; p Kd These represent the proportional gain coefficient and the differential gain coefficient of the spinal curvature controller, respectively. Indicates the backswing compensation angle required by the unloaded side arm to provide balancing torque; K comp This represents the compensation coefficient for the arm swing motion. This represents the rate of change of the lateral offset of the centroid with respect to time.
[0070] (3) Dynamic balance safety arbitrator
[0071] This module ensures the dynamic feasibility of the planned actions: based on the current robot posture (bipedal support area) and the estimated dynamic displacement of the center of mass, a stability boundary model is used to verify the feasibility.
[0072] Stability Verification: Based on the current robot posture (bipedal support area) and the estimated dynamic centroid offset, a feasibility verification is performed using a stability boundary model. This model is built upon Lyapunov stability theory.
[0073] Lyapunov function:
[0074]
[0075] Stability margin calculation:
[0076]
[0077] Where x represents the state vector of the robot system, x=[θ,ω,P] COM V COM ]ᵀ, where θ and ω represent the joint angle vector and joint angular velocity vector of all joints of the robot, respectively, P COM and V COM Let V(x) and Q(x) represent the position vector and velocity vector of the robot's center of mass in three-dimensional space, respectively; V(x) represents the Lyapunov function constructed based on Lyapunov stability theory, used to measure the total energy or generalized distance of the system; P represents a positive definite weight matrix used to configure the weights of the state vector x in the Lyapunov function; Q i θ represents the joint angle deviation weighting coefficient of the i-th joint; i θ represents the joint angle of the i-th joint. i-ref The reference angle of the i-th joint is represented by S; S represents the stability margin, where S > 0 indicates stability and S < 0 indicates a risk of instability. Let represent the partial derivative of the Lyapunov function with respect to time. Where min Take the minimum value within the next forecast time window.
[0078] In the specific calculation process, the system state vector includes information such as joint angles, angular velocities, center of mass position, and velocity. The stability margin is calculated using a positive definite weight matrix and joint angle deviation weighting coefficients. When the stability margin falls below the safety threshold, the planning module is required to replan, optimizing the trajectory by introducing virtual torso tilt and adjusting leg stiffness, until a safe joint control command that satisfies both task semantics and dynamic safety is output.
[0079] (4) Microkernel security scheduler
[0080] The microkernel safety scheduler serves as the system's "real-time execution and protection center." Connected to the dynamic balance safety arbitrator, it distributes safety joint control commands to the corresponding joint drivers at a fixed hard real-time scheduling cycle. Simultaneously, it monitors the real-time status of each joint (e.g., position, torque) and the overall balance (e.g., actual center of mass shift, plantar pressure center), and immediately triggers priority-based motion shutdown or degradation strategies when it detects the system approaching its stability boundary. Its role is to ensure seamless and secure integration between high-level intelligent planning and low-level hard real-time control, achieving full-cycle safety protection.
[0081] The microkernel safety scheduler implements a hierarchical safety protection mechanism with three distinct levels: Level 1 protection reduces system kinetic energy by decreasing the movement speed of each joint, serving as the most basic means of stabilization and recovery; Level 2 protection immediately stops the gait and switches to a stationary balance mode when a sustained risk of instability is detected, at which point the robot enters a high-stiffness posture maintenance state; Level 3 protection, as the final safety guarantee, immediately locks all joints and activates mechanical brakes when faced with an emergency instability threat, completely freezing the robot's motion state. This hierarchical design ensures that the safety response has both appropriate progressiveness and provides the highest level of protection.
[0082] (5) Interactive Context Learning Module
[0083] Preferably, the system of the present invention further includes an interactive context learning module, which continuously optimizes system performance by recording key parameters of the actual motion trajectory (including joint angle sequences, torque curves, and center of mass trajectories) and user satisfaction feedback provided through a preset interface after each task execution. Based on this data, the module uses an incremental learning algorithm to fine-tune the control parameters in the semantic-action mapping knowledge base, for example, adjusting the end-effector velocity ratio coefficient based on the actual execution effect of the "careful" command, or adjusting the upper limit of joint velocity based on user satisfaction with the "fast" command. This continuous learning mechanism enables the system to gradually adapt to the user's personal preferences and specific environmental characteristics, achieving truly personalized service.
[0084] Accordingly, the present invention also provides a control method employing the above-described system, such as... Figure 3 As shown, it includes the following steps:
[0085] Step S1: The intent understanding step is implemented through the multimodal intent understanding module, transforming ambiguous natural language instructions into precise, executable robot task descriptions. Specifically, this includes: speech recognition and text conversion, environmental information perception, semantic mapping query, and task objective and constraint extraction. Its output is a structured task description, laying the foundation for subsequent planning.
[0086] Step S2: The motion planning steps are implemented through the whole-body motion planning module, transforming the task description into specific, time-series-based whole-body joint motion trajectories. This includes: initial trajectory generation based on inverse kinematics, forward simulation based on a dynamic model, center-of-mass shift prediction, and asymmetric load balancing compensation. Its output is a sequence of whole-body joint trajectories containing timestamps and their predicted dynamic characteristics.
[0087] Step S3: Implement a safety arbitration process using a dynamic equilibrium safety arbitrator to verify and optimize the planned actions, ensuring their dynamic safety. This includes: stability boundary calculation, stability margin assessment, trajectory feasibility judgment, and iterative optimization adjustment. Its output is the joint control command that passes the safety verification.
[0088] Step S4: The microkernel secure scheduler implements the scheduling execution steps, which not only executes actions efficiently and accurately, but also provides real-time security guarantees throughout the entire execution process. Specifically, this includes: hard real-time instruction distribution, multi-source state information monitoring, real-time stability boundary assessment, and dynamic security policy triggering. This ensures the security of the entire task execution process.
[0089] As a preferred embodiment of the present invention, it further includes: step S5, using the interactive context learning module to adaptively adjust the control parameters in the semantic-action mapping knowledge base based on the actual motion trajectory and user feedback during task execution.
[0090] Example 1
[0091] In such Figure 4 In the typical application scenario shown, "Please carefully pick up the ball under the table," the various modules of the system work collaboratively according to the following process:
[0092] (1) The voice recognition system first converts the user's instructions into text information. Then, the visual perception system scans the working environment to accurately locate the spatial position of the table, the precise coordinates of the ball, and identify potential obstacles such as table legs;
[0093] (2) After receiving this information, the multimodal intent understanding module initiates a deep semantic parsing process: the verb "pick up" is mapped to a grasping motion primitive; the adjective "carefully" is parsed into specific motion constraint parameters, including reducing the end effector speed to 60% of the standard value and increasing the torque detection sensitivity to 1.5 times the standard value; the noun "ball" is combined with the visual recognition results to determine the grasping mode of the palm and set a medium gripping force; the environmental description "under the table" is parsed into an obstacle avoidance trajectory that needs to be executed and the required bending amplitude and torso forward tilt angle are determined.
[0094] (3) Based on the above analysis results, the whole-body motion planning module begins to generate a whole-body coordinated motion trajectory. This module constructs a composite trajectory containing multiple sub-movements through inverse kinematics calculations: the leg joints perform a bending motion to reduce the overall height, the torso tilts forward to approach the target, and the single arm extends and performs a grasping motion. At the same time, the planning module estimates through forward dynamic simulation that executing this trajectory will cause the center of mass to shift forward by 12 cm;
[0095] (4) Upon receiving the planned trajectory, the dynamic balance safety arbitrator immediately initiates a safety verification procedure. Calculations using the stability boundary model reveal that the original trajectory causes the stability margin S to drop to 0.05, below the safety threshold of 0.1. Therefore, the arbitrator requires the planning module to optimize the trajectory. After two iterations of optimization, the planning module adds a hip-sitting motion and an extension of the other arm as a balancing weight to the final solution, increasing the stability margin to 0.15, thus meeting the safety requirements.
[0096] (5) The microkernel safety scheduler begins to execute the optimized safety action sequence. During execution, the scheduler monitors the status of each joint and the pressure distribution on the sole of the foot in real time at a period of 1 millisecond. When it detects that the pressure center of the right foot has shifted abnormally forward by 3 centimeters (possibly due to slight slippage on the ground), the first-level protection mechanism is immediately activated, reducing the movement speed to 70% of the original plan. The speed is gradually restored to the original plan after the pressure distribution returns to normal, and the ball retrieval task is successfully completed.
[0097] Example 2
[0098] In elderly and disabled assistance applications, when faced with a user's instruction to "it's a little hot, bring it over steadily," the system demonstrates its ability to handle complex tasks:
[0099] (1) The multimodal intent understanding module performs multi-level parsing of the instructions: mapping "end" to the horizontal load motion primitive of the two arms; "a little hot" is parsed to extract two sets of key parameters, including requiring the end effector to always maintain a horizontal posture and to enable the vibration suppression algorithm to reduce liquid sloshing; "steady" corresponds to gait parameter adjustment, including shortening the stride to 40% of the standard value, reducing the stride speed to 30% of the standard value, and increasing the stiffness of the whole body joints to 1.8 times the normal value;
[0100] (2) The whole-body motion planning module generates motion trajectories specifically for liquid transport based on these constraints. The planning module uses an asymmetric load balancing algorithm to calculate the centroid fluctuations caused by water surface sloshing in real time and compensates for them by the slight bending of the virtual spine. At the same time, it plans a stable gait with small strides and low height to ensure the stability of the bipedal support phase during movement.
[0101] (3) The dynamic balance safety arbitrator performs full safety monitoring of the end-take process; through the stability boundary model, the arbitrator evaluates the impact of water surface sloshing on system stability in real time. When it is detected that the periodic sloshing of the liquid caused by walking causes the stability margin to fluctuate in the range of 0.08-0.12, the arbitrator requires the planning module to add compensation actions to the trajectory, including adjusting the phase difference of the two arms to counteract the periodic disturbance, and ensuring that the stability margin is always kept above the safety threshold of 0.1;
[0102] (4) The microkernel security scheduler executes the optimized end-to-end task and monitors the system status in real time at a 2-millisecond interval. When a large external disturbance causes the water surface to shake violently, causing the stability margin to drop sharply to 0.05, the scheduler immediately triggers the secondary protection mechanism, instantly stops the movement gait, switches to the stationary balance mode, and increases the stiffness of all joints to 90% of the maximum value. After the water surface returns to stability and the stability margin returns to 0.15, the system automatically resumes movement and finally delivers hot water to the user in a highly stable manner.
[0103] This invention achieves a direct conversion from natural language commands to safe, stable, and human-like full-body coordinated movements through the closed-loop collaboration of four core modules. The system architecture balances the complexity of intelligent decision-making with the real-time nature of underlying control. A semantic mapping mechanism enables the robot to understand rich linguistic connotations, while a dynamic balance guarantee mechanism ensures the safety of complex task execution. This integrated solution, combining interaction, planning, and balance control, lays the foundation for the practical application of humanoid robots in complex scenarios such as home services and elderly and disabled assistance.
[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention in any way. All equivalent transformations or modifications made in accordance with the essence of the present invention should be covered within the protection scope of the present invention.
Claims
1. A humanoid robot full-body collaborative voice control system, characterized in that, Integrated into the humanoid robot body, including: The multimodal intent understanding module is used to receive and recognize the user's voice commands, obtain natural language commands, and combine them with the physical attributes of environmental objects and the user's relative pose information obtained by the visual sensor. By querying the semantic-action mapping knowledge base, the natural language commands are parsed into specific task objectives and their implicit whole-body motion constraints. The full-body motion planning module, connected to the multimodal intent understanding module, is used to generate an initial joint trajectory sequence containing coordinated movements of the "torso-arm-hand-leg" based on the task objective and full-body motion constraints, estimate the range of dynamic centroid offset caused by executing the sequence, and use an asymmetric load balancing algorithm for motion compensation in tasks involving asymmetric loads on both arms. The dynamic balance safety arbitrator, connected to the whole-body motion planning module, is used to perform feasibility verification and smoothing optimization of the initial joint trajectory sequence based on the current posture of the humanoid robot and the estimated dynamic offset range of the center of mass, using a preset stability boundary model, and outputs safe joint control commands. The microkernel safety scheduler, connected to the dynamic balance safety arbiter, responds to the safety joint control command and distributes it to the corresponding joint driver at a fixed scheduling cycle. At the same time, it monitors the status of each joint and the overall balance in real time, and immediately executes the predefined safety protection mechanism when it detects actual center of mass shift or the plantar pressure center approaching the stability boundary. The multimodal intent understanding module, the whole-body motion planning module, the dynamic balance safety arbitrator, and the microkernel safety scheduler are connected in sequence to form a closed-loop control system from voice interaction and whole-body motion planning to dynamic balance control. The core formula of the asymmetric load balancing algorithm is: Among them, △COM x This represents the offset between the actual and desired center of mass positions in the lateral direction due to the load on both arms; m1 and m2 represent the masses of the loads held by the left and right hands, respectively; x1 and x2 represent the lateral coordinates of the loads held by the left and right hands in the robot's torso coordinate system; COM x0 θ represents the desired lateral position coordinates of the robot's center of mass when it is unloaded. spine K represents the virtual spinal curvature angle calculated to compensate for centroid shift; p K d These represent the proportional gain coefficient and the differential gain coefficient of the spinal curvature controller, respectively. Indicates the backswing compensation angle required by the unloaded side arm to provide balancing torque; K comp This represents the compensation coefficient for the arm swing motion. This represents the rate of change of the lateral offset of the centroid with respect to time. The dynamic equilibrium safety arbitrator's built-in stability boundary model is based on Lyapunov stability theory, and its stability is quantified by calculating the stability margin S. Where x represents the state vector of the robot system. Where θ and ω represent the joint angle vector and joint angular velocity vector of all joints of the robot, respectively, and P COM and V COM Let V(x) and Q(x) represent the position vector and velocity vector of the robot's center of mass in three-dimensional space, respectively; V(x) represents the Lyapunov function constructed based on Lyapunov stability theory, used to measure the total energy or generalized distance of the system; P represents a positive definite weight matrix used to configure the weights of the state vector x in the Lyapunov function; Q i θ represents the joint angle deviation weighting coefficient of the i-th joint, 1≤i≤n, where n represents the total number of joints; i θ represents the joint angle of the i-th joint. i-ref The reference angle of the i-th joint is represented by S; S represents the stability margin, where S > 0 indicates stability and S < 0 indicates a risk of instability. This represents the partial derivative of the Lyapunov function with respect to time.
2. The system according to claim 1, characterized in that, The semantic-action mapping knowledge base is built by parsing massive amounts of human-computer interaction data, and it can map adverbs and adjectives in natural language into specific motion control parameters.
3. The system according to claim 1, characterized in that, The security protection mechanism triggered by the microkernel security scheduler includes three levels: Level 1 slows down the movement speed, Level 2 stops the gait and switches to stationary balance mode, and Level 3 immediately locks all joints and activates mechanical brakes.
4. The system according to claim 1, characterized in that, It also includes an interactive context learning module, which records the actual motion trajectory and user feedback after the task is executed, and fine-tunes the parameters of the semantic-action mapping knowledge base based on this to achieve adaptive optimization of system performance.
5. A method for control using the humanoid robot full-body collaborative voice control system according to any one of claims 1-4, characterized in that, Includes the following steps: Step S1, Intent Understanding: The multimodal intent understanding module receives and recognizes the user's voice commands to obtain natural language commands. Combined with environmental visual information, the natural language commands are parsed into task objectives containing whole-body motion constraints through the semantic-action mapping knowledge base. Step S2, Motion Planning: Based on the task objective and whole-body motion constraints, the whole-body motion planning module generates an initial whole-body joint trajectory sequence and estimates the dynamic shift of the center of mass during the execution of the sequence. Step S3, Safety Arbitration: Using a dynamic balance safety arbitrator, based on the dynamic offset of the centroid and the current state of the robot, the initial joint trajectory sequence is verified and optimized using a stability boundary model to generate safe joint control commands; wherein, the stability boundary model is used to calculate the stability margin of the robot; Step S4, Scheduling and Execution: The safety joint control instructions are executed through the microkernel safety scheduler, and the robot's posture stability is monitored in real time. When the stability margin determination system approaches the stability boundary, the predefined safety protection mechanism is automatically executed.
6. The method according to claim 5, characterized in that, In step S1, for instructions containing delivery semantics, the multimodal intent understanding module automatically loads a socially compliant motion template from the semantic-action mapping knowledge base. This template at least constrains the robot during the delivery process. Maintain a smooth motion trajectory and stable speed for the end effector; And after confirming that the item has been received, perform a small-scale withdrawal action.
7. The method according to claim 5, characterized in that, Following step S4, the following is also included: Step S5, Learning Optimization: Through the interactive context learning module, the control parameters in the semantic-action mapping knowledge base are adaptively adjusted based on the actual motion trajectory and user feedback during task execution.
Citation Information
Patent Citations
Teleoperation system and control method for synergistically regulating and controlling single-leg operation and machine body translation of multi-legged robot
CN108791560A
Body posture slope self-adaptive control method of four-foot bionic robot
CN111891252A