An ar / vr-based humanoid robot and a control method thereof
By using an AR/VR hybrid interactive terminal and a low-latency communication network, combined with multimodal sensors and edge computing, multimodal interaction and real-time risk assessment are achieved. This solves the problems of low operating efficiency and poor safety of humanoid robots in existing technologies, supports multi-user collaborative operation, and improves the ability to perform tasks in complex environments.
Patent Information
- Application Number
- CN202511620830.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2045-11-07
AI Technical Summary
Existing humanoid robots suffer from low efficiency and poor safety in remote control operation, lack autonomous intelligence, are difficult to adapt to complex and ever-changing task environments, and have a single form of interaction, failing to achieve multimodal input and multi-user collaboration.
It adopts an AR/VR hybrid interactive terminal and a low-latency two-way communication network, combined with multimodal sensors and edge computing, to realize multimodal interaction and real-time risk assessment. It has an integrated intelligent decision-making module to dynamically switch permissions and support multi-user collaborative operation.
It improves operational efficiency and security, achieves a near-realistic operating experience, ensures task continuity and security, and supports efficient multi-user collaboration.
Smart Images

Figure CN121061906B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of humanoid robots, in particular to a humanoid robot based on AR / VR and a control method thereof. BACKGROUND
[0002] Existing humanoid robots for remote control mostly rely on traditional remote control handles or pre-programmed operations, which have poor intuitive operation and high learning threshold, and are difficult to adapt to complex and variable task environments. In high-risk complex scenes such as industry, mine, smelting workshop, disaster rescue, remote control faces many problems: low operation efficiency, single-mode input (such as handle or simple voice command) is not enough to complete complex tasks, and frequent step-by-step operation leads to low efficiency; obvious response lag, limited by communication and processing delay, the user's control action is difficult to synchronize to the robot execution in time; insufficient safety, the robot is easy to cause secondary accidents in dangerous environment due to instruction execution delay or misoperation; lack of autonomous intelligence, traditional robots strictly depend on user input, lack of autonomous judgment and emergency handling ability to environmental risks.
[0003] In recent years, AR / VR technology has been gradually introduced into robot remote control, providing an immersive interactive experience. However, existing solutions still have shortcomings: single interaction form, difficult to support multi-modal input such as voice, gesture, eye movement at the same time; lack of artificial intelligence-based assisted execution and risk prediction; insufficient virtual and real feedback, users have difficulty in obtaining a close-to-real operation experience; unable to realize multi-user collaboration and dynamic permission management. SUMMARY
[0004] The purpose of the present application is to provide a humanoid robot based on AR / VR and a control method thereof, which can solve the technical problems of low efficiency caused by single interaction mode of existing operation method, lack of autonomous judgment and emergency measures for environmental risks, and poor safety.
[0005] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows.
[0006] A humanoid robot based on AR / VR, comprising a humanoid robot body, an AR / VR hybrid interaction terminal, and a low-latency bidirectional communication network for realizing bidirectional communication between the robot body and the interaction terminal.
[0007] The robot body comprises a mobile chassis, a double-arm operating mechanism, and a main control system, a joint servo driving system, a multi-modal sensor assembly, a multi-spectral vision system, an edge computing unit, and an embodied intelligent decision module.
[0008] The multi-modal sensor assembly is used for real-time sensing of original contact and posture information of the robot body, the multi-spectral vision system is used for real-time acquisition of original environment three-dimensional structure, object state and distance information, the edge computing unit is used for fusion processing of data generated by the multi-modal sensor assembly and the multi-spectral vision system, the embodied intelligent decision module can evaluate the risk value of the state and the instruction according to the data fused by the edge computing unit, and generate an interrupt signal to start a preset risk avoidance strategy when the risk value exceeds a safety threshold, and the main control system can perform dynamic permission switching according to the risk value, and partially or wholly take over the control permission of the user, and respond to the interrupt signal of the embodied intelligent decision module, and autonomously interrupt the control flow and execute the risk avoidance strategy, for ensuring operation safety and task continuity.
[0009] The hybrid interaction terminal can recognize the voice and gesture instructions of the user and transmit them to the robot body to realize the control of the robot body by the user, and can also accept the data fused by the edge computing unit to provide somatosensory feedback to the user.
[0010] Further, the embodied intelligent decision module includes a state perception submodule, a risk evaluation submodule, and an emergency decision submodule.
[0011] The state perception submodule is used for fusing force sensation, posture, vision, and distance data generated by the multi-modal sensor assembly and the multi-spectral vision system, and constructing a current context vector.
[0012] The risk evaluation submodule is used for calculating an execution risk value by using a deep reinforcement learning model according to the current context vector and the user input instruction.
[0013] The emergency decision submodule is used for sending an interrupt signal to the main control system to start a preset risk avoidance strategy when the risk value exceeds a safety threshold.
[0014] Further, the main control system can switch between a following mode, a collaborative monitoring mode, and an autonomous emergency mode according to the risk value, in the following mode, the user has complete control over the robot body, in the collaborative monitoring mode, the key instructions of the user need to be verified by the embodied intelligent decision module before being executed, and in the autonomous emergency mode, the robot body suspends the execution of the user instructions and prioritizes the execution of the risk avoidance strategy.
[0015] Further, the preset risk avoidance strategy of the main control system includes emergency braking, retreat avoidance, posture stabilization, and sound and light alarm.
[0016] Further, the multi-modal sensor assembly includes a finger six-dimensional moment sensor, a palm tactile array, and a trunk pressure sensing skin, and the multi-spectral vision system includes a visible light camera, an infrared camera, a depth camera, and a laser radar.
[0017] Further, the hybrid interaction terminal includes a head-mounted display device, a tactile feedback device, a hand tracking module, a microphone array, and a local AI collaborative engine.
[0018] The head-mounted display device is used to generate a three-dimensional environment reconstruction, a safety heat map and a risk prompt according to data returned by the multi-spectrum vision system, and convey to the user through an AR transparent display or a VR immersion mode, and the tactile feedback device is used to generate a vibration waveform and spatial audio to convey the user according to data returned by the multi-modal sensor assembly, so as to obtain a remote operation perception.
[0019] The local AI cooperative engine is used to identify gestures and voice interaction signals of the user through the hand tracking module and the microphone array, and transmit to the robot body to realize control of the robot body by the user.
[0020] Further, the local AI cooperative engine is preconfigured with a motion template library including a plurality of standardized operation processes, and the user can call corresponding standardized operation processes according to a task scene, for simplifying operation.
[0021] A control method of an AR / VR-based humanoid robot, comprising the following steps.
[0022] S1, starting the system and establishing an encrypted communication connection between the hybrid interaction terminal and the robot body through a low-delay bidirectional communication network.
[0023] S2, the user selects an operation mode and inputs an instruction through voice or gesture.
[0024] In the voice instruction mode, the local AI cooperative engine identifies the user's voice through the microphone array and calls corresponding standardized operation processes to generate an executable instruction package and transmit to the main control system.
[0025] In the gesture instruction mode, the local AI cooperative engine collects pose data through the hand tracking module, and generates an execution instruction after filtering and coordinate mapping and transmits to the main control system.
[0026] S4, the embodied intelligent decision-making module continuously calculates a risk value according to data fused by the edge computing unit and the received execution instruction.
[0027] S5, the main control system switches the following modes, the cooperative monitoring mode and the autonomous emergency mode according to the risk value, to realize normal execution, prompt the user to confirm and autonomously take over.
[0028] S6, the hybrid interaction terminal displays a task progress, an environment three-dimensional view and a risk heat map through the head-mounted display device according to data returned by the robot body, and generates a vibration waveform and spatial audio through the tactile feedback device to convey the user, to realize real-time feedback.
[0029] Further, in S4, the calculation formula of the risk value R is
[0030] ,
[0031] wherein, is the collision probability, is the stability index, is the environmental hazard degree, is the communication delay risk, and the weight coefficient is dynamically adjusted by an AI model.
[0032] Further, multi-user permission allocation and collaborative work are supported, multi-users access the same robot body through the mixed interaction terminal, and the main control system can provide different users with differentiated environment perception data views based on the role responsibilities, so as to realize efficient, safe and intelligent remote human-machine collaboration.
[0033] After the above technical scheme is adopted, the present application has the following beneficial effects:
[0034] 1. The AR / VR mixed interaction terminal can complete complex tasks by comprehensively using voice, gesture, eye movement and other multi-modal operation modes, improve operation efficiency, and enable users to obtain close-to-real operation experience through AR / VR display and somatosensory feedback, thereby improving operation efficiency and safety.
[0035] 2. The present application can continuously construct the current situation vector by monitoring the contact and posture information of the robot body, the three-dimensional structure of the environment, the object state and distance information, and calculate the execution risk value by using a deep reinforcement learning model to realize risk prediction and take emergency measures when the risk exceeds the safety threshold.
[0036] 3. The present application can dynamically switch permissions according to the risk value, partially or completely take over the user's control authority, realize normal execution, prompt user confirmation, and autonomous takeover, and ensure operation safety and task continuity.
[0037] 4. The present application uses a 5G and optical fiber hybrid link, supports data encryption and QoS priority transmission, ensures the timeliness of operation, and avoids risks caused by communication delay.
[0038] 5. The present application supports multi-user operation of one robot body, realizes multi-user collaborative operation through permission allocation for different divisions, and improves reliability and efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0039] Fig. 1 is a basic operation process schematic diagram of the present application.
[0040] Fig. 2 is a whole structure schematic diagram of the present application.
[0041] Fig. 3 is a functional module schematic diagram of the AR / VR mixed interaction terminal in the present application.
[0042] Fig. 4 is a workflow schematic diagram of the embodied intelligent decision module.
[0043] Fig. 5 is a flowchart schematic diagram of the control method in the present application.
[0044] Fig. 6 is a flowchart schematic diagram of multi-user collaboration in the present application. DETAILED DESCRIPTION
[0045] In order to make the purpose, technical solutions and advantages of the present application clearer, the characteristics and performance of a humanoid robot based on AR / VR and a control method thereof in the present application are further described in detail below in combination with the drawings and examples.
[0046] Please refer to the accompanying Figs. 1-6 A humanoid robot based on AR / VR, comprising a humanoid robot body, an AR / VR hybrid interaction terminal, and a low-latency bidirectional communication network for realizing bidirectional communication between the robot body and the interaction terminal. The low-latency bidirectional communication network adopts a 5G and optical fiber hybrid link and supports data encryption and QoS priority transmission.
[0047] The robot body comprises a mobile chassis, a dual-arm operating mechanism, and a main control system, a joint servo driving system, a multi-modal sensor assembly, a multi-spectral vision system, an edge computing unit, and an embodied intelligent decision module.
[0048] The multi-modal sensor assembly is used to perceive real-time original contact and posture information of the robot body, and the multi-spectral vision system is used to collect real-time original environmental three-dimensional structure, object state and distance information.
[0049] Specifically, the multi-modal sensor assembly comprises a finger six-moment sensor, a palm tactile array, and a trunk pressure sensing skin, and the multi-spectral vision system comprises a visible light, an infrared, a depth camera and a laser radar.
[0050] The edge computing unit is used to fuse and process data generated by the multi-modal sensor assembly and the multi-spectral vision system, and the embodied intelligent decision module based on a deep reinforcement learning model can evaluate the risk value of the state and the instruction according to the fused data of the edge computing unit, and generate an interrupt signal to start a preset risk avoidance strategy when the risk value exceeds a safety threshold.
[0051] Specifically, the embodied intelligent decision module includes a state perception submodule, a risk assessment submodule, and an emergency decision submodule. The state perception submodule is configured to fuse force sensation, posture, vision, and distance data generated by a multi-modal sensor assembly and a multi-spectrum vision system, and construct a current situation vector. The risk assessment submodule is configured to calculate an execution risk value based on the current situation vector and a user input instruction, using a deep reinforcement learning model. The deep reinforcement learning model adopts a pre-trained deep Q-network (DQN) model, and is trained in a simulation environment through a large number of human-machine interaction accident scenarios, and has a generalization judgment capability for unknown dangers. The emergency decision submodule is configured to send an interrupt signal to the main control system to start a preset risk avoidance strategy when the risk value exceeds a safety threshold.
[0052] The main control system can perform dynamic permission switching according to the risk value, partially or wholly take over the control permission of the user, and respond to the interrupt signal of the embodied intelligent decision module to autonomously interrupt the control flow and execute the risk avoidance strategy, for ensuring operation safety and task continuity.
[0053] Specifically, the main control system can switch between a following mode, a cooperative monitoring mode, and an autonomous emergency mode according to the risk value. When the risk value is below a first threshold, the main control system switches to the following mode, in which the user has full control over the robot body. When the risk value is between the first threshold and a second threshold, the main control system switches to the cooperative monitoring mode, in which key instructions of the user need to be verified by the embodied intelligent decision module before being executed. When the risk value reaches or exceeds the second threshold, the main control system switches to the autonomous emergency mode, in which the robot body suspends execution of the user instructions and preferentially executes the risk avoidance strategy, which includes emergency braking, retreat avoidance, posture stabilization, sound and light alarm, etc.
[0054] The hybrid interaction terminal can recognize voice and gesture instructions of the user and transmit them to the robot body to realize control of the robot body by the user, and simultaneously receive data fused by the edge computing unit to provide somatosensory feedback to the user. The hybrid interaction terminal includes a head-mounted display device, a tactile feedback device, a hand tracking module, an eye tracker, a microphone array, and a local AI collaboration engine.
[0055] Specifically, the head-mounted display device is configured to generate a three-dimensional environment reconstruction, a safety heat map, and a risk prompt based on data returned by the multi-spectrum vision system, and convey them to the user through AR transparent display or VR immersive mode. The tactile feedback device is configured to generate vibration waveforms and spatial audio based on data returned by the multi-modal sensor assembly, and convey them to the user to obtain a remote operation perception. During execution of the robot body, state and environmental information of the robot body are returned to the hybrid interaction terminal in real time, and task progress and key node prompts are displayed in the head-mounted display device.
[0056] The local AI collaborative engine is used to recognize the user's gestures, eye movements, and voice interaction signals through the hand tracking module, eye tracker, and microphone array, and transmit them to the robot body to realize the user's control over the robot body. The local AI collaborative engine uses a Transformer architecture to jointly model voice-action pairs, has the ability to understand ambiguous instructions, and can optimize action recommendations in combination with contextual environmental information. In particular, the local AI collaborative engine has a preset action template library, including various standardized operation processes. Users can call the corresponding standardized operation processes through voice instructions according to the task scenario. The microphone array collects user voice instructions, which are recognized and semantically analyzed by the local AI collaborative engine, matched with the standardized operation processes in the action template library, and an executable instruction package is generated to realize "one-key" task execution. The action template library includes standardized operation processes such as "precision assembly", "emergency power-off", "injured person carrying", etc., and supports adding templates through demonstration playback.
[0057] A control method of an AR / VR-based humanoid robot, comprising the following steps.
[0058] S1, start the system and establish an encrypted communication connection between the hybrid interaction terminal and the robot body through a low-delay bidirectional communication network.
[0059] S2, the user selects an operation mode and inputs instructions through voice or gestures.
[0060] S3, in the voice instruction mode, the local AI collaborative engine recognizes user voice through the microphone array and calls the corresponding standardized operation process to generate an executable instruction package and transmit it to the main control system.
[0061] In the gesture instruction mode, the local AI collaborative engine collects pose data through the hand tracking module, generates an execution instruction after filtering and coordinate mapping, and transmits it to the main control system.
[0062] S4, the embodied intelligent decision-making module continuously calculates the risk value based on the data fused by the edge computing unit and the received execution instruction. The calculation formula of the risk value R is,
[0063] ,
[0064] wherein, is the collision probability, is the stability index, is the environmental hazard degree, is the communication delay risk, and the weight coefficient is dynamically adjusted by the AI model.
[0065] S5, the master system switches the following mode, the cooperative monitoring mode, the autonomous emergency mode according to the risk value, realizes normal execution, prompts the user to confirm, and autonomous takeover.
[0066] S6, the mixed interaction terminal displays task progress, environmental three-dimensional view, risk heat map through the head-mounted display device according to the data returned by the robot body, generates vibration waveform and space audio through the tactile feedback device to convey users, and realizes real-time feedback.
[0067] The system also supports multi-user permission allocation and collaborative work, and multiple users access the same robot body through mixed interaction terminals, and the master system can provide differentiated environmental perception data views to different users based on roles and responsibilities, to realize efficient, safe and intelligent remote human-machine collaboration.
[0068] Specifically, multiple users access the master system of the same robot body through mixed interaction terminals, and respectively assume the roles of navigator, operator, monitor and commander. The master system dynamically adjusts the control authority priority of each role according to the task stage and environmental risk state.
[0069] At the same time, differentiated environmental perception data views are provided to different users based on roles and responsibilities, wherein local environmental information related to the task of the leading role is provided to the leading role, and wide-area environmental information is provided to non-leading roles for path planning or risk prediction. The operation suggestion or recommended path generated by the non-leading role is projected to the AR display interface of the leading role in the form of augmented reality annotation, providing visual guidance.
[0070] The monitor is used for human supervision of risk trends, and initiates manual intervention when AI does not trigger takeover but there is potential danger.
[0071] The commander has global data access authority, and provides voice guidance under normal circumstances, and can initiate a manual takeover request under certain conditions, and suspends the operation authority of other users after takeover.
[0072] The master system can record permission changes, data distribution and collaborative interaction events, and generate audit logs containing context information.
[0073] In specific implementation, as shown in Figs. 1-4 A new type of humanoid robot device based on AR / VR, which is composed of a mixed interaction terminal, a robot body and a low-latency bidirectional communication network, and each part realizes bidirectional data interaction through an encrypted communication link.
[0074] Mixed interaction terminal: including a head-mounted display device that supports free switching between augmented reality (AR) transparent display and virtual reality (VR) immersive mode, integrates microphone array, hand tracking module, eye tracking device, tactile feedback device and local AI collaborative engine. The terminal has multi-modal input capability, can receive voice, gesture, eye movement and other interaction signals, and realizes accurate recognition of user intent through multi-channel fusion algorithm, improving the immersion and interaction efficiency of remote operation.
[0075] The head-mounted display device adopts Microsoft HoloLens 2 or Meta Quest Pro, supports real-time switching between AR transparent display and VR immersive mode, has high-resolution display and low-distortion optical system, and provides users with clear virtual and real fusion field of view. The hand tracking module is based on binocular infrared camera and deep learning algorithm, realizes sub-millimeter level hand pose estimation, supports bare hand interaction and gesture recognition. The microphone array is composed of 6 directional microphones, supports far-field voice acquisition and sound source positioning, and the voice recognition accuracy is more than 95%. The eye tracking device is integrated in the head-mounted display, used to identify the user's attention focus, and assist the AI collaborative engine in intent prediction. The tactile feedback device is a wristband type vibration motor array, supporting spatial directional feedback, providing multi-level vibration intensity and frequency adjustment, enhancing the user's operation sense of presence.
[0076] The local AI collaborative engine is deployed on the local GPU of the terminal, adopts a lightweight Transformer architecture, and processes multi-modal input. The model parameter quantity is less than 5 million, the inference delay is less than 50 ms, supports joint semantic analysis and intent recognition of voice-gesture-eye movement, and has the ability to understand ambiguous instructions such as "pick up that red box". Combined with environmental context information such as object position and lighting conditions, the action recommendation is dynamically optimized. Users can call the preset "action template library" through natural language to realize "one-key" execution of high-complexity tasks such as "precision assembly", "emergency power-off", "wounded person carrying" and other standardized operation processes. The template library supports online updating and user-defined extension, improving the adaptability and scalability of the system.
[0077] The action template library is stored in the cloud server and supports OTA updating. Each template contains metadata such as task name, action sequence, execution condition and safety constraint. For example, the "precision assembly" template contains steps such as "grab part A", "align hole position", "slowly insert", "detect position signal", etc. Each step is bound to force control threshold and visual verification condition. Users can trigger the template by voice command "execute precision assembly", and the system automatically checks the environment conditions before starting to execute the standardized operation process.
[0078] Robot body: a biped / wheeled hybrid mobile platform, about 1.7 meters tall, with a total of 32 degrees of freedom, including a master control system based on a ROS 2 distributed architecture, a joint servo drive system supporting high-precision force control and position closed loop, a mobile chassis with omnidirectional wheels and active obstacle avoidance function, a dual seven-degree-of-freedom redundant manipulator, a multi-spectral vision system, a multi-modal sensor assembly, an edge computing unit, and a body-aware intelligent decision module. The multi-spectral vision system includes visible light, infrared, depth camera and laser radar integrated in the head of the robot body. The multi-modal sensor assembly includes finger six-axis torque sensor, palm tactile array, and trunk pressure sensing skin. The robot has a highly anthropomorphic structure and supports high-degree-of-freedom motion and multi-task execution in complex environments.
[0079] The master control system is based on the NVIDIA Jetson AGX Orin platform, running ROS 2 Humble version, realizing modular task scheduling and communication management, supporting multi-thread concurrency and task priority scheduling. The mobile chassis adopts Mecanum wheel design, supporting omnidirectional movement and in-place rotation, with a maximum movement speed of 2m / s, and has active obstacle avoidance and path planning capabilities. The dual arms are seven-degree-of-freedom redundant manipulators, equipped with replaceable force control grippers at the end, with built-in six-axis torque sensors and tactile arrays, supporting 0.1N level force control accuracy. The trunk is covered with flexible pressure sensing skin for detecting external collisions and supporting multi-point tactile feedback. The head carries a multi-spectral vision system, including an RGB camera, an infrared thermal imager, a ToF depth camera, and a 2D laser radar, supporting all-weather environmental perception and three-dimensional reconstruction.
[0080] The master control system supports dynamic switching of three control modes to adapt to different task scenarios and risk levels, including,
[0081] Normal following mode: the user has full control, suitable for low-risk environments;
[0082] Collaborative monitoring mode: key actions (such as approaching high-temperature equipment, crossing obstacles) need to be verified by the body-aware intelligent module before execution;
[0083] Autonomous emergency mode: when high-risk situations are detected, such as imminent collision or structural instability, the system takes over control to prioritize safety.
[0084] After the risk is resolved, the system automatically restores user control rights and generates an event log containing timestamp, risk type, and takeover reason, supporting post-mortem analysis.
[0085] The body intelligence decision module is deployed in the edge computing unit of the robot body, and a risk assessment and emergency decision mechanism is constructed based on a deep reinforcement learning model such as DQN or Proximal Policy Optimization (PPO). The state perception sub-module fuses multi-source sensor data such as vision, force sense, distance, and attitude in real time to construct a high-dimensional "context vector" S t ∈R n , which includes information such as robot pose, environmental obstacle distribution, object material, and communication delay. The risk assessment sub-module inputs S t and user instructions A t into a pre-trained risk assessment model to calculate the comprehensive risk value R∈[0,1] of the current instruction execution. The emergency decision sub-module supports a configurable multi-level risk response mechanism. When R<R1, the system is in follow mode, and the user can freely control it; when R1≤R<R2, it enters the collaborative monitoring mode, and key instructions need to be verified and executed by AI; when R≥R2, the system automatically interrupts the user control flow and switches to the autonomous emergency mode, and performs emergency braking, retreat avoidance, or attitude stabilization actions, which are executed by the main control system. The model is trained through reinforcement learning in the Isaac Gym simulation environment through a large number of human-machine interaction accident scenarios, and has the ability to generalize the judgment of unknown dangers.
[0086] Low-latency bidirectional communication network: 5G and optical fiber hybrid link is adopted, control instructions and key sensor data are transmitted preferentially through QoS strategy, end-to-end encryption and packet retransmission mechanism are supported, real-time and safety of remote control are ensured. The communication protocol supports dynamic switching between TCP / UDP, the control instruction transmission delay is less than 50ms, the video stream transmission delay is less than 100ms, and the high real-time task demand is met.
[0087] During task execution, the multi-modal sensor data of the robot body is fused and processed by the edge computing unit, and then transmitted back to the hybrid interaction terminal through the low-latency network. The visual information imaging unit in the hybrid interaction terminal generates a three-dimensional environment reconstruction map, a safety heat map, and multi-dimensional feedback signals, and displays them in real time on the visualization interface of the head-mounted display device. The tactile feedback device generates spatial vibration waveform and directional spatial audio according to the returned force sense and contact information, so that the user can obtain a nearly real remote operation perception. When in autonomous emergency mode, the system takes control, outputs high-intensity pulse vibration and directional alarm sound, prompts the user that the current is in emergency state, and improves the operation sense of presence and safety perception ability.
[0088] As shown in Fig. 5 , after the system is powered on, the hybrid interaction terminal and the robot body establish a TLS encrypted connection through a 5G / optical fiber hybrid link, complete identity authentication and parameter synchronization.
[0089] The user wears a head-mounted display device, selects an operation mode such as "manual control" or "template execution", and inputs instructions through gestures or voice.
[0090] If it is a gesture instruction, the hand tracking module collects pose data, which is denoised through sliding mean filtering and Kalman filtering, and then mapped to the robot coordinate system through a homogeneous transformation matrix to generate a motion instruction. If it is a voice instruction, the local AI collaboration engine performs voice recognition and semantic analysis, and matches the action template library.
[0091] The embodied intelligent decision-making module continuously calculates the risk value R, whose calculation formula is:
[0092] ,
[0093] Among them, is the collision probability, is the stability index, is the environmental hazard degree, is the communication delay risk, and the weight is dynamically adjusted by the AI model according to the task type, for example, in the "wounded person carrying" task, the (stability) weight is significantly improved.
[0094] The main control system switches between the following modes according to the risk value: following mode, collaborative monitoring mode, and autonomous emergency mode, realizing normal execution, prompting user confirmation, and autonomous takeover.
[0095] The hybrid interaction terminal displays the task progress, environmental three-dimensional view, and risk heat map through the head-mounted display device based on the data returned by the robot body, and generates vibration waveforms and spatial audio through the tactile feedback device to convey the user, realizing real-time feedback.
[0096] As shown in Fig. 6 , the system supports multi-user collaborative operation mode, allowing multiple operators to access the same robot body through independent hybrid interaction terminals, forming a "human-robot-human" collaborative control architecture. The system defines four collaborative roles: navigator, operator, monitor, and commander. After identity authentication, each role is bound to a permission level and receives perception data and feedback information matching its responsibilities, realizing efficient, safe, and intelligent remote human-robot collaboration.
[0097] The system defines four collaborative roles: navigator, operator, monitor, and commander. After identity authentication, each role is bound to a corresponding permission level and receives perception data and feedback information matching its responsibilities, forming an efficient and safe human-robot collaboration architecture. The core responsibilities and permission mechanisms of each role are as follows.
[0098] The navigator is responsible for environmental exploration and path planning, has the highest operating authority in the "exploration" and "evacuation" task stages, can call sensors such as laser radar and depth camera to generate a global map and plan the optimal route, and can be called in the target operation stage. The role of auxiliary, but still can provide virtual track guidance for operators through the AR interface.
[0099] The operator is responsible for mechanical arm control and specific task execution, and is the leading role in the "precision assembly" and "wounded transport" operation stages, has direct control over the end effector, and receives local operation information such as near-field three-dimensional reconstruction and contact force feedback. In the non-operation stage, it only has observation authority.
[0100] The monitor is responsible for real-time risk monitoring and tactical intervention, and continuously assesses the trend changes of system stability index and risk value R. When potential dangers such as increased communication delay and continuously growing R are detected, warnings can be issued, and when unstructured risks such as building structure hazards are confirmed, the control system can be taken over, and the task can be suspended or the path can be re-planned.
[0101] The commander has global data access authority and final decision-making authority, is responsible for task strategy formulation and cross-stage coordination, and can initiate a manual takeover request at major decision-making nodes, continuous AI failure or emergency situations. After confirmation by all personnel or 3 seconds, it obtains the highest control authority, and the remaining roles switch to read-only mode.
[0102] The system dynamically adjusts the control authority priority of each role according to the task stage and environmental risk state. Typical permission switching strategies include:
[0103] A. In the environmental exploration stage, the navigator has the highest authority;
[0104] B. In the target operation stage, the operator's authority is elevated to the main control role;
[0105] C. When the risk value R ≥ R2, the embodied intelligent module automatically takes over, suspends all human-machine input, and executes emergency risk avoidance;
[0106] D. The commander can apply for takeover at any time, triggering centralized situation control.
[0107] All permission changes are synchronized with visual prompts and multi-modal feedback to notify all users, and generate timestamped encrypted audit logs to support post-mortem and system optimization.
[0108] Multi-user collaboration not only involves permission allocation, but also includes a differentiated data distribution mechanism based on roles. The system selectively processes and directs the multi-modal sensor data collected by the robot itself, ensuring that each user obtains the most relevant information to their responsibilities.
[0109] The operator terminal displays a partial operation view, focusing on the robot status, contact force feedback, and key surrounding information; the navigator receives large-scale spatial structure data, generates a recommended path, and superimposes a virtual track (such as a translucent light strip) in the operator's AR field of view, providing visual guidance similar to "task guidance"; the monitor views the risk heat map, R value trend curve, and stability index to actively predict risks; and the commander terminal presents a full-fusion view, including environment reconstruction, task progress, communication quality, and operation logs of each role, forming a "battle sand table" style decision interface.
[0110] In addition, the system supports a reverse information support mechanism: for example, when the operator is blocked in a narrow passage, the navigator can be requested to assist in evaluating the passable area around, and the system immediately grants the operator access to the environment perception module, generates alternative paths, and pushes them to the operator.
[0111] In summary, the present application provides a new type of humanoid robot device and control method based on AR / VR, which realizes efficient, safe, and intelligent remote human-machine collaborative work through mechanisms such as multi-modal interaction, AI-assisted decision-making, risk dynamic evaluation, and permission switching control, and is suitable for various high-risk complex scenarios such as industry, rescue, and medical treatment, having wide application prospects and promotional value.
[0112] It should be noted that the parts not described in detail in the present scheme are all prior art, and the above examples are only used to illustrate the present application, but the present application is not limited to the above examples. Any simple modification, equivalent change and modification according to the technical essence of the present application to the above examples all fall within the protection scope of the present application.
Claims
1. A humanoid robot based on AR / VR, characterized in that: This includes the humanoid robot body, an AR / VR hybrid interactive terminal, and a low-latency two-way communication network for enabling bidirectional communication between the robot body and the interactive terminal. The robot body includes a mobile chassis, a dual-arm manipulator mechanism, a main control system, a joint servo drive system, a multimodal sensor assembly, a multispectral vision system, an edge computing unit, and an embodied intelligent decision-making module. The multimodal sensor assembly is used to perceive the robot's raw contact and posture information in real time; the multispectral vision system is used to acquire raw environmental 3D structure, object state, and distance information in real time; and the edge computing unit is used to fuse and process the data generated by the multimodal sensor assembly and the multispectral vision system. The embodied intelligent decision-making module can assess the risk value of the state and instructions based on data fused by the edge computing unit, and generate an interrupt signal to activate a preset risk avoidance strategy when the risk value exceeds a safety threshold. The embodied intelligent decision-making module includes a state perception submodule, a risk assessment submodule, and an emergency decision-making submodule. The state perception submodule is used to fuse force, posture, vision, and distance data generated by multimodal sensor components and multispectral vision systems to construct the current situation vector. The risk assessment submodule is used to calculate the execution risk value using a deep reinforcement learning model based on the current situation vector and user input instructions. The emergency decision-making submodule is used to send an interrupt signal to the main control system to activate the preset risk avoidance strategy when the risk value exceeds the safety threshold. The main control system can dynamically switch permissions based on risk values, switching between follow mode, collaborative monitoring mode, and autonomous emergency mode. It can partially or completely take over the user's control permissions and respond to interrupt signals from the embedded intelligent decision-making module, autonomously interrupting the control flow and executing risk avoidance strategies to ensure operational safety and mission continuity. When the risk value is below the first threshold, the main control system can enter follow mode, where the user has complete control over the robot. When the risk value is between the first and second thresholds, it can enter collaborative monitoring mode, where the user's key commands must be verified as risk-free by the embodied intelligent decision-making module before execution. When the risk value reaches or exceeds the second threshold, it can enter autonomous emergency mode, where the robot suspends the execution of user commands and prioritizes risk avoidance strategies. The hybrid interactive terminal can recognize the user's voice and gesture commands and transmit them to the robot body to enable the user to control the robot. At the same time, it receives data fused by the edge computing unit to provide haptic feedback to the user. The hybrid interactive terminal includes a head-mounted display device, a haptic feedback device, a hand tracking module, a microphone array, and a local AI collaborative engine. The head-mounted display device is used to generate 3D environment reconstruction, safety heat map and risk warning based on the data returned by the multispectral vision system, and convey them to the user through AR transparent display or VR immersive mode. The haptic feedback device is used to generate vibration waveforms and spatial audio based on the data returned by the multimodal sensor components to convey to the user to obtain remote operation perception. The local AI collaborative engine is used to recognize the user's gestures and voice interaction signals through the hand tracking module and microphone array, and transmit them to the robot body to enable the user to control the robot body.
2. The AR / VR-based humanoid robot as described in claim 1, characterized in that: The main control system's preset risk avoidance strategies include emergency braking, reverse avoidance, attitude stabilization, and audible and visual alarms.
3. The AR / VR-based humanoid robot as described in claim 2, characterized in that: The multimodal sensor components include a six-dimensional finger torque sensor, a palm tactile array, and a torso pressure-sensing skin. The multispectral vision system includes visible light, infrared, depth cameras, and lidar.
4. The AR / VR-based humanoid robot as described in claim 3, characterized in that: The local AI collaboration engine has a pre-set action template library, which includes a variety of standardized operation procedures. Users can call the corresponding standardized operation procedures according to the task scenario to simplify the operation.
5. The control method for a humanoid robot based on AR / VR as described in claim 4, characterized in that: Includes the following steps, S1. Start the system and establish an encrypted communication connection between the hybrid interactive terminal and the robot body through a low-latency two-way communication network; S2. The user selects the operation mode and inputs commands via voice or gesture; In S3, voice command mode, the local AI collaboration engine recognizes the user's voice through the microphone array and calls the corresponding standardized operation process to generate an executable command package and transmit it to the main control system. In gesture command mode, the local AI collaborative engine collects pose data through the hand tracking module, and after filtering and coordinate mapping, generates execution commands that are transmitted to the main control system. S4. The embodied intelligent decision-making module continuously calculates risk values based on the data fused by the edge computing unit and the received execution instructions; S5. The main control system switches between follow mode, collaborative monitoring mode and autonomous emergency mode according to the risk value to achieve normal execution, prompt user confirmation and autonomous takeover. S6. The hybrid interactive terminal displays task progress, 3D environmental view, and risk heat map through a head-mounted display device based on the data transmitted back by the robot body. It also generates vibration waveforms and spatial audio through a tactile feedback device to convey feedback to the user in real time.
6. The control method for a humanoid robot based on AR / VR as described in claim 5, characterized in that: In S4, the formula for calculating the risk value R is: , in, Let be the collision probability. As a stability index, Based on the degree of environmental harm, Weighting coefficients to account for communication delay risk It is dynamically adjusted by the AI model.
7. The control method for a humanoid robot based on AR / VR as described in claim 5, characterized in that: It supports multi-user permission allocation and collaborative operation. Multiple users can access the same robot body through a hybrid interactive terminal. The main control system can provide different users with differentiated environmental perception data views based on their roles and responsibilities, so as to achieve efficient, safe and intelligent remote human-machine collaboration.
Citation Information
Patent Citations
Artificial intelligence based intelligent virtual reality navigation system for enhanced user experience
US20250044866A1
Multi-modal shared teleoperation system and method for three-arm space robot
WO2025179628A1