Double-arm robot autonomous control system and method based on remote operation and visual features
The intelligent control system, which integrates remote data acquisition and multi-view visual perception, solves the problems of autonomous decision-making and safety of dual-arm robots in dynamic environments, enabling efficient and safe execution of complex tasks.
Patent Information
- Application Number
- CN202511046610.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing dual-arm robot systems have difficulty adapting to changes in unstructured environments in dynamic environments or scenes with diverse objects. They suffer from severe visual noise interference, poor generalization performance of offline training strategies, and lack real-time safety monitoring and human-machine collaboration mechanisms, resulting in a high operation failure rate.
An intelligent control system employing remote data acquisition, multi-view visual perception, cross-modal feature joint encoding, and human-machine hybrid control achieves efficient operation of the autonomous control system through dynamic safety assessment and autonomous decision-making.
It improves the accuracy of autonomous decision-making and operational safety in complex tasks, supports high-degree-of-freedom operations such as grasping, handling, and assembly in various complex environments, and has task generalization and online learning capabilities to ensure system stability and security.
Smart Images

Figure CN120816484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dual-arm robot control systems, and in particular to a dual-arm robot autonomous control system and method based on teleoperation and visual features. Background Art
[0002] Dual-arm robots demonstrate significant advantages in industrial assembly, logistics, and service sectors. Their high degrees of freedom and complex manipulation capabilities support multi-step tasks such as grasping, assembly, and positioning. However, existing systems face significant limitations in dynamic environments or scenarios with diverse objects: traditional planning algorithms struggle to adapt to changes in unstructured environments, deep learning models based on unimodal perception are susceptible to visual noise, offline training strategies experience a sharp drop in generalization performance when faced with unseen objects, and most systems lack real-time safety monitoring and human-robot collaboration, resulting in high operational failure rates. Current technology faces three challenges: First, demonstration data collection relies on a complex calibration process, requiring operators to repeat standardized actions dozens of times; second, a single visual encoding system struggles to reliably link spatial object attributes with operational strategies; and finally, static deployment models are unable to handle execution deviations, and delays in manual takeover in emergencies often exceed safety thresholds. These issues severely hinder the practical application of dual-arm robots in real-world scenarios.
[0003] To address the above bottlenecks, the present invention proposes an intelligent control system that integrates multi-perspective spatiotemporal perception and dynamic safety assessment. Through three technological innovations: efficient acquisition of remote operation data, cross-modal feature joint encoding, and real-time arbitration of human-machine hybrid control, it achieves the simultaneous improvement of autonomous decision-making accuracy and operational safety in complex tasks. Summary of the Invention
[0004] Based on the problems raised in the background technology, the present invention proposes an autonomous control system and method for a dual-arm robot based on teleoperation and visual features.
[0005] The present invention proposes an autonomous control system for a dual-arm robot based on teleoperation and visual features. The autonomous control system is suitable for high-degree-of-freedom operations such as grasping, handling, assembly, placement, and handover in a variety of complex environments. It has task generalization and online learning capabilities. The system includes: The teleoperation acquisition module is used to obtain the operator's upper limb posture information and end-control commands through the motion capture suit and handle, and map them to the dual-arm robot to achieve teleoperation of the robot's arms; The multi-view visual perception module includes a global camera installed on the robot head and local cameras installed on the two end effectors, which are used to collect global and local image information of the environment and target objects; The data synchronization and demonstration acquisition module is used to synchronously record image information, robot joint movements, gripper status, and operator control input during teleoperation to construct a time series demonstration dataset; A policy model training module, which trains a multimodal policy model based on the demonstration data, wherein the model is a deep reinforcement learning model based on Diffusion Policy; An autonomous strategy execution module, configured to receive real-time image and state inputs and generate autonomous action instructions based on the strategy model; The trajectory deviation detection and takeover module is used to compare the current trajectory with the training distribution during the robot's execution of the action. If the deviation exceeds the set threshold, the remote operation takeover interface is automatically triggered; The hybrid control interface module is used to support seamless switching between autonomous and manual control and realize dynamic adjustment of the robot's operating mode.
[0006] Preferably, the multi-view visual perception module further comprises: Multi-view spatiotemporal joint coding technology is used to align the video streams of the global and local cameras in the temporal dimension and jointly represent them through feature fusion methods; The feature fusion method includes a deformable cross-view attention mechanism to address the scale differences and pose inconsistencies between images from different viewpoints.
[0007] Preferably, the strategy model training module further includes: Methods for extracting muscle activation patterns from motion capture data as control priors for policy models; A method for automatically labeling task difficulty based on the operator's hesitation duration, which is the delay time for the operator to initiate action at a key node, is used to assist training scheduling or data weighting.
[0008] Preferably, the trajectory deviation detection and takeover module includes: Dynamic autonomy gating mechanism, which dynamically adjusts the degree of autonomous control based on the current scene understanding results, autonomous trajectory stability indicators and estimated risk scores; The dynamic autonomous gating mechanism determines whether to allow the robot to maintain autonomous action or switch to human takeover based on multimodal perception results.
[0009] Preferably, the strategy model training module further includes: Embedded security learning mechanism, by introducing embedded security loss term: in, is the predicted value of collision probability, is the acceptable risk threshold; The risk score is given in real time by the collision probability prediction model and is used to dynamically adjust the control intensity or autonomous confidence during the strategy output stage.
[0010] Preferably, the hybrid control interface module supports automatic switching of the following three control modes: Full teleoperation mode: The robot's movements are fully controlled by humans through motion capture and handles; Semi-autonomous mode: The robot generates action proposals based on autonomous strategies, and the operator makes decisions by confirming or fine-tuning them. Fully autonomous mode: The robot automatically completes grasping and manipulation tasks based on visual perception and training strategies.
[0011] The present invention also proposes a control method for a dual-arm robot autonomous control system based on teleoperation and visual features, comprising the following steps: Step 1: Teleoperation Demonstration Data Collection The operator wears a motion capture suit and holds a virtual reality (VR) controller to enter teleoperation mode; Control the dual-arm robot to complete a series of operational tasks including grasping, moving, rotating, handing over, placing, etc. The system synchronously collects the following multimodal information and packages it into time series: three-view image streams from the head and left and right end cameras, the operator's skeletal joint angles and gripper status, muscle activation signals estimated by the motion capture suit, and the operator's hesitation time at key movement nodes; The collected data is constructed as a standard multimodal demonstration trajectory dataset and stored in the training module; Step 2: Policy model training Perform data preprocessing and alignment encoding on the collected multimodal trajectory data; The input is fed into an end-to-end control strategy model based on the Diffusion Policy structure for training. Specifically, the model uses: multi-view images fused via a deformable cross-view attention module; muscle activation signals as a priori guidance for the control strategy; and action hesitation time as a weighting factor for task difficulty. Introducing security loss function: Used to constrain security risks in the process of model strategy generation; After model training is completed, it is deployed in the autonomous control module; Step 3: Autonomous Task Execution The system enters "autonomous control mode" or "semi-autonomous mode": it receives images from three cameras and performs spatiotemporal fusion encoding; inputs them into the strategy model to obtain continuous control instructions (such as end-point velocity or joint angles); control instructions are issued and executed by the robot's underlying driver module; real-time status feedback is used for the next step of control and trajectory evaluation; The control loop is closed, autonomously completing the entire process of perception-planning-execution-evaluation; Step 4: Security Judgment and Takeover Mechanism During the execution cycle, the system evaluates the following key indicators in real time: The degree of deviation between the current trajectory and the training data distribution; The risk score output by the collision prediction module; The importance rating of the current task phase (e.g., grasping phase vs. moving phase); When any indicator exceeds the preset threshold, the dynamic autonomous gating mechanism is triggered and the control mode is switched; Degrade to "semi-autonomous" mode with human assistance for judgment; Or switch directly to "teleoperation" mode, where the operator takes full control; At the same time, the system records the takeover trajectory for subsequent model fine-tuning to achieve online learning; Step 5: Mode switching and closed-loop training The system supports dynamic switching of control modes based on the following three trigger sources at any time during task execution; Trajectory anomalies (distribution deviations); Risk prediction exceeds the limit; Manually take over instructions; The takeover trajectory data is added to the data pool as a new demonstration, forming a closed-loop data collection-training-deployment-evaluation-update process, supporting the continuous iterative evolution of the system and achieving long-term stable autonomous operation.
[0012] Compared with the prior art, the present invention has the following advantages and technical effects: High-quality demonstration data generation capability: By combining motion capture with controllers, human-level operation demonstrations with spatiotemporal continuity and muscle intent expression can be quickly captured. Multi-view alignment perception strategy: The cross-view deformable attention mechanism improves the robustness of visual understanding and significantly enhances target recognition and task adaptability. Human strategy prior integration and task difficulty modeling: Muscle activation patterns and hesitation duration are embedded in the strategy training process to effectively improve task generalization capabilities; Trajectory anomaly warning and remote control takeover mechanism: This builds a closed-loop path from "autonomous execution → anomaly detection → human takeover" to improve system stability; Safe learning mechanisms ensure deployment reliability: Collision prediction and embedded loss functions ensure the safe execution of policy actions in complex environments. Seamless collaboration between autonomy and manual control: A dynamic gating mechanism supports three-mode switching, allowing the robot to be fully autonomous or quickly introduce human assistance. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 This is a schematic structural diagram of the dual-arm robot of the present invention; In the figure: 1. Global camera; 2. Six-degree-of-freedom robotic arm; 3. Bracket; 4. Base; 5. Local camera; 6. Five-finger dexterous hand. DETAILED DESCRIPTION
[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0015] like Figure 1 As shown, the system of the present invention mainly includes the following hardware modules: Dual-arm robotic platform: Two robotic arms with 6 degrees of freedom, each equipped with a force-controlled gripper and an RGB-D camera; Head global camera: installed on the robot head or space bracket to provide a panoramic view of the task environment; Motion capture suit and VR controller: worn / held by the operator to capture upper limb posture and movement intention in real time; Data acquisition server: used to synchronously collect image streams, control signals, trajectory status, gripper force control information, etc. Training and deployment computing platform: used to run demonstration strategy training, large model reasoning, autonomous execution logic and risk prediction modules; Control switching module: embedded in the robot's main control system, used to switch between "teleoperation-semi-autonomous-autonomous" modes.
[0016] 1. Teleoperation data collection and strategy training phase To enable the robot's autonomous operation in complex environments, the system first collects high-quality multimodal behavioral data through human teleoperation demonstrations and then learns the operational logic based on an end-to-end deep learning policy model. This phase consists of two main steps: teleoperation demonstration data collection and policy model training.
[0017] 1. Teleoperation Demonstration Collection Process During the demonstration, operators donned high-precision motion capture suits and used virtual reality (VR) controllers to control a dual-arm robot to perform predetermined tasks. These tasks included, but were not limited to, multi-stage processes such as grasping, carrying, rotating, precisely placing objects, handing over two arms, and complex collaborative manipulation. The goal was to provide diverse, high-fidelity examples of human manipulation.
[0018] During this process, the system simultaneously collects the following multimodal information: (1) Multi-view image data: The system is equipped with three cameras, one mounted at the operator's head (global camera) and the other at the center of the end-of-hand tool (left-hand camera and right-hand camera), to simultaneously capture image streams. This setup can simultaneously obtain global layout information of the task environment and local detailed images of the end-of-hand operation site, improving perception accuracy and semantic integrity.
[0019] (2) Motion capture joint data: The motion capture system records the three-dimensional posture changes of the operator's key joints (including shoulders, elbows, wrists, etc.) in real time, and simultaneously obtains the opening and closing angles and operating status of the grippers of both hands to form a precise human motion trajectory.
[0020] (3) Muscle activation information estimation: Based on the data collected by the motion capture suit, the activation degree of the operator's muscle groups (especially the upper limbs) is estimated using biomechanical modeling methods as a priori indicators reflecting the operation load, movement intention and control force.
[0021] (4) Action hesitation time labeling: The system automatically records the operator's stillness or hesitation time before key operation steps, especially the response delay before key frames such as grasping, placing, and handing over, as an indirect quantitative signal of task complexity or operation uncertainty, which is used for sample weight scheduling in subsequent model training.
[0022] The above data is uniformly aligned and synchronized along the timeline, packaged into time series at a fixed frame rate, and constructed into a structured multimodal teleoperation demonstration trajectory sample library. The collected data has high realism and strong physical consistency, making it suitable for end-to-end policy learning.
[0023] 2. Strategy Model Training Method In the strategy learning stage, a generative deep reinforcement learning model with Diffusion Policy as the core is used to perform behavioral cloning and strategy generalization training on the teleoperation trajectory, realizing a direct mapping from perception input to control output.
[0024] The policy model has the following input modules: (1) Multi-view image feature input: Global and local visual features are extracted from three images through a unified visual encoding network (such as a multi-stream convolutional network or Transformer architecture). To address the scale variation and pose mismatch issues between views, a deformable cross-view attention mechanism is introduced into the model to automatically align the image semantic space and fuse them at the feature layer, improving the policy network's ability to understand the linkage between spatial context and detailed operations.
[0025] (2) Muscle activation control prior: The operator's muscle activation estimates are input into the policy network as a set of dynamic vectors, which assist the model in judging the intention intensity and control inertia of the current operation, and guide the generation of more reasonable control trajectories in specific scenarios (such as heavy object grasping).
[0026] (3) Hesitation time scheduling label: The difficulty of samples is weighted and scheduled based on the hesitation time of actions. During model training, higher attention is paid to high-hesitation samples, thereby enhancing the robustness of the strategy to complex task scenarios.
[0027] To ensure the execution security of the policy output actions, an embedded safety constraint loss function is introduced during the training process, which is defined as follows: Introducing embedded security loss terms: in: Indicates the risk probability of collision caused by the current strategy action, which is output by an independent collision prediction module; : Indicates the maximum allowable risk threshold set by the system; This safety loss term is embedded in the total loss function of the policy network and optimized together with the behavior imitation loss, so that the policy has basic collision avoidance and behavioral consistency while maintaining imitation accuracy, so that it can be safely deployed in actual robot systems.
[0028] II. Autonomous Control Execution and Takeover Mechanism After the policy model is trained and deployed to the robot control system, the system enters the autonomous operation phase. Considering the dynamic nature of real-world scenarios and the complexity of tasks, this paper introduces autonomous execution and safe takeover mechanisms to achieve seamless switching between autonomous control and human-machine collaboration, ensuring efficient system operation while maintaining sufficient safety and controllability.
[0029] 1. Multi-view visual strategy execution process In autonomous execution mode, the robotic system uses three cameras to perceive the operating environment in real time: A global camera mounted at the operator's head or on the robot itself provides complete task scene layout and target recognition information. Two local cameras mounted on the left and right end effectors provide high-resolution grasping details and hand-eye registration feedback. This multi-view image stream is first input into a pre-trained visual encoding module to extract multi-scale image features. The three-stream image stream is semantically aligned and spatially fused using the deformable cross-view attention mechanism proposed in this paper to generate a unified visual representation.
[0030] The fused multimodal features serve as one of the inputs to the policy model, which, combined with current state information (such as end-effector position and joint status), outputs a series of continuous control commands. These control variables may include, but are not limited to, velocity or acceleration commands (Cartesian or Joint Space) for the end effector; target pose trajectory points; and the degree of closure of the gripper. These commands are transmitted to the robot's underlying actuators via the control interface module. These commands then update the current state through a closed-loop system using multiple sensors, including force feedback, visual feedback, and position encoders, forming a dynamic and adaptive execution loop.
[0031] 2. Dynamic autonomous gating mechanism and takeover judgment Taking into account the uncertainty of the environment and the generalization boundary of the strategy model, in order to ensure the robustness and security of the system, the present invention introduces a dynamic autonomous gating mechanism during the execution process to evaluate the reliability of the current action in real time and automatically switch to the manual intervention control mode when necessary.
[0032] The system evaluates the following three core indicators in each control loop period (for example, 100ms): (1) Trajectory Divergence: This metric quantifies the degree of deviation between the current policy execution and the learning experience by comparing the distribution distance (KL divergence, Mahalanobis distance, confidence score, etc.) between the current autonomous trajectory and the most similar demonstration trajectory in the training set.
[0033] (2) The collision risk score (Risk Score) is given by the embedded collision probability prediction module. It combines environmental obstacle modeling and action expected trajectory to predict the collision probability in the next few steps and output the risk score value .
[0034] (3) Task Criticality Dynamically adjust the system's fault tolerance threshold during task execution based on task progress graphs, keyframe recognition models, or human annotations. For example, the safety factor during the grasping and placing phase is higher than that during the moving phase.
[0035] The system inputs the above indicators into the gate function for comprehensive judgment: (1) the trajectory deviation exceeds the preset threshold; (2) or the collision risk score is higher than the acceptable threshold; (3) or the current mission is in a critical stage and the operation is uncertain; This triggers an autonomous takeover mechanism, switching the system's operating mode to "semi-autonomous" or "teleoperated." Control is then handed over to a human operator to prevent dangerous behavior or mission failure. This takeover occurs through the following methods: the system automatically suspends autonomous operation to await teleoperation intervention; automatic voice / vibration prompts the operator to intervene; or the operator actively takes over the system through controller input or voice wake-up.
[0036] 3. Incremental learning and data reflow after takeover After a human takes over and completes the target subtask, the system automatically incorporates the newly generated operational data into its database as a new demonstration sample. Through periodic fine-tuning or online updates, the policy model's capabilities are continuously expanded, enabling online incremental learning and improving the system's adaptability to complex and ever-changing tasks.
[0037] The above-mentioned takeover mechanism and learning mechanism effectively build a closed-loop control system from "full autonomy" to "human-machine collaboration" to "full remote control", realizing the unity of mission safety and flexibility.
[0038] 3. Control mode switching mechanism In order to achieve robust operation and human-machine collaborative control of the system in complex and uncertain environments, the present invention designs and implements a dynamic switching mechanism for control modes, which supports the robot to flexibly switch operating strategies at different autonomy levels, thereby taking into account operational efficiency, safety assurance and task generalization capabilities.
[0039] The control mechanism includes the following three operating modes: (1) Full remote operation mode In this mode, full control of the robotic system rests with a human operator. The operator, wearing a motion capture suit, uses a VR controller to control the dual-arm robot in real time. The system does not output any policy instructions during this process; it only collects operational data at a high frequency for subsequent model training or validation.
[0040] This mode is suitable for: the initial demonstration acquisition stage; the unstable stage of early policy training; and the conservative operation stage when the task or scenario is new and has not been learned.
[0041] When the system is running, it supports auxiliary interfaces such as visual feedback and human-machine motion feedback to ensure the accuracy and stability of remote operation.
[0042] (2) Semi-autonomous mode In this mode, the robot system uses a trained policy model to perform target perception and trajectory planning, generating candidate execution plans. However, the final decision-making power remains in the hands of the human operator. The operator can: directly confirm the system-recommended trajectory and issue it for execution; fine-tune the trajectory based on the model's recommendations; or reject the system's recommendations and manually enter a new trajectory.
[0043] The system control interface in this mode supports human-computer interaction, taking into account both autonomy and operational flexibility. It is suitable for scenarios with medium strategy confidence, complex tasks, and drastic environmental changes.
[0044] The advantages of this model include: most decisions are made by the system, reducing the human burden; final human confirmation reduces the risk of misoperation, thereby controlling the safety window; human-adjusted trajectories can also be used for online fine-tuning training, thereby achieving incremental data collection.
[0045] (3) Fully autonomous mode In this mode, the robot system relies entirely on visual input, policy models, and control modules to complete tasks without human intervention. This includes automatic identification of task targets (e.g., objects to be grasped); automatic motion path planning (based on the Diffusion Policy); action execution and closed-loop gripper control; and internal system-wide safety assessment, collision prediction, and trajectory optimization. This mode is suitable for scenarios where the system has fully learned, risks are manageable, and task stability is high, such as repetitive tasks like automated assembly line picking and standard pick-and-place.
[0046] Mode switching trigger mechanism The system supports dynamic switching of the above control modes according to multi-dimensional indicators during task execution. The specific switching basis includes but is not limited to the following three conditions: (1) The trajectory deviation criterion evaluates the distribution difference between the current strategy trajectory and the historical demonstration trajectory in real time. If it exceeds the preset tolerance interval (such as the confidence lower bound, statistical distribution boundary, etc.), the system switches to "semi-autonomous" or "teleoperation" mode.
[0047] (2) Risk scoring criterion The collision risk prediction module outputs a risk value τrisk\tau_{risk}τrisk. If it exceeds the allowable threshold τthreshold\tau_{threshold}τthreshold, the control level is automatically downgraded and the operator is prompted to intervene.
[0048] (3) Manual intervention signal The operator can actively issue a takeover request through the handle button, voice wake-up or system interface, and the system responds and switches to remote operation mode in real time.
[0049] In addition, a finite state machine or Bayesian switching control graph is set up inside the system to fuse multi-modal input signals to achieve dynamic, adaptive and recoverable control mode scheduling.
[0050] The characteristics of the switching mechanism include strong real-time performance: it can respond in milliseconds within the control cycle; clear hierarchy: the control transfer mechanism is clear to avoid command conflicts; natural collaboration: it supports smooth human-machine switching and maintains trajectory continuity; data reuse: the data during the takeover process can be used for strategy updates.
[0051] Through this control mode switching mechanism, the system can maximize the robot's autonomous capabilities while ensuring safety, achieving flexible deployment and generalization capabilities for industrial-grade tasks.
[0052] 4. System Process The overall operational process of the system proposed in this paper covers the entire process from human teleoperation demonstration acquisition to autonomous control strategy learning, execution control, and mode switching, aiming to achieve a high-precision, strong generalization, and safe takeover dual-arm robot autonomous grasping system. The process is mainly divided into the following five stages: Step 1: Teleoperation Demonstration Data Collection 1. The operator wears a motion capture suit and holds a virtual reality (VR) controller to enter teleoperation mode; 2. Control the dual-arm robot to complete a series of operational tasks including grasping, moving, rotating, handing over, placing, etc. 3 The system synchronously collects the following multimodal information and packages it into a time series: three-view image stream (head + left and right end cameras); operator's skeletal joint angles and gripper status; muscle activation signals estimated by the motion capture suit; and the operator's hesitation time at key movement nodes.
[0053] The collected data are constructed as a standard multimodal demonstration trajectory dataset and stored in the training module.
[0054] Step 2: Policy model training 1. Perform data preprocessing and alignment encoding on the collected multimodal trajectory data; 2. Input into the end-to-end control policy model based on the Diffusion Policy structure for training. Specifically, it includes: multi-view images are fused through a deformable cross-view attention module; muscle activation signals are used as a priori guidance for the control policy; and action hesitation time is used as a weighting factor for task difficulty. 3. Introducing security loss function: Used to constrain security risks during model strategy generation.
[0055] After the model training is completed, it is deployed in the autonomous control module.
[0056] Step 3: Autonomous Task Execution 1. The system enters "autonomous control mode" or "semi-autonomous mode": it receives images from three cameras and performs spatiotemporal fusion encoding; it inputs them into the strategy model to obtain continuous control instructions (such as end-point velocity or joint angle); the control instructions are issued and executed by the robot's underlying driver module; and real-time status feedback is used for the next step of control and trajectory evaluation.
[0057] 2 The control loop is closed, autonomously completing the entire process of perception-planning-execution-evaluation.
[0058] Step 4: Security Judgment and Takeover Mechanism During the execution cycle, the system evaluates the following key indicators in real time: The degree of deviation between the current trajectory and the training data distribution; The risk score output by the collision prediction module; The importance rating of the current task phase (e.g., grasping phase vs. moving phase); When any indicator exceeds the preset threshold, the dynamic autonomous gating mechanism is triggered and the control mode is switched: Degrade to "semi-autonomous" mode with human assistance for judgment; Or switch directly to "teleoperation" mode, where the operator takes full control.
[0059] At the same time, the system records the takeover trajectory for subsequent model fine-tuning to achieve online learning.
[0060] Step 5: Mode Switching and Closed-Loop Training The system supports dynamic switching of control modes at any time during task execution based on the following three trigger sources: Trajectory anomalies (distribution deviations); Risk prediction exceeds the limit; Manually take over the command.
[0061] The takeover trajectory data is added to the data pool as a new demonstration, forming a closed-loop data collection-training-deployment-evaluation-update process, supporting the continuous iterative evolution of the system and achieving long-term stable autonomous operation.
[0062] Through the above five-stage process, the system realizes a closed-loop operation from demonstration learning → strategy formation → autonomous execution → safe takeover → continuous iteration, which has the following advantages: Driven by high-quality human demonstration data; Strong space-time fusion perception capability; Multimodal control prior guidance; Embedded security mechanism protection; Seamless switching between autonomous and manual collaborative control.
[0063] The system can be widely used in high-robustness robot operation tasks in complex scenarios such as industrial sorting, medical collaboration, and operations in hazardous environments.
[0064] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A dual-arm robot autonomous control system based on teleoperation and visual features. The autonomous control system is suitable for high-degree-of-freedom operation tasks such as grasping, handling, assembly, placement, and handover in a variety of complex environments. It has task generalization and online learning capabilities, and is characterized by: The system includes: The teleoperation acquisition module is used to obtain the operator's upper limb posture information and end-control commands through the motion capture suit and handle, and map them to the dual-arm robot to achieve teleoperation of the robot's arms; The multi-view visual perception module includes a global camera installed on the robot head and local cameras installed on the two end effectors, which are used to collect global and local image information of the environment and target objects; The data synchronization and demonstration acquisition module is used to synchronously record image information, robot joint movements, gripper status, and operator control input during teleoperation to construct a time series demonstration dataset; A policy model training module, which trains a multimodal policy model based on the demonstration data, wherein the model is a deep reinforcement learning model based on Diffusion Policy; An autonomous strategy execution module, configured to receive real-time image and state inputs and generate autonomous action instructions based on the strategy model; The trajectory deviation detection and takeover module is used to compare the current trajectory with the training distribution during the robot's execution of the action. If the deviation exceeds the set threshold, the remote operation takeover interface is automatically triggered; The hybrid control interface module is used to support seamless switching between autonomous and manual control and realize dynamic adjustment of the robot's operating mode.
2. The dual-arm robot autonomous control system based on teleoperation and visual features according to claim 1 is characterized in that: The multi-view visual perception module further includes: Multi-view spatiotemporal joint coding technology is used to align the video streams of the global and local cameras in the temporal dimension and jointly represent them through feature fusion methods; The feature fusion method includes a deformable cross-view attention mechanism to address the scale differences and pose inconsistencies between images from different viewpoints.
3. The dual-arm robot autonomous control system based on teleoperation and visual features according to claim 1 is characterized in that: The strategy model training module further includes: Methods for extracting muscle activation patterns from motion capture data as control priors for policy models; A method for automatically labeling task difficulty based on the operator's hesitation duration, which is the delay time for the operator to initiate action at a key node, is used to assist training scheduling or data weighting.
4. The dual-arm robot autonomous control system based on teleoperation and visual features according to claim 1 is characterized in that: The trajectory deviation detection and takeover module includes: Dynamic autonomy gating mechanism, which dynamically adjusts the degree of autonomous control based on the current scene understanding results, autonomous trajectory stability indicators and estimated risk scores; The dynamic autonomous gating mechanism determines whether to allow the robot to maintain autonomous action or switch to human takeover based on multimodal perception results.
5. The dual-arm robot autonomous control system based on teleoperation and visual features according to claim 1 is characterized in that: The strategy model training module also includes: Embedded security learning mechanism, by introducing embedded security loss term: in, is the predicted value of collision probability, is the acceptable risk threshold; The risk score is given in real time by the collision probability prediction model and is used to dynamically adjust the control intensity or autonomous confidence during the strategy output stage.
6. The dual-arm robot autonomous control system based on teleoperation and visual features according to claim 1 is characterized in that: The hybrid control interface module supports automatic switching of the following three control modes: Full teleoperation mode: The robot's movements are fully controlled by humans through motion capture and handles; Semi-autonomous mode: The robot generates action proposals based on autonomous strategies, and the operator makes decisions by confirming or fine-tuning them. Fully autonomous mode: The robot automatically completes grasping and manipulation tasks based on visual perception and training strategies.
7. The control method of the dual-arm robot autonomous control system based on teleoperation and visual features according to any one of claims 1 to 6, characterized in that: The following steps are involved: Step 1: Teleoperation Demonstration Data Collection The operator wears a motion capture suit and holds a virtual reality (VR) controller to enter teleoperation mode; Control the dual-arm robot to complete a series of operational tasks including grasping, moving, rotating, handing over, placing, etc. The system synchronously collects the following multimodal information and packages it into time series: three-view image streams from the head and left and right end cameras, the operator's skeletal joint angles and gripper status, muscle activation signals estimated by the motion capture suit, and the operator's hesitation time at key movement nodes; The collected data is constructed as a standard multimodal demonstration trajectory dataset and stored in the training module; Step 2: Policy model training Perform data preprocessing and alignment encoding on the collected multimodal trajectory data; The input is fed into an end-to-end control strategy model based on the Diffusion Policy structure for training. Specifically, the model uses: multi-view images fused via a deformable cross-view attention module; muscle activation signals as a priori guidance for the control strategy; and action hesitation time as a weighting factor for task difficulty. Introducing security loss function: Used to constrain security risks in the process of model strategy generation; After model training is completed, it is deployed in the autonomous control module; Step 3: Autonomous Task Execution The system enters "autonomous control mode" or "semi-autonomous mode": it receives images from three cameras and performs spatiotemporal fusion encoding. These images are then fed into the strategy model to obtain continuous control instructions (such as end-point velocity or joint angles). These control instructions are then issued and executed by the robot's underlying driver module. Real-time status feedback is used for further control and trajectory evaluation. The control loop is closed, autonomously completing the entire process of perception-planning-execution-evaluation; Step 4: Security Judgment and Takeover Mechanism During the execution cycle, the system evaluates the following key indicators in real time: The degree of deviation between the current trajectory and the training data distribution; The risk score output by the collision prediction module; The importance rating of the current task phase (e.g., grasping phase vs. moving phase); When any indicator exceeds the preset threshold, the dynamic autonomous gating mechanism is triggered and the control mode is switched; Degrade to "semi-autonomous" mode with human assistance for judgment; Or switch directly to "teleoperation" mode, where the operator takes full control; At the same time, the system records the takeover trajectory for subsequent model fine-tuning to achieve online learning; Step 5: Mode switching and closed-loop training The system supports dynamic switching of control modes based on the following three trigger sources at any time during task execution; Trajectory anomalies (distribution deviations); Risk prediction exceeds the limit; Manually take over instructions; The takeover trajectory data is added to the data pool as a new demonstration, forming a closed-loop data collection-training-deployment-evaluation-update process, supporting the continuous iterative evolution of the system and achieving long-term stable autonomous operation.
Citation Information
Patent Citations
Teleoperation robot auxiliary system and method based on vision and imitation learning
CN120023813A
Adaptive robot trajectory planning method and system based on deep reinforcement learning
CN120095834A
Multi-mode sensing multi-mechanical-arm cooperative control method and system and robot
CN120244968A
Systems and methods for minimanipulation library adjustments and calibrations of multi-functional robotic platforms with supported subsystem interactions
US20210069910A1
Cited By
Automatic closing robot control system for gas cylinder leakage in dangerous chemical accident site
CN121245845A
Automatic closing robot control system for gas cylinder leakage in hazardous chemical accident site
CN121245845B
Robot control method and device based on uncertainty perception and diffusion strategy fusion
CN121432895A
A robot control method and device based on uncertainty perception and diffusion strategy fusion
CN121432895B
Pose closed-loop planning system and method based on PLC six-surface visual quality inspection
CN121625174A