Dual-mode unified operating system based on vision-language-haptic fusion
By integrating vision, language, and force perception into a unified dual-mode operating system, the problems of lack of force perception and hardware inconsistency in high-precision contact tasks of robot systems are solved, achieving efficient and safe operation control and improving learning efficiency and operation success rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ROSIWIT TECHNOLOGY CO LTD
- Filing Date
- 2026-01-09
- Publication Date
- 2026-04-28
AI Technical Summary
Existing robot systems lack force information fusion, cannot handle high-precision contact tasks, lack interpretability in the conversion from language to physical control, lack energy safety constraints, and have inconsistent teaching and execution hardware, resulting in operational risks and domain migration issues.
A dual-mode unified operating system based on vision-language-force fusion is adopted. Multimodal data acquisition and processing are realized through UMI hardware structure and LWC-VLA architecture. Energy safety constraint control algorithm is designed to realize closed-loop operation from human teaching to robot autonomous execution.
It achieves high-precision simultaneous visual and force sensing acquisition, eliminates human-machine domain differences, supports natural language interaction, ensures safe and stable operation, improves learning efficiency and operation success rate, and reduces system complexity and cost.
Smart Images

Figure CN121928546A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of robotics and artificial intelligence, and in particular to a dual-mode unified operating system based on the fusion of vision, language and force. Background Technology
[0002] This invention relates to the field of robot intelligent control technology, specifically to a dual-mode unified operating system based on vision-language-force fusion, which is particularly suitable for teaching and autonomous execution scenarios of contact-intensive tasks such as precision assembly, insertion and removal, and tightening.
[0003] Existing robots largely rely on vision and pose control for operation, lacking tactile and force information, making it difficult to handle high-precision contact tasks. While traditional VLA models can parse language and vision, they only output positional actions, failing to generate interpretable force control commands and lacking safety or energy constraint mechanisms. Traditional force control algorithms, while providing compliance, cannot automatically generate control parameters from language or vision. Existing teaching systems often have inconsistent data acquisition and execution devices, leading to human-machine domain differences and affecting learning effectiveness. Chinese Patent Publication No. CN118061176A, entitled "A Multimodal Shared Teleoperation System and Method for a Three-Armed Space Robot," discloses a multimodal shared teleoperation system and method for a three-armed space robot. Its shortcomings include the inability of existing technology to meet the demands of complex extravehicular activities such as on-orbit assembly, on-orbit installation, and solar cell replacement. Traditional single-armed robots cannot meet the requirements of operational complexity and safety, placing a significant burden on operators controlling multi-armed robot systems. Chinese Patent Publication No. CN111376263A, entitled "A Composite Robot Human-Machine Collaboration System and Its Cross-Coupling Force Control Method," proposes... A composite robot human-robot collaboration system is provided, but its shortcomings include a lack of overall collaborative functionality in existing composite robots. The robotic arm can only collaborate independently and cannot achieve movement of the composite robot chassis through traction, thus failing to fully utilize its performance advantages. Chinese Patent Publication No. CN120439256A, entitled "Data Acquisition System and Data Processing Method," proposes a data acquisition system and data processing method. Its shortcomings include non-intuitive human-robot interaction, lack of multimodal data, insufficient operational intuition, and environmental sensitivity issues in existing industrial robot skill learning, leading to low skill transfer efficiency and insufficient reliability in complex operating scenarios. Furthermore, existing technologies lack force information fusion, cannot handle contact tasks requiring precise force control, lack interpretability in the conversion from language to physical control, lack energy safety constraint mechanisms, pose operational risks, have inconsistent teaching and execution hardware, suffer from domain migration problems, and lack a unified geometric representation framework for pose and force in different coordinate systems.
[0004] This invention addresses the shortcomings of the existing technology by proposing a unified interface and algorithm that integrates vision, language, and force perception, and has energy safety constraints, thereby achieving a complete closed loop from human instruction to autonomous robot execution. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of the prior art by proposing a dual-mode unified operating system based on vision-language-force fusion, which is suitable for human-computer teaching, embodied intelligence training, and force-controlled operation robots.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A dual-mode unified operating system based on vision-language-force fusion is disclosed. The system includes a dual-mode UMI hardware structure and its control algorithm. The dual-mode UMI hardware structure provides the physical acquisition and execution carrier for the system, and the control algorithm provides data processing and control logic support for the system. Together, they realize a closed-loop operation of human hand teaching, data training and robot autonomous execution. The control algorithm includes LWC-VLA architecture and Force-VLA, where UMI represents a unified operating interface, LWC-VLA represents a language-to-force compilation algorithm for the energy protection layer, and Force-VLA represents a unified pose-force representation method.
[0008] The specific workflow of the system is as follows:
[0009] Step S1, the teaching phase, the operator holds the UMI to complete the target task, and the UMI simultaneously records visual, force, pose and language data to generate a time-aligned multimodal dataset;
[0010] Step S2, training phase: The control algorithm performs Wrenchlet processing on the force data in the multimodal data, trains the LWC-VLA network to learn the mapping relationship from language-vision to force control, ensures energy safety through DPL constraints, and optimizes the model using a dual consistent loss function containing pose geometry loss, wrench parameter loss, power consistency regularization and energy feasibility penalty, where Wrenchlet represents the discretized force symbol unit and DPL represents the differentiable passive energy layer;
[0011] Step S3, execution phase: The UMI is installed on the robot end effector. The control algorithm outputs control parameters based on real-time visual data and task instructions. The UMI is driven to perform operations through a 1kHz force control inner loop, and the DPL provides real-time constraints for safe, stable, and precise force control operations.
[0012] Furthermore, the dual-mode UMI hardware structure includes an integrated sensor system, a mechanical interface design, and a dual-mode switching mechanism;
[0013] The sensor system includes a six-dimensional force and torque sensor, an ultra-wide-angle vision camera, and a spatial positioning marker; the six-dimensional force and torque sensor measures in real time. , , and , , ,in , , Represents forces in different directions in space. , , It represents torques in different directions in space; the ultra-wide-angle vision camera provides first-person perspective observation; the spatial positioning marker uses AR tags and April tags for pose tracking, where AR represents augmented reality tags and April tags represent high-precision two-dimensional coded tags specifically designed for robot positioning and camera calibration;
[0014] The mechanical interface design includes a convertible mechanical interface and a synchronous acquisition and timestamp module; the convertible mechanical interface supports handheld handle mode and robot flange connection mode; the synchronous acquisition and timestamp module is used to ensure the accurate alignment of visual, force, and pose data.
[0015] The dual-mode switching mechanism includes a teaching mode and an execution mode. The teaching mode is operated by a human hand and records multimodal data; the execution mode is installed at the end of the robot to reproduce the operation task.
[0016] Furthermore, the LWC-VLA architecture includes a WT module, an LWC module, a DPL, a CEP, and a TRC, where the WT module represents the force discretization module; the LWC module represents the language-wrench compiler module; the CEP represents the contact event predictor; and the TRC represents the teleoperation-robot coordinate unifier.
[0017] The WT module uses a rule based on the force amplitude mutation threshold, direction vector clustering, and joint partitioning of contact event boundaries to project continuous force signals onto a finite symbol set, preserving contact semantics and force control dynamic features. This is used to discretize continuous force signals into force tokens, establish a semantic representation of force signals, and facilitate neural network processing and learning.
[0018] The LWC module takes into account natural language commands, visual features extracted from the ultra-wide-angle vision camera, and force history context output from the WT module. Its output includes time-varying target force. Impedance parameters , Pose increment This enables intelligent compilation from semantics to physical control parameters;
[0019] The DPL is integrated into the training and execution chain of LWC-VLA. It employs a feasible region determined based on the device's rated power, friction cone parameters, and safe operating torque boundaries. A least-squares differentiable projection algorithm is used to jointly project the screw wrench and torque onto the energy safety set. During the training phase, energy feasibility constraints are applied to the parameters output by the LWC module. During the execution phase, control parameters are projected onto the energy safety set in real time, and the power is adjusted according to the constraints. During training and execution, power constraints, pitch range and friction stability are guaranteed in real time, and a differentiable projection mechanism is used to ensure the passivity and operational stability of the system.
[0020] The CEP predicts the stage of the task in real time based on changes in visual features, force history context and preset scene features, and triggers the corresponding control strategy switching to adapt to the dynamic characteristics requirements of different stages.
[0021] The TRC uses a coordinate transformation algorithm to unify the force and pose coordinate system during human hand teaching with the base coordinate system during robot execution, correcting the deviation of force control parameters caused by differences in the dynamics of the operating subject, so that the force perception law in the teaching data can be directly transferred to the robot execution process, improving the learning fidelity.
[0022] Furthermore, the Force-VLA provides a unified geometric representation specification and physical logic basis for pose and force for LWC-VLA. The core content specifically includes exponential coordinate representation based on SE(3), screw wrench parameterization, unified state space, Force-VLA strategy head, passive energy layer constraint projection, and 1kHz force-pose coupling control loop; where SE(3) is an abbreviation for special Euclidean group, that is, three-dimensional special Euclidean group, also called Lie group, which is used to represent rigid body motion.
[0023] Furthermore, the specific content of the exponential coordinate representation based on SE(3) is as follows:
[0024] The exponent coordinates of SE(3) are calculated using a numerical computation process based on Lie algebras, and the rotation is completed using Rodrigues expansion. The translation part is obtained by block analysis based on the Adjoint structure, where the logarithmic mapping... Used for error feedback and incremental pose solving, where Rodrigues represents an efficient numerical method for calculating the exponential mapping of rotation matrices in robotics, and Adjoint represents the adjoint action on the SE(3) Lie group;
[0025] For pose representation, exponential coordinates are used. For rigid body motion, it can be represented as:
[0026]
[0027] in, Let be the motion spiral axis in the Lie algebra se(3), and let se(3) denote the Lie algebra of SE(3). For the scale parameter, se(3) and SE(3) are related to each other through exponential and logarithmic mappings.
[0028] Furthermore, the specific parameterization of the spiral wrench is as follows:
[0029] The force-torque helix and its parameterized representation are as follows:
[0030]
[0031] in, Indicates torque, Indicates force, Indicates the direction axis of the force. Represents a point on the central axis. Indicates the pitch. Indicates the force amplitude;
[0032] The relationship is represented as follows:
[0033] .
[0034] Furthermore, the specific content of the unified state space is as follows:
[0035]
[0036] in, For torque, This is a passive impedance parameter;
[0037] Maintaining power invariance, a uniform helical force-velocity inner product is used to represent the direction and magnitude of the instantaneous energy flow between the robot's end effector and the environment. This serves as the basis for the system's passive constraint and power invariance, specifically expressed as follows:
[0038]
[0039] in, It represents the rate at which a force does work on a rigid body. This represents the original screw wrench before the transformation, i.e., the original force-torque vector.
[0040] Furthermore, the specific content of the Force-VLA policy header is as follows:
[0041] Its inputs are language description, visual features, Wrenchlet history, and state context;
[0042] Its output includes:
[0043] The desired pose index coordinates on the pose side With pose increment ,in, Indicates the desired motion of the helical axis. This represents the desired angular velocity parameter;
[0044] Force side screw wrench With the passive impedance parameter on the stability side ,in, The axis representing the desired force direction. Indicates the desired force spiral center axis point. This represents the desired force helical pitch. This represents the expected force amplitude. Indicates positional stiffness. Indicates damping, Indicates force stiffness, Indicates maximum power limit. Indicates a friction cone constraint;
[0045] Furthermore, the specific content of the passive energy layer constraint projection is as follows:
[0046] The power remains constant during the transformation, specifically expressed as follows:
[0047]
[0048] in, For the transformed screw wrench, it represents the integrated vector of force and torque. Let be the transpose inverse of the adjoint transformation, and denote the linear transformation operator on SE(3). The transformed torque represents the motion spiral of the rigid body. This indicates that the adjoint transformation is a linear transformation operator on the SE(3) group;
[0049] The set of constraints for the passive energy layer constraint projection includes:
[0050] Power constraints:
[0051] Force amplitude constraint:
[0052] Pitch boundary:
[0053] Friction cone:
[0054] Torque boundary:
[0055] Differentiable projection is specifically expressed as: and Jointly project onto the feasible region.
[0056] Furthermore, the specific content of the 1kHz force-pose coupling control loop is as follows:
[0057] The outer layer generates the target helical pose velocity and target wrench parameters using Force-VLA;
[0058] Inner layer reconstruction target spiral wrench:
[0059]
[0060] in, Indicates the target screw wrench. Represents the desired torque vector. Represents the desired force vector;
[0061] An impedance compliance control law that aligns with the power generated by torque V is used, real-time DPL online constraints ensure safety, and seamless switching between position-dominant and force-dominant phases is achieved.
[0062] Compared with the prior art, the present invention, employing the above technical solution, has the following beneficial effects:
[0063] 1) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, and invented a unified interface that can switch between dual modes to achieve high-precision vision-force synchronous acquisition, eliminate human-machine domain difference, and has strong versatility. A single UMI device can both acquire human teaching and directly serve as a robot end effector, eliminating the domain difference problem caused by inconsistency between teaching and execution hardware, and greatly reducing system complexity and cost.
[0064] 2) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, and designs an intelligent compilation algorithm that can translate natural language into target force and impedance parameters, realizing interpretable mapping from "semantics to force control". It can directly translate natural language instructions and visual features into physically interpretable force control parameters, realizing end-to-end intelligent control from "semantics to force control", without the need for manual parameter tuning, supporting natural language interaction, significantly reducing the threshold for use, and improving the intelligence of the system;
[0065] 3) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, which introduces a differentiable passive energy layer in the control loop to ensure stable and safe operation and prevent energy runaway. Through DPL, it ensures that the system is always in the power-limited area, effectively preventing shock, oscillation and damage, meeting the human-machine collaboration safety standard, and ensuring safe and reliable operation.
[0066] 4) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion. The consistency of acquisition and execution hardware enables high-fidelity imitation learning, resulting in high data quality and significantly improved training efficiency. It supports continuous system learning and online optimization, forming a data closed-loop advantage.
[0067] 5) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, which solves the problems of existing VLA models being limited to pose control, lacking physical constraints, and lacking safety boundaries. Under the unified SE(3) framework of pose and force, it ensures the geometric consistency and power invariance of coordinate transformation, realizes the unified control of force and pose, and its axis, pitch, amplitude and other parameters directly map to the contact geometry, with clear physical meaning, which is convenient for cross-task migration.
[0068] 6) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, which improves the success rate by about 30% and reduces the peak force by about 40% and shortens the operation time by about 25% in contact tasks such as plugging, unplugging and assembly. It is applicable to a variety of contact task scenarios and has outstanding practical value.
[0069] 7) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion, which can access various vision-language-motion models and simulation training platforms, supports multi-robot collaboration, has the potential to serve as a unified interface standard for embodied intelligence, and has strong scalability.
[0070] 8) The present invention proposes a dual-mode unified operating system based on vision-language-force fusion. The output spiral parameters have clear physical meanings, which facilitates fault diagnosis and performance optimization, supports the integration of human expert knowledge, and improves the interpretability and maintainability of the system. Attached Figure Description
[0071] Figure 1 This is a detailed flowchart of the system of the present invention;
[0072] Figure 2 This is a flowchart of the Force-UMI data collection phase of the present invention;
[0073] Figure 3 This is a flowchart of the Force-UM deployment phase of the present invention. Detailed Implementation
[0074] The present invention is described below based on embodiments, but the invention is not limited to these embodiments. In the detailed description of the invention below, certain specific details are described in detail. Those skilled in the art can fully understand the invention even without these detailed descriptions. To avoid obscuring the essence of the invention, well-known methods, processes, flows, elements, and circuits are not described in detail. To make the objectives, technical solutions, and advantages of the invention clearer, the invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention.
[0075] This embodiment relates to a dual-mode unified operating system based on vision-language-force fusion. The system includes a dual-mode UMI hardware structure and its control algorithm. The dual-mode UMI hardware structure provides the system with a physical acquisition and execution carrier, and the control algorithm provides the system with data processing and control logic support, collaboratively realizing a closed-loop operation of human teaching-data training-robot autonomous execution. The control algorithm includes LWC-VLA architecture and Force-VLA, where UMI represents a unified operating interface, LWC-VLA represents a language-to-force compilation algorithm for the energy protection layer, and Force-VLA represents a unified pose-force representation method.
[0076] like Figure 1 As shown, the specific workflow of the system is as follows:
[0077] Step S1, the teaching phase, involves the operator using the UMI to complete the target task. The UMI simultaneously records visual, force, pose, and language data, generating a time-aligned multimodal dataset. A flowchart of the Force-UM data collection phase is shown below. Figure 2 As shown;
[0078] Step S2, training phase: The control algorithm performs Wrenchlet processing on the force data in the multimodal data, trains the LWC-VLA network to learn the mapping relationship from language-vision to force control, ensures energy safety through DPL constraints, and optimizes the model using a dual consistent loss function containing pose geometry loss, wrench parameter loss, power consistency regularization and energy feasibility penalty, where Wrenchlet represents the discretized force symbol unit and DPL represents the differentiable passive energy layer;
[0079] Step S3, the execution phase, involves installing the UMI onto the robot's end effector. The control algorithm outputs control parameters based on real-time visual data and task instructions, driving the UMI to perform operations via a 1kHz force control inner loop. Real-time DPL constraints ensure safe, stable, and precise force control operations. The specific Force-UM deployment flowchart is shown below. Figure 3 As shown.
[0080] Furthermore, the dual-mode UMI hardware structure includes an integrated sensor system, a mechanical interface design, and a dual-mode switching mechanism;
[0081] The sensor system includes a six-dimensional force and torque sensor, an ultra-wide-angle vision camera, and a spatial positioning marker; the six-dimensional force and torque sensor measures in real time. , , and , , ,in , , Represents forces in different directions in space. , , It represents torques in different directions in space; the ultra-wide-angle vision camera provides first-person perspective observation; the spatial positioning marker uses AR tags and April tags for pose tracking, where AR represents augmented reality tags and April tags represent high-precision two-dimensional coded tags specifically designed for robot positioning and camera calibration;
[0082] The mechanical interface design includes a convertible mechanical interface and a synchronous acquisition and timestamp module; the convertible mechanical interface supports handheld handle mode and robot flange connection mode; the synchronous acquisition and timestamp module is used to ensure the accurate alignment of visual, force, and pose data.
[0083] The dual-mode switching mechanism includes a teaching mode and an execution mode. The teaching mode is operated by a human hand and records multimodal data; the execution mode is installed at the end of the robot to reproduce the operation task.
[0084] Furthermore, the LWC-VLA architecture includes a WT module, an LWC module, a DPL, a CEP, and a TRC, where the WT module represents the force discretization module; the LWC module represents the language-wrench compiler module; the CEP represents the contact event predictor; and the TRC represents the teleoperation-robot coordinate unifier.
[0085] The WT module uses a rule based on the force amplitude mutation threshold, direction vector clustering, and joint partitioning of contact event boundaries to project continuous force signals onto a finite symbol set, preserving contact semantics and force control dynamic features. This is used to discretize continuous force signals into force tokens, establish a semantic representation of force signals, and facilitate neural network processing and learning.
[0086] The LWC module takes into account natural language commands, visual features extracted from the ultra-wide-angle vision camera, and force history context output from the WT module. Its output includes time-varying target force. Impedance parameters , Pose increment This enables intelligent compilation from semantics to physical control parameters;
[0087] The DPL is integrated into the training and execution chain of LWC-VLA. It employs a feasible region determined based on the device's rated power, friction cone parameters, and safe operating torque boundaries. A least-squares differentiable projection algorithm is used to jointly project the screw wrench and torque onto the energy safety set. During the training phase, energy feasibility constraints are applied to the parameters output by the LWC module. During the execution phase, control parameters are projected onto the energy safety set in real time, and the power is adjusted according to the constraints. During training and execution, power constraints, pitch range and friction stability are guaranteed in real time, and a differentiable projection mechanism is used to ensure the passivity and operational stability of the system.
[0088] The CEP predicts the stage of the task in real time based on changes in visual features, force history context and preset scene features, and triggers the corresponding control strategy switching to adapt to the dynamic characteristics requirements of different stages.
[0089] The TRC uses a coordinate transformation algorithm to unify the force and pose coordinate system during human hand teaching with the base coordinate system during robot execution, correcting the deviation of force control parameters caused by differences in the dynamics of the operating subject, so that the force perception law in the teaching data can be directly transferred to the robot execution process, improving the learning fidelity.
[0090] Furthermore, the Force-VLA provides a unified geometric representation specification and physical logic basis for pose and force for LWC-VLA. The core content specifically includes exponential coordinate representation based on SE(3), screw wrench parameterization, unified state space, Force-VLA strategy head, passive energy layer constraint projection, and 1kHz force-pose coupling control loop; where SE(3) is an abbreviation for special Euclidean group, that is, three-dimensional special Euclidean group, also called Lie group, which is used to represent rigid body motion.
[0091] Furthermore, the specific content of the exponential coordinate representation based on SE(3) is as follows:
[0092] The exponent coordinates of SE(3) are calculated using a numerical computation process based on Lie algebras, and the rotation is completed using Rodrigues expansion. The translation part is obtained by block analysis based on the Adjoint structure, where the logarithmic mapping... Used for error feedback and incremental pose solving, where Rodrigues represents an efficient numerical method for calculating the exponential mapping of rotation matrices in robotics, and Adjoint represents the adjoint action on the SE(3) Lie group;
[0093] For pose representation, exponential coordinates are used. For rigid body motion, it can be represented as:
[0094]
[0095] in, Let be the motion spiral axis in the Lie algebra se(3), and let se(3) denote the Lie algebra of SE(3). For the scale parameter, se(3) and SE(3) are related to each other through exponential and logarithmic mappings.
[0096] Furthermore, the specific parameterization of the spiral wrench is as follows:
[0097] The force-torque helix and its parameterized representation are as follows:
[0098]
[0099] in, Indicates torque, Indicates force, Indicates the direction axis of the force. Represents a point on the central axis. Indicates the pitch. Indicates the force amplitude;
[0100] The relationship is represented as follows:
[0101] .
[0102] Furthermore, the specific content of the unified state space is as follows:
[0103]
[0104] in, For torque, This is a passive impedance parameter;
[0105] Maintaining power invariance, a uniform helical force-velocity inner product is used to represent the direction and magnitude of the instantaneous energy flow between the robot's end effector and the environment. This serves as the basis for the system's passive constraint and power invariance, specifically expressed as follows:
[0106]
[0107] in, It represents the rate at which a force does work on a rigid body. This represents the original screw wrench before the transformation, i.e., the original force-torque vector.
[0108] Furthermore, the specific content of the Force-VLA policy header is as follows:
[0109] Its inputs are language description, visual features, Wrenchlet history, and state context;
[0110] Its output includes:
[0111] The desired pose index coordinates on the pose side With pose increment ,in, Indicates the desired motion of the helical axis. This represents the desired angular velocity parameter;
[0112] Force side screw wrench With the passive impedance parameter on the stability side ,in, The axis representing the desired force direction. Indicates the desired force spiral center axis point. This represents the desired force helical pitch. This represents the expected force amplitude. Indicates positional stiffness. Indicates damping, Indicates force stiffness, Indicates maximum power limit. Indicates a friction cone constraint;
[0113] Furthermore, the specific content of the passive energy layer constraint projection is as follows:
[0114] The power remains constant during the transformation, specifically expressed as follows:
[0115]
[0116] in, For the transformed screw wrench, it represents the integrated vector of force and torque. Let be the transpose inverse of the adjoint transformation, and denote the linear transformation operator on SE(3). The transformed torque represents the motion spiral of the rigid body. This indicates that the adjoint transformation is a linear transformation operator on the SE(3) group;
[0117] The set of constraints for the passive energy layer constraint projection includes:
[0118] Power constraints:
[0119] Force amplitude constraint:
[0120] Pitch boundary:
[0121] Friction cone:
[0122] Torque boundary:
[0123] Differentiable projection is specifically expressed as: and Jointly project onto the feasible region.
[0124] The outer layer generates the target helical pose velocity and target wrench parameters using Force-VLA;
[0125] Inner layer reconstruction target spiral wrench:
[0126]
[0127] in, Indicates the target screw wrench. Represents the desired torque vector. Represents the desired force vector;
[0128] An impedance compliance control law that aligns with the power generated by torque V is used, real-time DPL online constraints ensure safety, and seamless switching between position-dominant and force-dominant phases is achieved.
[0129] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A dual-mode unified operating system based on vision-language-force fusion, characterized in that, The system includes a dual-mode UMI hardware structure and its control algorithm. The dual-mode UMI hardware structure provides the physical acquisition and execution carrier for the system, and the control algorithm provides data processing and control logic support for the system, working together to realize a closed-loop operation of human teaching-data training-robot autonomous execution. The control algorithm includes LWC-VLA architecture and Force-VLA, where UMI represents a unified operation interface, LWC-VLA represents a language-to-force compilation algorithm for the energy protection layer, and Force-VLA represents a unified pose-force representation method. The specific workflow of the system is as follows: Step S1, the teaching phase, the operator holds the UMI to complete the target task, and the UMI simultaneously records visual, force, pose and language data to generate a time-aligned multimodal dataset; Step S2, training phase: The control algorithm performs Wrenchlet processing on the force data in the multimodal data, trains the LWC-VLA network to learn the mapping relationship from language-vision to force control, ensures energy safety through DPL constraints, and optimizes the model using a dual consistent loss function containing pose geometry loss, wrench parameter loss, power consistency regularization and energy feasibility penalty, where Wrenchlet represents the discretized force symbol unit and DPL represents the differentiable passive energy layer; Step S3, execution phase: The UMI is installed on the robot end effector. The control algorithm outputs control parameters based on real-time visual data and task instructions. The UMI is driven to perform operations through a 1kHz force control inner loop, and the DPL provides real-time constraints for safe, stable, and precise force control operations.
2. The dual-mode unified operating system based on vision-language-force fusion according to claim 1, characterized in that, The dual-mode UMI hardware structure includes an integrated sensor system, a mechanical interface design, and a dual-mode switching mechanism. The sensor system includes a six-dimensional force and torque sensor, an ultra-wide-angle vision camera, and a spatial positioning marker; the six-dimensional force and torque sensor measures in real time. , , and , , ,in , , Represents forces in different directions in space. , , It represents torques in different directions in space; the ultra-wide-angle vision camera provides first-person perspective observation; the spatial positioning marker uses AR tags and April tags for pose tracking, where AR represents augmented reality tags and April tags represent high-precision two-dimensional coded tags specifically designed for robot positioning and camera calibration; The mechanical interface design includes a convertible mechanical interface and a synchronous acquisition and timestamp module; the convertible mechanical interface supports handheld handle mode and robot flange connection mode; the synchronous acquisition and timestamp module is used to ensure the accurate alignment of visual, force, and pose data. The dual-mode switching mechanism includes a teaching mode and an execution mode. The teaching mode is operated by a human hand and records multimodal data; the execution mode is installed at the end of the robot to reproduce the operation task.
3. The dual-mode unified operating system based on vision-language-force fusion according to claim 1, characterized in that, The LWC-VLA architecture includes the WT module, LWC module, DPL, CEP, and TRC, where the WT module represents the force discretization module; the LWC module represents the language-wrench compiler module; the CEP represents the contact event predictor; and the TRC represents the teleoperation-robot coordinate unifier. The WT module uses a rule based on the force amplitude mutation threshold, direction vector clustering, and joint partitioning of contact event boundaries to project continuous force signals onto a finite symbol set, preserving contact semantics and force control dynamic features. This is used to discretize continuous force signals into force tokens, establish a semantic representation of force signals, and facilitate neural network processing and learning. The LWC module takes into account natural language commands, visual features extracted from the ultra-wide-angle vision camera, and force history context output from the WT module. Its output includes time-varying target force. Impedance parameters and Pose increment This enables intelligent compilation from semantics to physical control parameters; The DPL is integrated into the training and execution link of LWC-VLA. It adopts a feasible region determined based on the device's rated power, friction cone parameters, and safe operating torque boundary, and uses the least squares differentiable projection algorithm to jointly project the screw wrench and torque to the energy safety set. During the training phase, energy feasibility constraints are applied to the parameters output by the LWC module; during the execution phase, control parameters are projected onto the energy safety set in real time, and power is adjusted according to the constraints. During training and execution, power constraints, pitch range and friction stability are guaranteed in real time, and a differentiable projection mechanism is used to ensure the passivity and operational stability of the system. The CEP predicts the stage of the task in real time based on changes in visual features, force history context and preset scene features, and triggers the corresponding control strategy switching to adapt to the dynamic characteristics requirements of different stages. The TRC uses a coordinate transformation algorithm to unify the force and pose coordinate system during human hand teaching with the base coordinate system during robot execution, correcting the deviation of force control parameters caused by differences in the dynamics of the operating subject, so that the force perception law in the teaching data can be directly transferred to the robot execution process, improving the learning fidelity.
4. A dual-mode unified operating system based on vision-language-force fusion as described in claim 3, characterized in that, The Force-VLA provides a unified geometric representation specification and physical logic basis for pose and force for LWC-VLA. The core contents specifically include exponential coordinate representation based on SE(3), screw wrench parameterization, unified state space, Force-VLA strategy head, passive energy layer constraint projection and 1kHz force-pose coupling control loop. SE(3) is an abbreviation for special Euclidean group, which is a three-dimensional special Euclidean group, also called Lie group, used to represent rigid body motion.
5. A dual-mode unified operating system based on vision-language-force fusion according to claim 4, characterized in that, The specific content of the exponential coordinate representation based on SE(3) is as follows: The exponent coordinates of SE(3) are calculated using a numerical computation process based on Lie algebras, and the rotation is completed using Rodrigues expansion. The translation part is obtained by block analysis based on the Adjoint structure, where the logarithmic mapping... Used for error feedback and incremental pose solving, where Rodrigues represents an efficient numerical method for calculating the exponential mapping of rotation matrices in robotics, and Adjoint represents the adjoint action on the SE(3) Lie group; For pose representation, exponential coordinates are used. For rigid body motion, it can be represented as: in, Let be the motion spiral axis in the Lie algebra se(3), and let se(3) denote the Lie algebra of SE(3). For the scale parameter, se(3) and SE(3) are related to each other through exponential and logarithmic mappings.
6. A dual-mode unified operating system based on vision-language-force fusion according to claim 5, characterized in that, The specific content of the parameterization of the spiral wrench is as follows: The force-torque helix and its parameterized representation are as follows: in, Indicates torque, Indicates force, Indicates the direction axis of the force. Represents a point on the central axis. Indicates the pitch. Indicates the force amplitude; The relationship is represented as follows: 。 7. A dual-mode unified operating system based on vision-language-force fusion according to claim 6, characterized in that, The specific contents of the unified state space are as follows: in, For torque, This is a passive impedance parameter; Maintaining power invariance, a uniform helical force-velocity inner product is used to represent the direction and magnitude of the instantaneous energy flow between the robot's end effector and the environment. This serves as the basis for the system's passive constraint and power invariance, specifically expressed as follows: in, It represents the rate at which a force does work on a rigid body. This represents the original screw wrench before the transformation, i.e., the original force-torque vector.
8. A dual-mode unified operating system based on vision-language-force fusion according to claim 7, characterized in that, The specific content of the Force-VLA policy header is as follows: Its inputs are language description, visual features, Wrenchlet history, and state context; Its output includes: The desired pose index coordinates on the pose side With pose increment ,in, Indicates the desired motion of the helical axis. Represents the desired angular velocity parameter; Force side screw wrench With the passive impedance parameter on the stability side ,in, The axis representing the desired force direction. Indicates the desired force spiral center axis point. This represents the desired force helical pitch. This represents the expected force amplitude. Indicates positional stiffness. Indicates damping, Indicates force stiffness, Indicates maximum power limit. Indicates a friction cone constraint; The output parameters are directly adapted to the compilation requirements of LWC-VLA.
9. A dual-mode unified operating system based on vision-language-force fusion as described in claim 8, characterized in that, The specific details of the passive energy layer constraint projection are as follows: The power remains constant during the transformation, specifically expressed as follows: in, For the transformed screw wrench, it represents the integrated vector of force and torque. Let be the transpose inverse of the adjoint transformation, and denote the linear transformation operator on SE(3). The transformed torque represents the motion spiral of the rigid body. This indicates that the adjoint transformation is a linear transformation operator on the SE(3) group; The set of constraints for the passive energy layer constraint projection includes: Power constraints: Force amplitude constraint: Pitch boundary: Friction cone: Torque boundary: Differentiable projection is specifically expressed as: and Jointly project onto the feasible region.
10. A dual-mode unified operating system based on vision-language-force fusion according to claim 9, characterized in that, The specific contents of the 1kHz force-pose coupling control loop are as follows: The outer layer generates the target helical pose velocity and target wrench parameters using Force-VLA; Inner layer reconstruction target spiral wrench: in, Indicates the target screw wrench. Represents the desired torque vector. Represents the desired force vector; An impedance compliance control law that aligns with the power generated by torque V is used, real-time DPL online constraints ensure safety, and seamless switching between position-dominant and force-dominant phases is achieved.
Citation Information
Patent Citations
Composite robot man-machine cooperative system and crossed coupling force control method thereof
CN111376263A
Multi-mode sharing teleoperation system and method for three-arm space robot
CN118061176A
Data acquisition system and data processing method
CN120439256A