Target-binding visual-haptic interaction method and system for robot teleoperation
By generating a lightweight target agent and combining visual and tactile feedback, the problems of unclear target recognition and lack of semantic force feedback in robot teleoperation systems are solved, and high-precision teleoperation in complex scenarios is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF SCI & TECH
- Filing Date
- 2026-06-25
- Publication Date
- 2026-07-31
AI Technical Summary
In multi-target environments, target occlusion scenarios, and dual-arm collaborative operation scenarios, existing robot teleoperation systems make it difficult for operators to accurately perceive the spatial depth relationship between the robot hand and the target object. This results in unclear target recognition, unclear task allocation between the left and right hands, and a lack of object correlation in force feedback, which can easily lead to misgrasping, accidental collisions, and unstable gripping.
By acquiring image and depth data of the remote robot's workspace, target pixels are selected, target segmentation is performed to generate a lightweight target agent, and the target agent's status is updated in real time by combining depth data and robot hand pose, providing visual and tactile feedback to assist the operator in teleoperation.
It improves the operator's spatial perception accuracy of target objects, reduces the risk of accidental grasping and collision, enhances the stability and safety of teleoperation, and is suitable for robot operation in complex environments.
Smart Images

Figure CN122480987A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot teleoperation technology, and in particular to a target-binding visual-mechanical interaction method and system for robot teleoperation. Background Technology
[0002] With the development of robotics, virtual reality, and teleoperation technologies, robot teleoperation systems have been widely applied in industrial maintenance, hazardous environment operations, remote experiments, service robots, and complex environments such as space and deep sea. Through head-mounted displays, master-slave control mechanisms, data gloves, or exoskeletons, operators can perform tasks such as grasping, handling, assembly, and collaborative operations remotely from the work site. To enhance the naturalness and immersion of teleoperation, existing systems typically employ a combination of video visual feedback and hand motion mapping, enabling operators to intuitively control the remote robot's arms and dexterous hands using their own movements.
[0003] In existing technologies, to enhance remote environmental perception capabilities, some solutions employ binocular vision, depth perception, and digital twin technologies to reconstruct the remote working environment in three dimensions and establish a remote scene model in virtual space. For example, existing technologies utilize binocular cameras to acquire environmental depth information and combine it with 3D reconstruction algorithms to construct a digital twin environment, thereby enhancing the operator's understanding of target positions, spatial relationships, and robot states. Other technologies enhance the robot's dual-arm collaborative operation capabilities by introducing virtual force guidance, closed-chain collaborative control, or dual-arm master-slave control strategies, and utilize force sensors to provide feedback on remote contact information, thereby improving teleoperation stability and operational accuracy.
[0004] However, when existing methods rely solely on remote video streams or traditional 3D reconstruction results, operators struggle to accurately perceive the spatial depth relationship between the robot hand and the target object. This is especially problematic in complex scenarios such as multi-object environments, target occlusion, and dual-arm collaborative operations, leading to issues like unclear target recognition and ambiguous task allocation between the left and right hands. Furthermore, existing force feedback mechanisms typically output remote contact force information directly, lacking semantic representation associated with the target object, making it difficult for operators to distinguish whether the contact source is the target object, environmental obstacles, or the work surface. When the robot hand occludes the target, visual segmentation results are unstable, or depth estimation fluctuates, fixed visual cues and raw force feedback can easily mislead operators, resulting in accidental grasping, collisions, and unstable gripping. Summary of the Invention
[0005] To address the challenges of existing methods that rely solely on remote video streams or traditional 3D reconstruction results, where operators struggle to accurately perceive the spatial depth relationship between the robot hand and the target object—especially in complex scenarios such as multi-object environments, target occlusion, and dual-arm collaborative operations—this invention provides a target-binding visual-force interaction method and system for robot teleoperation. This is particularly relevant in complex scenarios involving multiple targets, target occlusion, and dual-arm collaborative operations, where unclear target identification and ambiguous task allocation between the left and right hands are common problems. Furthermore, existing force feedback mechanisms typically output remote contact force information directly, lacking semantic representation associated with the target object, making it difficult for operators to distinguish whether the contact originates from the target object, environmental obstacles, or the work surface. When the robot hand occludes the target, visual segmentation results are unstable, or depth estimation fluctuates, fixed visual cues and raw force feedback can easily mislead operators, leading to issues such as accidental grasping, accidental collisions, and unstable gripping.
[0006] The technical solutions provided by the embodiments of the present invention are as follows: First aspect: This invention provides a target-binding visual-force interaction method for robot teleoperation, comprising: S1: Acquire image data and depth data of the remote robot's workspace, and display the image data on a head-mounted display; S2: Select the target pixel in the image data displayed on the head-mounted display, and determine the target binding channel based on the operator's hand input; S3: Based on the target pixels, the target object is segmented using a target segmentation model to generate a target object mask; S4: Generate a lightweight target agent based on the target object mask and the depth data; S5: Based on the end-effector pose of the remote robotic arm and the wrist pose of the data glove, the lightweight target agent is transformed to the local coordinate system of the corresponding hand to obtain the positional relationship, depth relationship and pose relationship of the target object relative to the corresponding hand; S6: Based on the positional relationship, the depth relationship, and the posture relationship, display a virtual auxiliary object corresponding to the target object in the head-mounted display to assist the operator in performing remote operation; S7: Collect hand and wrist movement data of the data glove, and control the corresponding robotic arm end effector and its end gripper to perform teleoperation based on the hand and wrist movement data; S8: During teleoperation, the lightweight target agent is updated online based on the target object mask, depth data, robotic arm end pose, robotic arm end motion state, and force sensor contact information. S9: Update the virtual auxiliary object and generate the tactile or force feedback signal of the data glove based on the updated lightweight target agent.
[0007] The second aspect: This invention provides a target-binding visual-force interaction system for robot teleoperation, comprising: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the target-binding visual-mechanical interaction method for robot teleoperation as described in the first aspect.
[0008] Third aspect: The present invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the target-binding visual-mechanical interaction method for robot teleoperation as described in the first aspect.
[0009] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: In this embodiment of the invention, a lightweight target agent is generated through target pixel selection, target segmentation, and depth fusion. This agent is then updated online by combining target object masks, depth data, robot hand pose, motion state, and contact information. This maintains the target state without reconstructing the entire remote scene, reducing computational and display burdens. By binding the target agent to the left-hand, right-hand, or dual-hand collaboration channel and obtaining the target's position, depth, and pose relative to the hand in the corresponding hand's local coordinate system, the operator can intuitively perceive the spatial relationship between the target and the execution end, effectively improving target confusion and unclear left-right hand allocation in multi-target, occluded, and dual-arm collaboration scenarios. Furthermore, this invention generates visual cues and data glove tactile or force feedback based on the target agent's state, giving the feedback a semantic connection to the target. This reduces the misleading effects of fixed cues and raw force feedback under conditions of occlusion, segmentation fluctuations, and depth anomalies, minimizing false grasping, false collisions, and gripping instability, and exhibits good platform applicability and scalability. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating a target-binding visual-mechanical interaction method for robot teleoperation, provided in an embodiment of the present invention.
[0012] Figure 2 This is a schematic diagram of the overall system architecture provided in an embodiment of the present invention.
[0013] Figure 3 This is a schematic diagram of a target agent generation process provided in an embodiment of the present invention.
[0014] Figure 4 This is a schematic diagram illustrating the online update process and target agent confidence calculation process provided in an embodiment of the present invention.
[0015] Figure 5 This is a schematic diagram of a contact consistency judgment process provided in an embodiment of the present invention.
[0016] Figure 6 This is a schematic diagram of the structure of a target-binding visual-mechanical interaction system for robot teleoperation provided in an embodiment of the present invention. Detailed Implementation
[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.
[0018] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0019] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0020] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0021] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.
[0022] Reference manual attached Figure 1 - Appendix Figure 5 .
[0023] This invention provides a target-binding visual-mechanical interaction method for robot teleoperation. This method can be implemented using a target-binding visual-mechanical interaction device for robot teleoperation, which can be a terminal or a server. The processing flow of the target-binding visual-mechanical interaction method for robot teleoperation may include the following steps: S1: Acquire image and depth data of the remote robot's workspace and display the image data on the head-mounted display.
[0024] It should be noted that the operating end employs a two-handed inertial data glove and a mixed reality head-mounted display. The data glove has a sampling rate of no less than 100Hz and can collect finger joint angles, wrist posture, and grip force. The head-mounted display is used to show the operator the video stream of the remote robot's workspace, the target virtual auxiliary object, the distance or direction cues of the hand relative to the target, and contact status cues. WebRTC is used for low-latency video transmission, and ZMQ ROS2 bridging is used for command and force feedback data transmission, with end-to-end latency controlled within 50ms. The dual-arm robot includes a left robotic arm, a right robotic arm, a left robotic hand, and a right robotic hand. The image and depth acquisition module may include one or more RGB-D cameras, binocular cameras, structured light cameras, time-of-flight cameras, or other sensors capable of acquiring image and depth information.
[0025] S2: Select the target pixel in the image data displayed on the head-mounted display and determine the target binding channel based on the operator's hand input.
[0026] The target binding channel is used to characterize the association between the target object and the robot's operating end, determining whether the currently selected target is operated by the left robot hand, right robot hand, or both hands working together. When the operator selects a target object from the image displayed on the head-mounted display, the system determines the corresponding target binding channel based on the hand input that triggered the target selection.
[0027] In one possible implementation, determining the target binding channel based on the operator's hand input specifically includes: Obtain the operator's hand input channel.
[0028] When the hand input channel is the left hand input channel, it is determined to be the left robotic arm binding channel.
[0029] When the hand input channel is the right-hand input channel, it is determined to be the right robotic arm binding channel.
[0030] When the hand input channel is a two-hand input channel, it is determined to be a two-hand collaboration binding channel.
[0031] In this embodiment of the invention, this step can establish a clear correspondence between the target and the operating end in the scenario of dual-arm robot teleoperation, improve the consistency between target selection and operation execution, reduce the risk of target confusion in the process of multi-target and multi-arm collaborative operation, thereby improving the interaction efficiency and operation accuracy of dual-arm teleoperation.
[0032] S3: Based on the target pixels, the target object is segmented using a target segmentation model to generate a target object mask.
[0033] In one possible implementation, image segmentation of the target object based on target pixels includes: Based on the target pixels, a cue-based segmentation model, such as point cues, bounding box cues, text cues, or mask cues, is used to generate a target object mask. Cue-based segmentation models include general instance segmentation models or video object segmentation models.
[0034] It should be noted that a general instance segmentation model refers to a visual model capable of detecting multiple target instances in an input image and outputting pixel-level segmentation results for each target. Its output typically includes the target category, target bounding box, and target instance mask. A video target segmentation model, on the other hand, refers to a visual model capable of temporal tracking and pixel-level segmentation of target objects across consecutive video frames. Video target segmentation models are typically based on the temporal continuity of the target's appearance, motion consistency, or feature correlations. After determining the target in the initial frame, they can continuously output target object masks for subsequent video frames.
[0035] By adopting the above model, the system can achieve rapid initialization and segmentation of target objects, and maintain stable updates of target agents during teleoperation, thereby improving the continuity and robustness of target representation.
[0036] S4: Generate a lightweight target agent based on the target object mask and depth data.
[0037] The lightweight target agent is used to provide a low-computational-overhead, real-time-updable structured representation of the current operational target. It maintains the spatial state, operational relationships, and interaction status information of the target object during robot teleoperation. Unlike complete 3D reconstruction models, dense point cloud models, or high-precision digital twin models, the lightweight target agent does not require the construction of a complete geometric shape or detailed surface model of the target object. Instead, it retains key target feature information relevant to the teleoperation task to reduce computational complexity and improve real-time interaction performance.
[0038] In one possible implementation, the lightweight target agent includes a target object mask, a 3D center of the target object, a spatial boundary of the target object, a pose relationship of the target object relative to the end effector of the left or right robotic arm, a target binding channel, a target agent confidence level, and a gripping state.
[0039] Optionally, the system uses the image segmentation model to generate a target object mask based on the target pixel. The initial target mask generation formula is: in, This represents the target mask at the initial moment. This indicates a single target segmentation operation based on point cues. Represents the initial frame image. This represents the target pixel.
[0040] The formula for updating the temporal frame mask tracking is: in, This represents the target mask at time t. This indicates a segmentation operation based on time-series tracking. This represents the frame image at time t. This represents the target mask at time t-1.
[0041] By combining depth data, the 3D center, spatial boundary, or directed bounding box of the target object can be obtained, thereby generating a lightweight target proxy.
[0042] In this embodiment of the invention, by constructing a lightweight target agent, the present invention adopts a structured target representation method to uniformly express the spatial information, operational relationships, and interactive state information of the target object. Compared with a complete 3D reconstruction model, a dense point cloud model, or a high-precision digital twin model, it can reduce computational complexity and data processing overhead, and improve the real-time response capability of the system. At the same time, by incorporating information such as target binding channels, target agent confidence, and grasping status into the unified target representation, the continuous maintenance capability of the target during occlusion, contact, and grasping processes can be enhanced, and the stability, robustness, and interactive reliability of the target representation during robot teleoperation can be improved.
[0043] S5: Based on the end-effector pose of the remote robotic arm and the wrist pose of the data glove, the lightweight target agent is transformed to the local coordinate system of the corresponding hand to obtain the positional, depth, and pose relationships of the target object relative to the corresponding hand.
[0044] Among them, the data glove serves as a human-computer interaction device for remote operation. On the one hand, it is used to collect the operator's finger movement data, hand posture data, and wrist movement data in real time, and generate control inputs for the dual-arm robot and its end effector hand. On the other hand, it is used to receive tactile or force feedback signals generated by the system, enabling the operator to perceive interactive information such as target contact status, grasping stability, slippage risk, or non-target collision, thereby forming a visual-tactile closed-loop remote operation interaction process.
[0045] In one possible implementation, S5 specifically includes: S501: Obtain the pose of the target object in the world coordinate system or the robot base coordinate system.
[0046] S502: Obtain the pose of the corresponding robotic arm end effector in the world coordinate system or robot base coordinate system.
[0047] S503: Based on the pose of the target object and the pose of the robotic arm end effector, determine the relative pose of the target object with respect to the corresponding robotic arm end effector through coordinate transformation calculation.
[0048] S504: Based on the relative pose, determine the positional, directional, depth, and orientation relationships of the target object relative to the corresponding hand.
[0049] Specifically, when generating the target agent, the system determines the target binding channel based on the operator's current input channel. If the operator selects the target using their left hand, the target agent is bound to the left robot hand; if the target is selected using their right hand, the target agent is bound to the right robot hand; and if the target is selected using both hands, the target agent is bound to the two-handed collaborative channel. Based on the robot hand pose or data glove pose, the system transforms the target agent from the camera coordinate system, world coordinate system, or robot base coordinate system to the corresponding hand's local coordinate system, obtaining the target object's position, depth, orientation, and pose relative to the corresponding hand. The pose transformation formula is as follows: in, This indicates the relative pose of the left robotic arm's end effector in the world coordinate system. This indicates the pose of the left robotic arm's end effector in the world coordinate system. This represents the pose of the target object in the world coordinate system. This indicates the relative pose of the right robotic arm's end effector in the world coordinate system. This indicates the pose of the right robotic arm's end effector in the world coordinate system.
[0050] In this embodiment of the invention, by establishing a local coordinate representation of the target object relative to the corresponding robot hand, the invention can realize relative spatial modeling between the target state and the robot's operating end, so that the target representation does not depend on a fixed world coordinate system and can adapt to robot motion, perspective changes, and multi-hand collaborative operation scenarios. At the same time, it provides a unified hand-centered reference framework for subsequent virtual auxiliary display, online target update, and semantic tactile / force feedback generation, thereby improving the system's real-time interaction capability, target tracking stability, and teleoperation accuracy.
[0051] S6: Based on positional, depth, and orientation relationships, display virtual auxiliary objects corresponding to the target object in the head-mounted display to assist the operator in teleoperation.
[0052] Among them, virtual auxiliary objects are used to visualize relevant information of target objects in mixed reality head-mounted displays. By mapping the positional, depth, posture, and interaction states of the target relative to the corresponding robotic hand into perceptible visual cues, the operator can quickly understand the target's spatial state and operational intent, reduce the spatial cognitive burden and target confusion risk during dual-arm teleoperation, and improve the intuitiveness, operational accuracy, and interaction efficiency of target approach, grasping, and collaborative operation.
[0053] Optionally, during teleoperation, the system updates the target agent online based on the segmentation mask, depth data, robot hand pose, hand motion state, historical target state, and force sensor contact information. The standardized data structure of the lightweight target agent is defined as follows: in, A standardized data structure representing a lightweight target agent. Representing the three-dimensional geometric data of the target, This indicates the overall confidence level of the target agent. Indicates the hand binding status. Indicates the state of contact force. This indicates the state during the fetching phase.
[0054] The target agent confidence score can be calculated based on at least one of segmentation stability, depth effectiveness, motion continuity, and contact consistency. The weighted fusion formula for the overall confidence score is: in, Indicates the stability score weight. This represents the segmentation stability score. Indicates the depth-based validity weight. Indicates the depth of effectiveness score. Indicates the score for continuity of motion. Indicates the score for continuity of motion. Indicates the contact consistency weight. This indicates the contact consistency score.
[0055] The formula for calculating segmentation stability is: in, Indicates the mask intersection ratio. This represents the currently predicted target object mask. This indicates the confidence score of the segmentation model output. Indicates inter-frame mask consistency. , , This represents a fixed weighting coefficient.
[0056] When the confidence level of the target agent is low, the system can weaken or hide the target virtual auxiliary object and reduce or suspend the target-related force feedback to avoid misleading the operator.
[0057] It should be noted that in the head-mounted display, the system can use different display styles for targets bound to the left hand, right hand, and both hands. For example, targets bound to the left hand can be displayed with a first color or a first marker, targets bound to the right hand can be displayed with a second color or a second marker, and targets bound to both hands can be displayed with a combination of markers.
[0058] In this way, the operator can intuitively distinguish between left-hand, right-hand, and two-handed task targets, reducing target confusion during two-arm teleoperation.
[0059] S7: Collects hand and wrist motion data from the data glove, and controls the corresponding robotic arm end effector and its end gripper to perform teleoperation based on the hand and wrist motion data.
[0060] S8: During teleoperation, the lightweight target agent is updated online based on the target object mask, depth data, robot arm end pose, robot arm end motion state, and force sensor contact information.
[0061] In one possible implementation, online updates to the lightweight target agent specifically include updates during the target object non-contact phase and updates during the hand occlusion phase.
[0062] During the phase before contact with the target object, the target agent is updated based on the target object's mask and depth data.
[0063] During the hand occlusion phase, the target agent is predicted or maintained based on the target agent, the end effector motion state of the robotic arm, and the target agent confidence level at the previous moment.
[0064] In one possible implementation, online updates to the lightweight target agent also include: When the force sensor of the robot's finger or end effector detects a contact event, the position of the fingertip where the contact event occurred or the candidate contact position of the end effector is obtained.
[0065] Calculate the spatial distance between the candidate location and the lightweight target agent.
[0066] Determine if the spatial distance is greater than or equal to a distance threshold. If so, classify it as target contact and increase the target agent's confidence level. Otherwise, classify it as non-target contact or uncertain contact and decrease the target agent's confidence level.
[0067] It should be noted that those skilled in the art can set the distance threshold according to actual needs, and this invention does not impose any limitations on it.
[0068] Optionally, force sensors can be positioned on the robot's fingertips, palm, wrist, or gripper contact surface to detect distal contact forces, gripping forces, tangential forces, sliding-related signals, or contact events. If the force sensor is a fingertip force sensor, the fingertip location where the force event occurs can be used as a candidate contact location. If the force sensor is a tactile array, the candidate contact location can be calculated based on the pressure center. If the force sensor is a wrist force / torque sensor, the candidate contact location can be estimated by combining it with the robot hand's geometric model.
[0069] The target agent employs a phased dynamic update strategy: in the untouched phase, visual temporal fusion is used. in, Indicates dynamic fusion weights. This represents the visual 3D geometric data of the current frame. This represents the geometric data at time t-1.
[0070] The occlusion phase has been completed. in, express, This represents the target mask at time t-1. This represents the target's velocity at time t-1.
[0071] The hand-object coupling update formula during the stable grasping phase is: in, This represents the pose of the object in the world coordinate system. This indicates the current pose of the robot's wrist in the world coordinate system. This indicates the relative pose of the target object with respect to the robot's wrist in the initial state.
[0072] It should be noted that the online update module periodically updates the lightweight target agent during teleoperation. The update process may include: updating the target object mask, updating the target's 3D center and spatial boundaries, updating the target's pose relative to the left or right hand, updating the target binding channel, updating the target agent confidence, and updating the grasping state. The target agent confidence indicates whether the current target agent is reliable. The target agent confidence is obtained by weighting segmentation stability, depth effectiveness, motion continuity, and contact consistency. Segmentation stability can be determined by the area change, center change, or overlap between the current target object mask and the previous target object mask. Depth effectiveness can be determined by the proportion of effective depth points within the target object mask. Motion continuity can be determined by the distance between the current target's 3D center and the predicted target's 3D center. Contact consistency can be determined by the consistency between force sensor contact events and the target agent's spatial position.
[0073] When the target object is not occluded by the robot hand, the online update module primarily updates the target agent based on the target object's mask and depth data. When the robot hand approaches the target object and causes image occlusion, if the target agent's confidence decreases but the robot hand remains near the target, the online update module can predict or maintain the target agent based on the previous target agent and robot hand movements. Once the system detects that the target object has been stably grasped, the online update module can update the target agent's pose based on changes in robot hand pose and contact information, allowing the target agent to move with the robot hand.
[0074] The visual cues in the head-mounted display can be adjusted based on the target agent's confidence level. For example, when the confidence level is high, a clear target square and distance arrow are displayed. When the confidence level is medium, the target is displayed in a semi-transparent or dashed form. When the confidence level is low, the operator is prompted to reselect the target or pause the target-related force feedback. When the force sensor detects a contact event, the online update module obtains the fingertip position or hand contact candidate position associated with the contact event. In a simplified implementation, if the force sensor output value of the i-th finger is greater than a preset force threshold, the fingertip position of the i-th finger is approximated as a contact candidate position. The system calculates the spatial distance between the contact candidate position and the lightweight target agent. If the contact candidate position is located within the target 3D square or within a preset expansion range of the target 3D square, the contact event is determined to be consistent with the target agent, i.e., target contact.
[0075] If the candidate contact location is far from the target agent, the contact event is determined to be a non-target contact. If a reliable determination cannot be made due to missing depth, calibration errors, or occlusion, it is determined to be an indeterminate contact.
[0076] It should be noted that, to improve the stability of the judgment, the system can also project the contact candidate position onto the image plane, determine whether the projected position falls within or near the target object mask, and compare the consistency between the depth of the contact candidate position and the depth at the target object mask. When at least two of the three-dimensional distance judgment, two-dimensional mask judgment, and depth consistency judgment meet the preset conditions, the system determines that the target is in contact.
[0077] The contact consistency assessment result can be fed back to the target agent's confidence calculation. When it is determined to be target contact, the target agent's confidence is increased and it enters a contact state or a captured state. When it is determined to be non-target contact, the target agent's confidence is decreased or a non-target collision warning is generated. When it is determined to be uncertain contact, the current state is maintained or a weakened prompt is output. The system can set the capture state machine according to the target agent's state. The capture state can include selected state, approach state, contact state, captured state, and target lost state.
[0078] In this embodiment of the invention, by establishing a lightweight target agent online update mechanism, the invention can use the target object mask, depth data, robot motion state, and contact perception information to correct and continuously maintain the target state in real time during robot teleoperation. This allows the target representation to be independent of a single visual observation result, and can maintain continuous target tracking even during robot hand occlusion, target movement, and stable grasping phases. Furthermore, by combining spatial distance judgment, two-dimensional mask consistency, and depth consistency for contact semantic determination, the accuracy of target contact recognition can be effectively improved, and the risk of false triggering and misjudgment of non-target collisions can be reduced, thereby enhancing the interaction reliability, feedback semantics, and operational safety during robot teleoperation.
[0079] S9: Update the tactile or force feedback signals of the virtual auxiliary object and the generated data glove based on the updated lightweight target agent.
[0080] In one possible implementation, S9 specifically includes: S901: Extract the target agent confidence from the updated lightweight target agent.
[0081] S902: If the target agent's confidence level is higher than the first threshold, display a clear visual aid. If the target agent's confidence level is between the first and second thresholds, display a semi-transparent or weakened visual aid. If the target agent's confidence level is lower than the second threshold, prompt the user to reselect the target object or reduce the feedback intensity.
[0082] S903: Detects contact events using a force sensor, outputting continuous contact feedback in case of target contact. In case of non-target contact, it outputs pulsed warning feedback, damping enhancement feedback, or vibration feedback. In case of uncertain contact, it outputs weakened feedback.
[0083] It should be noted that those skilled in the art can set the size of the first threshold and the second threshold according to actual needs, and this invention does not limit this.
[0084] Specifically, the tactile or force feedback generation module generates different feedback based on the grasping state and contact consistency.
[0085] In the approach phase, the system can display a directional arrow, distance value, or depth indicator relative to the hand, and can output weak feedback to the data glove. In the target contact phase, if contact consistency indicates target contact, the system outputs continuous contact force feedback or feedback related to the magnitude of the contact force to the data glove. In the non-target contact phase, the system outputs pulsed warning feedback, damping enhancement feedback, or vibration feedback to the data glove, and displays red or other warning indicators on the head-mounted display. In the grasped phase, the system determines whether the target is stably grasped based on multi-finger contact, gripping force stability, and the target agent's hand movement. When over-gripping is detected, the system can increase the data glove's feedback resistance or issue an over-force warning. When slippage risk is detected, the system can output intermittent pulses or directional vibration warnings. Through this method, the force signal is not simply played back directly to the operator, but rather serves as observation information updated online by the target agent, and simultaneously as feedback output modulated by the target agent's state, thus giving the tactile or force feedback target semantics. The robot control module generates control commands for the dual-arm robot based on the hand and wrist movement data collected by the data glove. Wrist motion data can be mapped to the robot arm end-effector pose, and hand motion data can be mapped to the robot finger joint angles or gripper opening and closing degrees.
[0086] In this embodiment of the invention, by establishing a visual-force joint feedback mechanism based on a lightweight target agent state, the invention dynamically modulates the feedback output using target agent confidence, contact consistency, and grasping state. This enables the system to generate visual aids and tactile or force feedback with target semantics based on different operational states such as target approach, target contact, non-target collision, stable grasping, and slip risk. Compared with the traditional direct force playback method, this invention can effectively reduce feedback ambiguity caused by noisy contact, false collisions, or uncertain observations, and improve the operator's ability to understand changes in target state, perceive grasping stability, and enhance operational accuracy, interactive robustness, and safety during dual-arm teleoperation.
[0087] Reference manual attached Figure 6 The diagram shows a target-binding visual-mechanical interaction system for robot teleoperation provided by the present invention.
[0088] The present invention also provides a target-binding visual-mechanical interaction system 20 for robot teleoperation, applied to the above-mentioned target-binding visual-mechanical interaction method for robot teleoperation, comprising: Processor 201.
[0089] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, they implement the target-binding visual-mechanical interaction method for robot teleoperation as described in the method embodiment.
[0090] The target-binding visual-mechanical interaction system 20 for robot teleoperation provided by the present invention can execute the above-mentioned target-binding visual-mechanical interaction method for robot teleoperation and achieve the same or similar technical effects. To avoid repetition, the present invention will not elaborate further.
[0091] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.
[0092] It should also be understood that the memory in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0093] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0094] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0095] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0096] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0097] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0098] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0099] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0100] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0101] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0102] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0103] This invention provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the target-binding visual-mechanical interaction method for robot teleoperation as described in the method embodiment.
[0104] The present invention provides a computer-readable storage medium that can implement the steps and effects of the target-binding visual-mechanical interaction method for robot teleoperation described in the above method embodiments. To avoid repetition, the present invention will not elaborate further.
[0105] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0106] The following points need to be explained: (1) The accompanying drawings of the embodiments of the present invention only involve the structures involved in the embodiments of the present invention. Other structures can refer to the general design.
[0107] (2) For clarity, the thickness of layers or regions is enlarged or reduced in the drawings used to describe embodiments of the invention, i.e., these drawings are not drawn to scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being “above” or “below” another element, the element may be “directly” located “above” or “below” the other element or there may be intermediate elements.
[0108] (3) Where there is no conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.
[0109] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. The scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A target-binding visual-haptic interaction method for robot teleoperation, characterized in that, include: S1: Acquire image data and depth data of the remote robot's workspace, and display the image data on a head-mounted display; S2: Select the target pixel in the image data displayed on the head-mounted display, and determine the target binding channel based on the operator's hand input; S3: Based on the target pixels, the target object is segmented using a target segmentation model to generate a target object mask; S4: Generate a lightweight target agent based on the target object mask and the depth data; S5: Based on the end-effector pose of the remote robotic arm and the wrist pose of the data glove, the lightweight target agent is transformed to the local coordinate system of the corresponding hand to obtain the positional relationship, depth relationship and pose relationship of the target object relative to the corresponding hand; S6: Based on the positional relationship, the depth relationship, and the posture relationship, display a virtual auxiliary object corresponding to the target object in the head-mounted display to assist the operator in performing remote operation; S7: Collect hand and wrist movement data of the data glove, and control the corresponding robotic arm end effector and its end gripper to perform teleoperation based on the hand and wrist movement data; S8: During teleoperation, the lightweight target agent is updated online based on the target object mask, depth data, robotic arm end pose, robotic arm end motion state, and force sensor contact information. S9: Update the virtual auxiliary object and generate the tactile or force feedback signal of the data glove based on the updated lightweight target agent.
2. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, The process of determining the target binding channel based on the operator's hand input specifically includes: Acquire the operator's hand input channel; When the hand input channel is the left hand input channel, it is determined to be the left robotic arm binding channel; When the hand input channel is a right-hand input channel, it is determined to be the right robotic arm binding channel; When the hand input channel is a two-hand input channel, it is determined to be a two-hand cooperative binding channel.
3. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, The image segmentation of the target object based on the target pixels includes: Based on the target pixels, a cue-based segmentation model, such as point cues, box cues, text cues, or mask cues, is used to generate a mask for the target object; the cue-based segmentation model includes a general instance segmentation model or a video target segmentation model.
4. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, The lightweight target agent includes a target object mask, a target object 3D center, a target object spatial boundary, a target object pose relative to the end effector of the left or right robotic arm, a target binding channel, a target agent confidence level, and a grasping state.
5. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, S5 specifically includes: S501: Obtain the pose of the target object in the world coordinate system or the robot base coordinate system; S502: Obtain the pose of the corresponding robotic arm end effector in the world coordinate system or robot base coordinate system; S503: Based on the pose of the target object and the pose of the robotic arm end effector, determine the relative pose of the target object with respect to the corresponding robotic arm end effector through coordinate transformation calculation; S504: Based on the relative pose, determine the positional relationship, orientation relationship, depth relationship and posture relationship of the target object relative to the corresponding hand.
6. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, The online update of the lightweight target agent specifically includes updates during the target object non-contact phase and updates during the hand occlusion phase; During the phase before contact with the target object, the target agent is updated based on the target object's mask and depth data. During the hand occlusion phase, the target agent is predicted or maintained based on the target agent, the end effector motion state of the robotic arm, and the target agent confidence level at the previous moment.
7. The target-binding visual-force interaction method for robot teleoperation according to claim 1, characterized in that, The online update of the lightweight target agent also includes: When the force sensor of the robot's finger or end effector detects a contact event, the position of the fingertip where the contact event occurred or the candidate contact position of the end effector is obtained; Calculate the spatial distance between the candidate location and the lightweight target agent; Determine whether the spatial distance is greater than or equal to a distance threshold; if so, determine it as target contact and increase the target agent confidence; otherwise, determine it as non-target contact or uncertain contact and decrease the target agent confidence.
8. The target-binding visual-force interaction method for robot teleoperation according to claim 7, characterized in that, S9 specifically includes: S901: Extract the target agent confidence from the updated lightweight target agent; S902: If the target agent confidence level is higher than the first threshold, display a clear visual aid object; if the target agent confidence level is between the first threshold and the second threshold, display a semi-transparent or weakened visual aid object; if the target agent confidence level is lower than the second threshold, prompt the user to reselect the target object or reduce the feedback intensity. S903: Detects contact events with the force sensor, outputs continuous contact feedback in the case of target contact; outputs pulse warning feedback, damping enhancement feedback, or vibration feedback in the case of non-target contact; outputs weakened feedback in the case of uncertain contact.
9. A target-binding visual-force interaction system for robot teleoperation, characterized in that, include: processor; A memory storing computer-readable instructions, which, when executed by the processor, implement the target-binding visual-mechanical interaction method for robot teleoperation as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the target-binding visual-mechanical interaction method for robot teleoperation as described in any one of claims 1 to 8.