Man-machine interaction system and method for robot training and data acquisition
By introducing wearable sensors and vibrating motor arrays into the robot system to provide tactile feedback, and combining multimodal data sets to train the diffusion model, the problem of insufficient force perception in complex scenarios is solved, and high-precision autonomous operation and environmental perception are improved.
Patent Information
- Application Number
- CN202511051786.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-07-29
AI Technical Summary
Existing robots lack effective force perception and tactile feedback in complex and fine operation scenarios, resulting in insufficient operating accuracy and reliability, making it difficult to meet the operation needs of high precision and high reliability.
The main equipment instruction system is used to capture the operator's hand posture through wearable sensors and generate attitude instructions. It combines the ring-distributed vibration motor array to provide tactile feedback. The data processing system is mapped as a robot arm control signal. The robot self-learning module trains the diffusion model based on the multimodal data set to realize closed-loop control of tactile feedback and autonomous operation.
It improves the robot's independent operation ability in complex scenarios, enhances the operator's environment perception ability, optimizes the human-computer interaction algorithm, reduces hardware costs, improves operation accuracy and efficiency, and realizes the optimization of the robot's independent learning and operation level.
Smart Images

Figure CN120552080A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to robots and human-computer interaction technology, and in particular to a human-computer interaction system and method for robot training and data collection. Background Art
[0002] Traditional robots exhibit significant shortcomings in scenarios requiring precise manipulation and adaptability to complex environments, such as remote surgery, precision industrial assembly, and operations in partially visible or obscured environments. These scenarios often place extremely high demands on the robot's force perception and control accuracy, operational flexibility, and reliability. They have limitations in force feedback, environmental perception, and autonomous decision-making, and are unable to make precise adjustments and flexibly respond to real-time conditions like human experts. For example, when cutting fruit, the lack of visibility of the core makes it difficult to accurately avoid it using vision alone. Operators often rely on force feedback provided by the robot to sense cutting resistance and adjust their operations accordingly.
[0003] To overcome this bottleneck and truly meet the demand for high-precision, high-reliability robotic operations in various fields, expert-guided operation has become an indispensable tool. Remote control of robots by experts can directly empower them to perform complex and precise tasks, while also accumulating valuable expert operation data and skills for autonomous learning. Therefore, the exploration and development of human-machine interaction technologies, particularly those emphasizing expert-guided operation and multimodal information fusion, have become key to solving the challenges of robotic application in complex and precise operation scenarios, achieving technological advancement, and expanding the scope of application. In precision assembly, hazardous environments, and environments lacking visual information, expert-guided robots can perform precise operations, ensure production safety, and improve efficiency and quality. The core significance of teleoperation technology lies in establishing a human-machine collaborative model that integrates human intelligence with robotic execution.
[0004] However, while expert-led remote operation demonstrates significant advantages for complex and delicate tasks, remote control models that rely solely on human operators still have limitations in terms of efficiency and scalability. The challenge facing the development of robotics technology is how to effectively transform the operational experience of human experts, particularly the force feedback information inherent in teleoperation, into autonomous robot operation.
[0005] Kofman et al. proposed a computer vision-based robotic teleoperation scheme. This method cleverly uses a camera to capture the operator's natural hand movements. Using advanced image processing and motion capture algorithms, it analyzes the operator's hand's six degrees of freedom in three-dimensional space in real time. These motion commands are then accurately transmitted to a remote robotic manipulator, enabling the completion of various object manipulation tasks. However, the performance of the vision system is highly dependent on ambient lighting conditions. If the operating environment experiences insufficient lighting or drastic changes, the system's motion capture accuracy and stability will significantly decrease, making teleoperation tasks difficult to carry out smoothly. This, to a certain extent, limits its application in complex environments.
[0006] Handheld devices such as Dexmo and CyberGrasp utilize a mechanical linkage design that captures hand motion and drives a motor to generate a reverse torque, simulating the resistance generated in the operating environment. These systems primarily provide kinesthetic feedback, transmitting force information through muscle and joint perception, allowing the operator to sense macroscopic properties of an object, such as hardness and inertia. However, kinesthetic feedback technology focuses on macroscopic force simulation and has limitations in transmitting fine tactile information, making it difficult to meet the growing demand for precise teleoperation.
[0007] Researchers such as Trinitatova designed a glove-based tactile feedback device that combines hand tracking with electromagnetic actuators to achieve finger tactile feedback. Using tactile devices such as vibration motors and electromagnetic actuators, the device can simulate microscopic tactile information such as contact and texture directly on the operator's skin. Furthermore, they proposed a method for providing tactile feedback by indirectly applying laser radiation to an elastic medium, which can produce continuous tactile movement perception on the skin. However, existing tactile feedback research still primarily focuses on hand areas such as the fingers, with relatively insufficient attention paid to other important tactile perception areas such as the palm and wrist, limiting the comprehensiveness and immersiveness of overall tactile feedback.
[0008] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0009] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide a human-computer interaction system and method for robot training and data collection.
[0010] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect of the present invention, a human-computer interaction system for robot training and data collection comprises: The main device command system includes a motion acquisition subsystem and a tactile feedback subsystem. It is used to capture the operator's hand posture through wearable sensors to generate posture commands, and transmit force direction and intensity tactile feedback to the operator through a ring-distributed vibration motor array; A data processing system, connected to the main device command system, maps posture commands into robotic arm control signals, and processes force data fed back by the robotic end; The slave device response system includes a multi-degree-of-freedom robotic arm and force sensors, which are used to execute control instructions and collect environmental interaction force data and transmit it back to the data processing system; A robot self-learning module that trains a diffusion model based on a multimodal dataset collected synchronously with teleoperation, including time-aligned visual data, robot trajectory data, and force feedback data; Among them, the tactile feedback subsystem activates the corresponding angle motor to indicate the force direction through a single motor direction mapping mechanism, and linearly adjusts the vibration intensity encoding force size through PWM wave voltage; the robot self-learning module uses the denoising diffusion implicit model DDIM to accelerate reasoning and generate the robotic arm action sequence.
[0011] Furthermore, the motion collection subsystem includes: Three IMU sensors are deployed on the operator's upper arm, forearm, and wrist; A USB hub for synchronously receiving data from the three IMU sensors; The posture calculation module receives the data transmitted by the USB hub, fuses the IMU data with the human arm kinematic model, and outputs six-degree-of-freedom posture instructions to the data processing system.
[0012] Furthermore, the tactile feedback subsystem includes: A ring-shaped wristband device evenly distributes multiple linear resonant actuators (LRAs); Multiple PWM output circuits to independently control each LRA motor; Among them, each LRA motor is evenly spaced in a ring, and a direction mapping is established with the positive direction of the y-axis as 0° and increasing counterclockwise. When the force direction approaches the nearest motor, the motor is activated.
[0013] Furthermore, the data processing system is configured to perform: Cartesian impedance control strategy for driving the robotic arm based on posture commands; Decoding and transmitting force feedback data to the tactile feedback subsystem; Multimodal data synchronization: Based on the visual sampling time, the robot trajectory data is linearly interpolated and aligned.
[0014] Furthermore, the multimodal dataset includes: RGB images captured by the environment RGB camera to provide visual context information; The three-dimensional position (xyz), quaternion posture (wxyz) and three-dimensional force / torque data of the end of the robotic arm; Vision-trajectory-force synchronized data pairs aligned by timestamps.
[0015] Furthermore, the robot self-learning module includes: The data preprocessing unit extracts the visual features of the RGB image through the residual network and fuses them with the force / trajectory data through attention pooling to form a multi-dimensional feature vector; The DDIM strategy unit takes the feature vector as a conditional input, generates a future action sequence through a denoising diffusion process, and only executes the first action in the sequence.
[0016] Furthermore, the attention pooling performs weighted fusion of the visual features extracted by the residual network and the low-dimensional force / trajectory data, and outputs a feature vector after dimensionality reduction.
[0017] In a second aspect of the present invention, a human-machine interaction and robot training method using the system comprises: S1. Collect the operator's hand posture data through the main device command system and generate the robot arm posture command; S2, the slave device response system executes the posture command and collects the environmental force data, which is decoded by the data processing system and drives the tactile feedback subsystem; S3, synchronously collect time-aligned RGB images, robotic arm trajectory and force data to construct a multimodal dataset; S4. Preprocessing of the dataset: The visual data is extracted using a residual network and then fused with the force / trajectory data through attention pooling to form a multi-dimensional feature vector. S5. Input the DDIM model with the feature vector as the condition, generate the action sequence through denoising diffusion and execute the first action to realize the self-learning of the robotic arm.
[0018] Furthermore, the data synchronization in step S3 specifically includes: Search for robot trajectory data points based on camera frame timestamps; If the timestamps match, pair directly; If the timestamp is between two trajectory points, the position, attitude, and force data are linearly interpolated according to the time interval weight.
[0019] Furthermore, the DDIM reasoning process of step S5 includes: Receive current and historical state information to build a condition vector; Generate future action sequences from noise through iterative denoising; Only the state is recollected after the first action in the sequence is executed, triggering the next round of diffusion calculation.
[0020] The present invention has the following beneficial effects: The present invention proposes a human-machine interactive remote operation system and method with tactile feedback function. The motion acquisition subsystem and tactile feedback subsystem in the main device command system capture the operator's hand posture and transmit tactile feedback of force direction and intensity. The data processing system converts the posture command into a robotic arm control signal and processes the force data. The slave device response system executes the command and collects environmental interaction force data. The robot self-learning module uses a multimodal data set to train a diffusion model. Furthermore, in terms of the construction of the human-computer interaction system, the system introduces a wrist-worn vibration feedback interaction device that can accurately sense and transmit contact force information, enhancing the operator's environmental perception ability. The integrated wrist-worn tactile vibration feedback system can assist the operator in adjusting instructions in visually restricted scenarios. The human-computer interaction algorithm is also optimized to ensure real-time response of the robotic arm. In terms of robot autonomous learning, it can complete complex operation tasks without human intervention and has a high level of operation. At the same time, it is designed with a low-cost, portable vibration tactile feedback device and an imitation learning network that integrates force information, which can improve operation accuracy and efficiency and realize autonomous optimization of the robot's operation level. In addition, the low-cost motion acquisition subsystem can meet the robot's demand for data and provide support for imitation learning. Compared with the existing technology, it effectively overcomes the shortcomings of fuzzy force transmission of tactile feedback devices, high cost and poor portability of remote operation equipment, and imitation learning that ignores force information.
[0021] In summary, the present invention uses tactile feedback to enhance the accuracy of human-machine collaboration, low-cost motion acquisition to support large-scale data acquisition, and multimodal diffusion learning to enhance robot autonomy, forming a "perception-interaction-learning" closed loop, improving the robot's autonomous operation ability in complex scenarios, and providing efficient and reliable technical support for high-end scenarios such as remote surgery and precision industry.
[0022] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Schematic diagram of the human-computer interaction system for robot training and data collection according to an embodiment of the present invention.
[0024] Figure 2 Schematic diagram of a motion acquisition and tactile feedback device according to an embodiment of the present invention.
[0025] Figure 3 Schematic diagram of the motion collection subsystem according to an embodiment of the present invention.
[0026] Figure 4 Schematic diagram of the principle of a hand tactile feedback device according to an embodiment of the present invention.
[0027] Figure 5Schematic diagram of synchronous sampling and processing of multimodal data according to an embodiment of the present invention. DETAILED DESCRIPTION
[0028] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.
[0029] It should be noted that when an element is referred to as being "fixed to" or "disposed on" another element, it can be directly on the other element or indirectly on the other element. When an element is referred to as being "connected to" another element, it can be directly connected to the other element or indirectly connected to the other element. In addition, connection can be used for both fixing and coupling or communication.
[0030] It should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the embodiments of the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0031] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0032] See Figure 1An embodiment of the present invention provides a human-computer interaction system for robot training and data acquisition, including a master device instruction system, a data processing system, a slave device response system, and a robot self-learning module. The master device instruction system includes a motion acquisition subsystem and a tactile feedback subsystem, which are used to capture the operator's hand posture through wearable sensors to generate posture instructions, and transmit force direction and intensity tactile feedback to the operator through a ring-distributed vibration motor array. The data processing system is connected to the master device instruction system, maps the posture instructions into robotic arm control signals, and processes the force data fed back by the mechanical end. The slave device response system includes a multi-degree-of-freedom (such as seven degrees of freedom) robotic arm and a force sensor, which is used to execute control instructions and collect environmental interaction force data and transmit it back to the data processing system. The robot self-learning module trains a diffusion model based on a multi-modal data set synchronously collected by teleoperation, and the data set includes time-aligned visual data, robot trajectory data, and force feedback data. Among them, the tactile feedback subsystem of the main device command system activates the corresponding angle motor to indicate the force direction through a single motor direction mapping mechanism, and linearly adjusts the vibration intensity encoding force size through PWM wave voltage. The robot self-learning module uses the denoising diffusion implicit model DDIM to accelerate reasoning and generate a robotic arm action sequence.
[0033] This invention constructs a human-machine interaction system comprising a master device command system, a data processing system, a slave device response system, and a robot self-learning module. This multi-module collaboration enables robot training and data collection, and utilizes a multimodal data training diffusion model to enhance the robot's autonomy. A high-precision human-machine interaction closed loop is constructed by synergizing the tactile feedback subsystem with the motion acquisition subsystem. The DDIM model is trained based on a time-aligned multimodal dataset, enabling the robot to transition from teleoperation to autonomous decision-making. This improves data collection efficiency while utilizing the DDIM model to accelerate reasoning and enhance the robot's autonomous operation performance, significantly enhancing the robot's ability to operate autonomously in complex scenarios.
[0034] In some embodiments, the motion acquisition subsystem includes: three IMU sensors (such as three ten-axis IMU sensors), which are respectively deployed on the operator's upper arm, forearm and wrist; a USB hub, which synchronously receives data from the three IMU sensors; a posture calculation module, which receives data transmitted by the USB hub, fuses the IMU data with the human arm kinematic model, and outputs six-degree-of-freedom posture instructions to the data processing system.
[0035] In some embodiments, the tactile feedback subsystem includes: a ring-shaped wristband device, with multiple linear resonant actuators (LRAs) evenly distributed (e.g., 12 LRAs); a multi-channel PWM output circuit, which independently controls each LRA motor; wherein each LRA motor is evenly spaced in a ring (e.g., at intervals of 30°), and a direction mapping is established with the positive direction of the y-axis as 0° and increasing counterclockwise, and the motor is activated when the force direction approaches the nearest motor.
[0036] In some embodiments, the data processing system is configured to execute: a Cartesian impedance control strategy for driving a robotic arm based on posture instructions; decoding and transmitting force feedback data to a tactile feedback subsystem; and multimodal data synchronization: linear interpolation and alignment of robot trajectory data based on visual sampling time.
[0037] In some embodiments, the multimodal dataset includes: RGB images captured by an environmental RGB camera to provide visual context information; three-dimensional position (xyz), quaternion posture (wxyz) and three-dimensional force / torque data of the end of the robotic arm; and visual-trajectory-force synchronized data pairs aligned by timestamps.
[0038] In some embodiments, the robot self-learning module includes: a data preprocessing unit, which extracts visual features of RGB images through a residual network, and fuses them with force / trajectory data through attention pooling to form a multi-dimensional feature vector; a DDIM strategy unit, which uses the feature vector as conditional input, generates future action sequences through a denoising diffusion process, and only executes the first action in the sequence.
[0039] In some embodiments, the attention pooling performs weighted fusion of the visual features extracted by the residual network and the low-dimensional force / trajectory data, and outputs a feature vector after dimensionality reduction.
[0040] See Figure 1 The present invention also provides a method for human-computer interaction and robot training using the human-computer interaction system of any of the above embodiments, comprising the following steps: Step S1: collecting operator's hand posture data through the main device instruction system to generate a robot arm posture instruction; Step S2: The slave device response system executes the posture command and collects environmental force data, which is decoded by the data processing system and drives the tactile feedback subsystem; Step S3: synchronously collect time-aligned RGB images, robotic arm trajectory and force data to construct a multimodal dataset; Step S4: Preprocess the dataset: extract features from the visual data using a residual network and fuse them with the force / trajectory data through attention pooling to form a multi-dimensional feature vector. Step S5: input the denoising diffusion implicit model DDIM based on the feature vector, generate an action sequence through denoising diffusion and execute the first action to realize self-learning of the robotic arm.
[0041] In some embodiments, the data synchronization of step S3 specifically includes: searching for robot trajectory data points based on the camera frame timestamp; directly pairing if the timestamps match; and linearly interpolating the position, posture, and force data according to the time interval weight if the timestamp is between the two trajectory points.
[0042] In some embodiments, the DDIM reasoning process of step S5 includes: receiving current and historical state information to construct a condition vector; generating a future action sequence from noise through iterative denoising; and re-collecting the state after executing only the first action in the sequence to trigger the next round of diffusion calculation.
[0043] Compared with the prior art, the significant advantages of the embodiments of the present invention are reflected in the following aspects: (1) Existing tactile feedback devices suffer from fuzzy force transmission and low precision: the force feedback accuracy of traditional nonlinear motor vibration units is insufficient, and the vibration intensity and direction are difficult to linearly map to the actual contact force, resulting in the operator being unable to accurately perceive the angle and magnitude of the force at the robot's end; the equipment is bulky, and the motor-type device is difficult to integrate into a lightweight wristband, limiting the flexibility of human-machine collaboration. In contrast, the present invention uses a linear resonant actuator (LRA) as the minimum vibration unit of a low-cost portable tactile feedback system. Not only is the vibration sense more obvious than that of nonlinear motors at the hardware level, but it can also provide accurate feedback on the angle and magnitude of the force in the form of vibration. At the same time, the equipment is lightweight and easy to integrate, improving the flexibility of human-machine collaboration.
[0044] (2) Data collection bottlenecks caused by the high cost and poor portability of existing teleoperation equipment: Professional motion capture systems (such as optical tracking and exoskeletons) are expensive and difficult to deploy on a large scale, and cannot meet the needs of robot autonomous learning for massive demonstration data. In addition, heavy equipment causes operator fatigue, limits continuous working time, and reduces data collection efficiency. In contrast, the present invention uses inertial units to combine with the human posture capture system to form a lightweight, wearable, and low-cost motion acquisition subsystem. It uses a flexible controller and a robotic arm for joint control to achieve real-time human control of the robot, significantly reducing hardware costs, meeting the needs of robot autonomous learning for massive, high-quality, multimodal data sets, reducing operator fatigue, and improving data collection efficiency.
[0045] (3) Existing imitation learning ignores force information, resulting in a high error rate in autonomous operation: Pure visual imitation learning has a high failure rate in contact scenarios. Existing multimodal networks have difficulty in synchronously processing force-visual heterogeneous data. The time alignment error between the force sensor signal and the image frame is large, causing control delay. In contrast, this paper proposes an imitation learning network that integrates force information. It innovatively adopts a time-aligned multimodal dataset (vision-trajectory-force synchronous data), integrates force / visual heterogeneous features through residual networks and attention pooling, and generates future action sequences based on the denoising diffusion implicit model (DDIM). It only executes the first action and then replans in real time. This effectively solves the high error rate problem of pure visual imitation in contact scenarios and significantly improves the reliability of operations in complex environments.
[0046] Specific embodiments of the present invention are further described below.
[0047] like Figure 1 As shown in the figure, a human-computer interaction system for robot training and data collection is proposed. Its basic implementation scheme includes: (1) The human-machine interaction system realizes the integration and coordination of the teleoperation master device at the operator end and the slave device at the manipulator end, and accurately maps and smoothly follows the manipulator end gesture through the operator's hand gesture movement; (2) The operator remotely controls the robotic arm through the human-machine interaction system to perform related tasks. During the process, visual data, robot trajectory data, and force feedback data are collected synchronously to construct a multimodal dataset. (3) Preprocess the data set, process the high-dimensional data (image) and low-dimensional data (force data and trajectory data) separately, and generate training input data through feature extraction; (4) Train the training input data generated in the previous step, use the denoising diffusion implicit model (DDIM) to accelerate the reasoning process, conduct autonomous learning experiments on the robotic arm, and verify the autonomous learning ability based on teleoperation data.
[0048] The specific implementation of each component is described in detail below.
[0049] Design of human-computer interaction teleoperation system The human-machine interactive teleoperation system framework is functionally divided into three core modules: the master device command system, the data processing system, and the slave device response system. These three modules work together like a "brain" and "nervous system," forming a complete and efficient teleoperation control loop. This allows the operator to remotely control the robot in an immersive way and truly experience the force and tactile information of the remote environment.
[0050] Main device instruction system The main device command system, serving as the human-machine interface, is located on the left side of the architecture diagram. It is further subdivided into a motion acquisition subsystem and a tactile feedback device. The left arm is equipped with a motion acquisition device that provides input for the robot's posture control in three translational and three rotational directions. The right arm is equipped with a tactile feedback device. This vibration tactile feedback device provides tactile feedback and remote control start and stop functions.
[0051] Figure 2 A schematic diagram of a motion acquisition and tactile feedback device according to an embodiment of the present invention is shown.
[0052] The core of the motion acquisition subsystem lies in the posture calculation processor, which is responsible for accurately capturing the operator's hand posture, converting it into posture instructions, and passing it to the algorithm layer to drive the remote robotic arm.
[0053] The tactile feedback subsystem receives force feedback data from the remote robotic arm through the force feedback processor, decodes it into tactile stimulation signals, drives the tactile feedback actuator, and ultimately transmits the force sensation in the virtual environment to the operator realistically.
[0054] Motion acquisition subsystem In order to build a hand gesture capture system with multi-sensor collaborative work, the present invention adopts a three-sensor parallel acquisition scheme to accurately capture the operator's hand motion gestures and convert them into robot operation instructions.
[0055] The present invention selects a ten-axis attitude sensor as the core component for hand attitude capture. Compared with traditional six-axis or nine-axis inertial measurement units (IMUs), the ten-axis IMU has significant advantages in suppressing data drift and improving heading angle accuracy due to its redundant heading gyroscope technology. It also has a high data output frequency and can meet the stringent requirements of remote operation systems for data stability and real-time performance.
[0056] Specifically, posture capture is based on the kinematic model of the human arm, simplifying the arm into a motion chain consisting of key joints such as the shoulder, elbow, and wrist, and deploying an IMU sensor at key locations of the upper arm, forearm, and wrist to capture the motion posture information of different parts of the hand, thereby achieving more comprehensive and accurate hand motion tracking, and sensing its own angular velocity and acceleration in three-dimensional space in real time, and then calculating its own posture information, providing a data basis for the accurate capture of hand movements.
[0057] In order to solve the problem of synchronous acquisition and high-speed transmission of multi-sensor data, a high-speed USB hub was innovatively introduced as a data aggregation center. The three ten-axis attitude sensors are connected to the USB hub at high speed through independent USB interfaces. The hub acts as a data bridge to integrate and aggregate the high-bandwidth data streams from the three sensors, and transmit them to the computer at a baud rate of up to 921600bps and high-frequency data frequency. It is then transmitted to the host computer (computer side) at high speed in a unified data format for real-time processing and attitude solution. In terms of attitude solution, the fusion of IMU sensor data and kinematic models enables accurate solution of the operator's hand posture, providing accurate posture information for subsequent remote operation command generation. The schematic diagram of the motion acquisition subsystem is shown in the figure. Figure 3 shown.
[0058] Haptic feedback subsystem In the design of the haptic feedback solution, a linear resonant actuator (LRA) was selected as the core vibration element. Its linear vibration characteristics, fast response, and small size better meet the system requirements. To achieve wrist force feedback, the device is strapped to the wrist, and 12 LRA motors are distributed in a ring around the wristband. To control these 12 LRA motors, a multi-channel PWM output circuit is configured to precisely adjust the motor's vibration intensity.
[0059] The tactile feedback device uses single-motor direction indication feedback to clearly and concisely indicate the primary direction of force by activating a single motor. First, the physical positions of the 12 motors on the wristband are mapped to their directions within a two-dimensional plane, thereby establishing a one-to-one correspondence between motors and directions. The specific mapping rules are shown in the figure. In the wristband device, the 12 motors are evenly arranged in a circular pattern, with the angular spacing between adjacent motors set to approximately 30 degrees. The positive direction of the y-axis is defined as 0 degrees, and the angle value increases in a counterclockwise direction.
[0060] During actual feedback, when the end of the robot's arm is subjected to environmental forces, the force sensor collects this information in real time and converts it into a force vector containing direction and magnitude. This force direction information is then mapped to the preset wristband motor direction standard. The system determines which motor's direction most closely matches the force direction mapping and activates that motor for vibration, allowing the user to perceive the primary direction of the force.
[0061] To further convey force information, the single-motor direction indication feedback solution also incorporates a force intensity encoding mechanism. As force feedback increases, the voltage generated by the PWM wave output by the driver increases linearly, and the motor's vibration intensity and frequency increase accordingly. Conversely, they decrease. This allows users to sense not only the direction of the force but also its magnitude through changes in motor vibration intensity, providing comprehensive force feedback.
[0062] In this way, by activating a single motor to indicate the main direction of the force and encoding the magnitude of the force through the vibration intensity, a tactile feedback array is formed that can simulate force information of different directions and intensities. Figure 4 As shown, the principle of the hand tactile feedback device according to an embodiment of the present invention is demonstrated.
[0063] Data processing system The data processing system, serving as a bridge between the master device command system and the slave device system, is located at the center of the architecture diagram and serves as the "control brain" of the entire system. The Cartesian impedance controller is the core of the algorithm layer. It receives posture commands from the posture solver and drives the robot arm's movement accordingly, achieving precise mapping and tracking of the operator's posture to the robot's posture. The system adopts an impedance control strategy, which makes the robot arm compliant when interacting with the environment, improving the safety and naturalness of operation. At the same time, the Cartesian impedance controller is also responsible for transmitting force feedback data from the slave device response system to the force feedback processor, establishing a data path for tactile feedback and ensuring the effective return of force and tactile information.
[0064] Slave device responds to system The slave system, serving as the teleoperation execution terminal, is located on the right side of the architecture diagram. The slave response system serves as both an "executor" of control commands and a "sensor" of environmental information. It receives control commands from the algorithm layer, drives the robotic arm to perform the corresponding actions, and collects real-time force feedback data generated by the interaction between the robotic arm and the environment. This data is then transmitted to the force feedback processor, forming a complete force feedback closed loop. The robotic arm utilizes a high-performance seven-degree-of-freedom collaborative robot. Its high-precision and sensitive force sensing capabilities enable it to more closely mimic the structure and movement of the human arm, providing reliable hardware support for high-fidelity tactile feedback. In summary, these three modules work together to create a closed-loop force and tactile feedback teleoperation system, providing the operator with an immersive, high-precision remote operation experience.
[0065] Dataset Generation Based on Human-Computer Interaction Teleoperation System Relying on the previously established teleoperation human-machine interaction system, the operator can remotely and precisely control the robotic arm, collecting teleoperation data during repeated operations. Vision data, robot trajectory data, and force feedback data are simultaneously collected to construct a multimodal dataset.
[0066] (1) Visual data acquisition primarily relies on an RGB camera installed in the robot's working environment. The camera continuously captures image information of the operating scene and stores the image data in RGB format, with a resolution of 128x128x3 per frame. RGB image data can provide the robot with rich visual context information, such as scene layout, object position, and object posture.
[0067] (2) The robot trajectory data acquisition covers the position, posture, and force feedback information of the end effector of the manipulator, and comprehensively records the robot's motion state and interaction with the environment during teleoperation. The position data is represented by three-dimensional Cartesian coordinates (xyz), which accurately describes the position of the end effector in space; the posture data is represented by quaternions (wxyz), which effectively avoids the singularity problem that may exist in the Euler angle representation and can accurately describe the orientation of the end effector in three-dimensional space.
[0068] (3) Force feedback data includes the magnitude of three-dimensional force and torque information. These data directly reflect the interaction force when the end of the robotic arm contacts the environment. They are important inputs to the tactile feedback system and also provide key information for the robot to learn force control strategies and compliant operations.
[0069] (4) In order to achieve effective fusion of multimodal data, data synchronization sampling processing is required to solve the problem of inconsistent sampling frequencies between visual data and robot trajectory data. Taking the camera sampling time as the reference timeline, for each frame of camera image data, the timestamp of its imaging moment is used as the reference to search on the time axis of the robot trajectory data. If there is a robot trajectory data point that is completely consistent with the camera frame timestamp, then the trajectory data point is directly paired with the current camera frame data as the synchronized sampling data. If the timestamp of the camera frame is between the timestamps of the two sampling points before and after the robot trajectory data, the linear weighted average interpolation method is used, with the time interval between the timestamps of the two trajectory points before and after and the camera frame timestamp as the weight, and the position, posture and force feedback data of the two trajectory points before and after are weighted and averaged. The data obtained by interpolation calculation is used as the robot trajectory data corresponding to the current camera frame moment. By combining this method of timeline synchronization sampling with linear interpolation, the visual data and robot trajectory data of different sampling frequencies are effectively and accurately aligned in time, and a time-synchronized multimodal dataset is constructed.
[0070] Figure 5 A schematic diagram of synchronous sampling and processing of multimodal data according to an embodiment of the present invention is shown.
[0071] Dataset preprocessing Data is processed before training to extract features from multimodal input, providing high-quality input for subsequent training stages. Multimodal information consists of two major types of data: high-dimensional image data: environmental visual data (images), labeled "RGB information," typically three-channel color images (dimensions [3, 128, 128], i.e., RGB three channels, 128×128 resolution); and low-dimensional data, such as "force data and trajectory data," which refers to the forces and torques generated by the robot's interaction with the environment, as well as the robot's own kinematic information, which are numerical structured data.
[0072] The visual data of environmental information is processed through a residual network for deep convolutional feature extraction. The residual network can extract high-level semantic features from a [3,128,128] image and output a feature map with dimensions of [512,4,4]. Here, "512" represents the number of feature channels, and "4×4" represents the spatial resolution of the feature map. Through the residual network's multi-layer convolution and residual connection mechanism, the originally high-dimensional and complex image information is refined into more compact deep features, paving the way for multimodal fusion.
[0073] The visual features [512, 4, 4] output by the residual network are fused with the low-dimensional data (force, trajectory, etc.) from another branch in the "attention pooling" module. Attention pooling combines the attention mechanism with conventional pooling operations, enabling the model to adaptively assign corresponding weights to different modalities or different regions, thereby highlighting key information and compressing feature dimensions. The final output is a
[512] -dimensional vector, which means that in the process of multimodal information fusion and dimensionality reduction, attention pooling provides a more refined feature representation for subsequent training. The
[512] -dimensional feature vector output by attention pooling points to the "training input", representing the multimodal features that can be directly used for model training after preprocessing. This vector combines visual information with motion (force, trajectory) information, retaining key elements while effectively compressing them through deep learning feature extraction and attention mechanism.
[0074] Robot self-learning training based on human-computer interaction data By introducing denoising diffusion probabilistic models (DDPMs) and the robot's learning strategy, the robot's learning steps based on the diffusion model are combined with the dataset construction and data processing in the above steps to implement model training and inference, thus building a complete robotic arm self-learning process.
[0075] The core idea of the Diffusion Policy is to prevent the robot from guessing single actions. Instead, it leverages its denoising capabilities to generate a short sequence of optimal future actions. The Diffusion Policy workflow is summarized as follows: 1. Look: Collect the current state and the state of a short period of time in the past.
[0076] 2. Imagine: Compress this state information into a "conditional vector" and feed it to the diffusion model (i.e., the preprocessed feature vector). Based on this "condition," the diffusion model starts from a cloud of noise and gradually removes the noise (reverse diffusion), imagining a short future action sequence that is the smoothest and most suitable for the current task scenario.
[0077] 3. Do: Take out only the first action in the imagined action sequence and execute it.
[0078] 4. Watch again after the action is completed: The environment changes due to the execution of the action, and the robot obtains a new state.
[0079] 5. Refresh ideas (re-planning): Reassemble the "state window" with the latest state information, start the next round of diffusion model calculation from scratch, and generate a new action sequence that adapts to the latest situation.
[0080] 6. Go back to step 3 and repeat.
[0081] In actual autonomous robot operation, the robot's operating speed often depends on the inference speed of the network model. This paper uses a denoising diffusion implicit model (DDIM) to accelerate the inference process. The core advantage of DDIM lies in its decoupling of the number of denoising iterations between the training and inference phases. While traditional diffusion model strategies require multiple iterations during training to ensure performance, the DDIM approach significantly reduces the number of iterations during inference, significantly improving inference speed. This shift in approach enables diffusion model-based robotic strategies to meet the response speed requirements of real-time closed-loop control, paving the way for the deployment of diffusion models in robotic applications with high real-time requirements.
[0082] Important features and innovative advantages of the embodiments of the present invention: (1) In the human-machine interaction system, a wrist-worn vibration feedback interaction device is introduced. This device can accurately sense the magnitude and direction of the contact force between the end of the robotic arm and the environment, and transmit tactile information to the operator in real time in the form of intuitive vibrations. Compared with the human-machine interaction mode that relies solely on visual information, the introduction of the tactile feedback device significantly enhances the operator's ability to perceive the environment, enabling them to more effectively deal with the uncertainty and blind spots in complex operation scenarios.
[0083] (2) The wristband tactile vibration feedback system is integrated into the human-computer interaction system to form a complete tactile feedback system for human-computer interaction system design and human-computer interaction experiments. This system focuses on solving operational difficulties in some observable environments. Tactile feedback assists operators in quickly and accurately adjusting human-computer interaction instructions in scenes with limited vision or many obstructions. To achieve the real-time nature of the system and smooth operation, the human-computer interaction software algorithm is used to ensure that the robotic arm can respond to the operator's command input in real time and can quickly adapt to complex operational tasks.
[0084] (3) The designed human-computer interaction system is applied to robot imitation learning training, and it successfully completes complex operation tasks independently without human intervention, demonstrating an operation level close to or even exceeding that of human operators.
[0085] In summary, the present invention designs a complete human-machine interactive teleoperation system with a super-redundant robotic arm operating in a complex and narrow space, so that the system has real-time, accuracy, safety, and user-friendliness. In terms of the construction of the human-machine interactive system, an innovative teleoperation scheme based on tactile feedback is proposed, and a low-cost, portable vibration tactile feedback device is designed to transmit the contact force information of the robot end in real time and effectively, thereby significantly improving the accuracy and efficiency of the operation. In terms of the research on autonomous learning algorithms, the present invention further proposes an imitation learning network that combines force information, effectively integrates the force and visual data collected during the human-machine interaction process, and applies it to drive the autonomous learning of the robot, ultimately achieving autonomous optimization and improvement of the robot's operation level. The significant significance and application value of the embodiments of the present invention are reflected in the following aspects: (1) Value of wristband tactile feedback subsystem: This invention uses a linear resonant driver as the minimum vibration unit to create a low-cost portable tactile feedback system. Not only is the vibration more obvious than that of a nonlinear motor at the hardware level, but it can also provide accurate feedback on the angle and magnitude of force in the form of vibration.
[0086] (2) Human-machine interaction system with a lightweight, wearable, and low-cost motion acquisition subsystem: This invention utilizes a low-cost inertial unit combination, known as a human posture capture system, and uses a compliant controller to jointly control the robot arm, enabling real-time human control of the robot. Furthermore, the significant cost reduction effectively addresses the growing demand for high-quality data for autonomous robot operation. Combined with the tactile feedback subsystem, it provides critical data support for subsequent robot imitation learning, making it an important component of the human-machine collaborative system.
[0087] (3) The significance of imitation learning that integrates force information: The proposed imitation learning network that integrates force information closely revolves around the goal of autonomous robot operation and makes full use of the high-quality data collected by the tactile feedback human-computer interaction system. It verifies the imitation learning method as an effective way to achieve autonomous robot operation and its advantages in integrating multimodal force information, providing strong support for autonomous robot learning.
[0088] The above description further details the present invention in conjunction with specific / preferred embodiments, and the specific implementation of the present invention should not be construed as being limited to these descriptions. Persons skilled in the art will appreciate that, without departing from the spirit of the present invention, they may make various substitutions or modifications to the described embodiments, and these substitutions or modifications should be considered to fall within the scope of protection of the present invention. Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "preferred embodiments," "examples," "specific examples," or "some examples" indicates that the specific features, structures, materials, or characteristics described in conjunction with such embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples. Persons skilled in the art may combine and assemble the different embodiments or examples described in this specification, as well as features of different embodiments or examples, without conflicting opinions. Although the embodiments of the present invention and their advantages have been described in detail, it should be understood that various changes, substitutions, and modifications may be made within the scope of protection of the patent application.
Claims
1. A human-computer interaction system for robot training and data collection, characterized in that: include: The main device command system includes a motion acquisition subsystem and a tactile feedback subsystem. It is used to capture the operator's hand posture through wearable sensors to generate posture commands, and transmit force direction and intensity tactile feedback to the operator through a ring-distributed vibration motor array; A data processing system, connected to the main device command system, maps posture commands into robotic arm control signals, and processes force data fed back by the robotic end; The slave device response system includes a multi-degree-of-freedom robotic arm and force sensors, which are used to execute control instructions and collect environmental interaction force data and transmit it back to the data processing system; A robot self-learning module that trains a diffusion model based on a multimodal dataset collected synchronously with teleoperation, including time-aligned visual data, robot trajectory data, and force feedback data; Among them, the tactile feedback subsystem activates the corresponding angle motor to indicate the force direction through a single motor direction mapping mechanism, and linearly adjusts the vibration intensity encoding force size through PWM wave voltage; the robot self-learning module uses the denoising diffusion implicit model DDIM to accelerate reasoning and generate the robotic arm action sequence.
2. The system according to claim 1, wherein: The motion acquisition subsystem includes: Three IMU sensors are deployed on the operator's upper arm, forearm, and wrist; A USB hub for synchronously receiving data from the three IMU sensors; The posture calculation module receives the data transmitted by the USB hub, fuses the IMU data with the human arm kinematic model, and outputs six-degree-of-freedom posture instructions to the data processing system.
3. The system according to claim 1 or 2, characterized in that The tactile feedback subsystem includes: A ring-shaped wristband device evenly distributes multiple linear resonant actuators (LRAs); Multiple PWM output circuits to independently control each LRA motor; Among them, each LRA motor is evenly spaced in a ring, and a direction mapping is established with the positive direction of the y-axis as 0° and increasing counterclockwise. When the force direction approaches the nearest motor, the motor is activated.
4. The system according to any one of claims 1 to 2, characterized in that The data processing system is configured to perform: Cartesian impedance control strategy for driving the robotic arm based on posture commands; Decoding and transmitting force feedback data to the tactile feedback subsystem; Multimodal data synchronization: Based on the visual sampling time, the robot trajectory data is linearly interpolated and aligned.
5. The system according to any one of claims 1 to 2, characterized in that: The multimodal dataset includes: RGB images captured by the environment RGB camera to provide visual context information; The three-dimensional position (xyz), quaternion posture (wxyz) and three-dimensional force / torque data of the end of the robotic arm; Vision-trajectory-force synchronized data pairs aligned by timestamps.
6. The system according to any one of claims 1 to 2, characterized in that: The robot self-learning module includes: The data preprocessing unit extracts the visual features of the RGB image through the residual network and fuses them with the force / trajectory data through attention pooling to form a multi-dimensional feature vector; The DDIM strategy unit takes the feature vector as a conditional input, generates a future action sequence through a denoising diffusion process, and only executes the first action in the sequence.
7. The system according to claim 6, characterized in that The attention pooling performs weighted fusion of the visual features extracted by the residual network and the low-dimensional force / trajectory data, and outputs a feature vector after dimensionality reduction.
8. A method for human-computer interaction and robot training using the system according to any one of claims 1 to 7, characterized in that: include: S1. Collect the operator's hand posture data through the main device command system and generate the robot arm posture command; S2, the slave device response system executes the posture command and collects the environmental force data, which is decoded by the data processing system and drives the tactile feedback subsystem; S3, synchronously collect time-aligned RGB images, robotic arm trajectory and force data to construct a multimodal dataset; S4. Preprocessing of the dataset: The visual data is extracted using a residual network and then fused with the force / trajectory data through attention pooling to form a multi-dimensional feature vector. S5. Input the denoising diffusion implicit model DDIM with the feature vector as the condition, generate the action sequence through denoising diffusion and execute the first action to realize the self-learning of the robotic arm.
9. The method according to claim 8, characterized in that The data synchronization in step S3 specifically includes: Search for robot trajectory data points based on camera frame timestamps; If the timestamps match, pair directly; If the timestamp is between two trajectory points, the position, attitude, and force data are linearly interpolated according to the time interval weight.
10. The method according to claim 8 or 9, characterized in that The DDIM reasoning process of step S5 includes: Receive current and historical state information to build a condition vector; Generate future action sequences from noise through iterative denoising; Only the state is recollected after the first action in the sequence is executed, triggering the next round of diffusion calculation.
Citation Information
Patent Citations
Robot control method and system based on multi-source information perception and electronic equipment
CN116966054A
Wearable control device and system for large mimicry bionic robot
CN118024229A
Teleoperating of robots with tasks by mapping to human operator pose
US10919152B1
Robotic arm control method and device, and human-machine cooperation model training method
WO2022088593A1
Cited By
Interactive mechanical interaction deduction method and system applied to exhibition hall
CN120891932A
Remote mechanical arm precise control system based on circular polarization signal sensing
CN120921410A
Self-adaptive massage equipment based on multi-mode perception
CN121401107A
Human body motion capture system based on multi-mode synchronous acquisition
CN121512506A