Humanoid robot teleoperation grabbing control method
By constructing a multimodal temporal coding sequence and tactile compliant control, combined with a whole-body coordination control model, the problems of insufficient prediction of action intent and unstable redundant solution of the teleoperation system in the agricultural chemical environment were solved, and stable grasping and compliant operation of agricultural precision glassware were realized.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU JINZHI HUMANOID ROBOT TECHNOLOGY CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-19
AI Technical Summary
Existing teleoperation systems suffer from problems such as insufficient prediction of action intent, insufficient tactile feedback, and unstable redundant solutions in agricultural chemical scenarios, making it difficult to meet the high safety and high compliance requirements for the operation of precision glassware in agricultural chemicals.
By collecting operator data to construct a multimodal temporal coding sequence for predicting action intent, and combining it with tactile array data for compliant control, a full-body coordination control model with zero-spatial projection is adopted, and the control parameters are adaptively adjusted by estimating the attributes of the grasped object online.
It achieves continuity and stability in robot grasping under high latency conditions, reduces the risk of vessel breakage, avoids posture jitter, and improves adaptability to various vessel types.
Smart Images

Figure CN122058346A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot teleoperation control technology, specifically to a teleoperation grasping control method for a humanoid robot. Background Technology
[0002] In the agrochemical industry, the production, dispensing, and quality testing of pesticide formulations commonly require the measurement, titration, and mixing of strong acids, strong alkalis, and highly corrosive reagents, necessitating frequent handling of precision glassware such as beakers, measuring cylinders, and burettes. These tasks demand not only high-precision control but also require minute operations within confined spaces such as fume hoods and narrow reagent racks, making them far more challenging than typical industrial handling scenarios.
[0003] In practical operations, the following technical challenges exist. Firstly, operators are repeatedly exposed to corrosive liquids in confined spaces, easily leading to skin burns, eye damage, and inhalation hazards. Secondly, glassware requires extremely high stability and force control compliance during handling, moving, tilting, and placement; even slight angular deviations can cause breakage or liquid splashing. Therefore, given the increasing demand for robot-assisted or human-replacement operations, existing remote operation methods are insufficient to meet the precision operational requirements of agricultural chemical scenarios.
[0004] Currently, most mainstream teleoperation systems adopt the direct position mapping method, that is... ,in For the operator's hand posture. For the mapping matrix, This is due to network latency. This simple mapping ignores the time lead of the operation intention and the prediction of the action trend. In latency scenarios, it is easy for the trajectory to diverge, causing the robot to shake when grasping and pouring liquid.
[0005] Furthermore, agrochemical glassware is extremely sensitive to the force applied during gripping, while traditional tactile control is mostly based on linear impedance models. The model struggles to handle the nonlinear contact requirements of fragile yet stable glassware, leading to overshoot during force control switching and further increasing the risk of breakage.
[0006] Meanwhile, the highly constrained operations of robots in confined spaces can lead to joint redundancy and instability in the solution. Commonly used pseudo-inverse methods... Near singular configurations, a velocity amplification effect occurs, amplifying minute adjustments at the end of the joint into violent joint movements, making it difficult to achieve the smooth micro-manipulation of a human hand.
[0007] In summary, existing remote telecontrol systems generally have the following characteristics: The lack of motion intention prediction leads to delayed amplification of errors; tactile feedback is insufficient to adjust for minute force changes; redundant solutions are unstable and cause posture jitter; and it is difficult to meet the high safety and high compliance requirements of operating precision glassware in agrochemicals.
[0008] Therefore, there is an urgent need for an algorithmic framework that integrates motion intention prediction and nonlinear tactile feedback compensation, enabling wheeled humanoid robots to achieve stable grasping, precise liquid pouring, and compliant contact operations in high-risk agricultural environments. Summary of the Invention
[0009] The present invention aims to solve at least one of the technical problems existing in the prior art, and to provide a method for remotely operating and grasping control of a humanoid robot.
[0010] To achieve the above objectives, the present invention provides a teleoperated grasping control method for a humanoid robot, comprising: The operator's wrist speed data, joint angular acceleration data, and pose trajectory data are collected, and a multimodal temporal coding sequence is constructed by temporal splicing. Feature extraction and action intention prediction are performed on the multimodal temporal coding sequence, and a predicted action command is output. The robot collects pressure distribution data from the tactile array of its hand, converts the pressure distribution data into contact force data according to the calibration model, calculates the tactile deviation between the contact force data and the target contact force, and processes the tactile deviation based on a compliant control law to generate a torque compensation command. A full-body coordinated control model is constructed based on null-space projection. An optimization objective function is defined, encompassing center of gravity stability, trunk uprightness, and collision avoidance. Upper limb task commands are fused with the redundant degrees of freedom optimization objective corresponding to the defined objective function, outputting full-body joint velocity commands. Based on the pressure distribution data and end force feedback data, the mass parameters and elastic modulus parameters of the grasped object are calculated, and the grasping force threshold and force control sensitivity parameters are adaptively adjusted according to the mass parameters and elastic modulus parameters.
[0011] Furthermore, the construction formula for the multimodal temporal coding sequence is as follows: ; in, From time At the time The multimodal temporal coding sequence, The length of the historical time window. for The wrist speed data at that moment, for The joint angular acceleration data at time [time], for The pose trajectory data at the given time.
[0012] Furthermore, feature extraction and action intent prediction are performed on the multimodal temporal encoded sequence using a spatiotemporal attention encoder, wherein the calculation formula for action intent prediction is: ; in, For the predicted future The predicted action instruction at the specified time, Let be the mapping function of the spatiotemporal attention encoder. To predict the number of time-domain steps, The value ranges from 1 to the preset maximum prediction step number.
[0013] Furthermore, the formula for calculating the contact force data is as follows: ; The formula for calculating the tactile deviation is: ; in, for The contact force data at that moment, The mapping function of the calibration model. for The pressure distribution data at time t, for The tactile deviation at that moment, The target contact force is denoted as .
[0014] Furthermore, the calculation formula for the torque compensation command is as follows: ; in, for The torque compensation command at that moment, It is a proportional gain matrix. The differential gain matrix is... The rate of change of the tactile deviation is denoted as .
[0015] Furthermore, the compliant control law also includes adaptive adjustment of the contact state: When an initial contact event is detected, the values of the proportional gain matrix and the differential gain matrix are reduced to form a buffered response; When a shift in the distribution center of the pressure distribution data is detected, it is determined to be a slippage trend, and the value of the proportional gain matrix is increased to improve the gripping force.
[0016] Furthermore, the calculation formula for the whole-body joint velocity command is as follows: ; in, This refers to the velocity command for all joints in the body. This is the pseudo-inverse of the Jacobian matrix. The end-effector velocity corresponding to the upper limb task command. It is the identity matrix. For Jacobian matrices, The gradient of the optimization objective function is... This is the current joint angle vector.
[0017] Furthermore, the optimization objective function includes: The objective for center-of-gravity stability is to ensure that the robot's center-of-gravity projection falls within the supporting polygon. For targets with upright torso posture, the deviation of torso posture angle is constrained to be less than a preset posture threshold. Collision avoidance targets are achieved by ensuring that the distance between the arms and torso is greater than a preset safe distance threshold.
[0018] Furthermore, the mass parameters and elastic modulus parameters of the grasped object are calculated using the following formulas: ; in, For the estimated quality parameters, For the estimated elastic modulus parameter, Let be the parameter vector to be estimated. For parameter-based The simulated contact force model output.
[0019] Furthermore, the adaptive adjustment includes: When the mass parameter is greater than the preset mass threshold, the upper limit of the gripping force threshold is increased and the angular velocity limit of the tilting operation is reduced; When the elastic modulus parameter is less than the preset elastic threshold, the force control sensitivity parameter is reduced and the low-force gripping mode is enabled.
[0020] The beneficial effects of this invention are as follows: This invention constructs a multimodal temporal coding sequence and uses a spatiotemporal attention encoder for action intent prediction. It can infer the operator's intent in advance with a network latency of 80 to 150 milliseconds, automatically fill in the missing action trajectory, and enable the robot to maintain continuity and stability in the grasping path, significantly reducing the latency perception of teleoperation.
[0021] This invention utilizes a compliant control law based on a tactile array to rapidly reduce output stiffness upon contact with the glassware, creating a buffer zone to prevent excessive force from causing breakage. Simultaneously, it can detect slippage trends and automatically increase gripping force, achieving a dual safety strategy of preventing breakage and drop.
[0022] This invention employs a whole-body coordination control model with zero-space projection, which comprehensively considers multiple optimization objectives such as center of gravity stability, torso uprightness, and collision avoidance. Under high-constraint operation scenarios, it avoids posture jitter caused by instability in joint redundancy solutions, and achieves compliant micro-operations similar to human hands.
[0023] This invention estimates the physical property parameters of the object to be grasped online and adaptively adjusts the control parameters, enabling it to handle containers of different weights and materials, thus improving the system's adaptability to various types of containers. Attached Figure Description
[0024] Figure 1 This is a flowchart of the teleoperation grasping control method for a humanoid robot according to the present invention; Figure 2 This is a schematic diagram of the module composition of the remote operation control system of the present invention; Figure 3 This is a comparison diagram of the teleoperation trajectory of the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and beneficial effects of this application clearer, the following detailed description, in conjunction with the accompanying drawings and specific embodiments, further illustrates this application. It should be understood that the specific embodiments described in this specification are merely for explaining this application and are not intended to limit it.
[0026] The humanoid robot teleoperation grasping control method of the present invention is based on a teleoperation control system.
[0027] See Figure 2 The teleoperation control system includes a motion acquisition module, an intent prediction module, a tactile perception module, a compliant force control module, a whole-body coordination module, and a parameter estimation module. The motion acquisition module collects real-time data on the operator's wrist speed, joint angular acceleration, and pose trajectory using VR devices and a motion capture system. The intent prediction module constructs a multimodal temporal coding sequence and predicts the action intent using a spatiotemporal attention encoder. The tactile perception module collects pressure distribution data from the robot's hand tactile array and determines the contact state. The compliant force control module generates torque compensation commands based on tactile deviations and compliant control laws. The whole-body coordination module coordinates upper limb tasks with overall stability based on zero-space projection. The parameter estimation module estimates the physical property parameters of the grasped object online and adaptively adjusts the control parameters.
[0028] Example 1 This embodiment uses the glassware grasping operation in an agricultural chemical laboratory as an application scenario. In this scenario, the robot needs to grasp, move, and place precision glassware such as beakers, measuring cylinders, and burettes in remote operation mode.
[0029] The robotic platform is a wheeled humanoid robot, 165 cm tall, with 7 degrees of freedom in each arm and a three-finger dexterity hand design in each hand. The chassis uses a differential drive structure, with a maximum movement speed of 1.2 meters per second. The sensor configuration includes: a dual-finger tactile array sensor, with 6 tactile sensing units in each hand, measuring a range of 0 to 30 Newtons and a resolution of 0.05 Newtons; a wrist six-dimensional force sensor, measuring force and torque in three axes; a head RGB-D camera with a resolution of 640 x 480 pixels and a frame rate of 30 frames per second; and joint encoders providing position, velocity, and torque feedback.
[0030] The remote control equipment includes: a VR controller for collecting the operator's wrist six-dimensional pose and finger opening and closing status; a motion capture system that uses optical marker tracking to collect the operator's upper limb joint movement data; and an operator station monitor that displays the robot's view and tactile feedback information in real time.
[0031] See Figure 1 The control method in this embodiment includes the following steps: Step S1: Operator motion data acquisition and preprocessing.
[0032] The VR controllers and motion capture system collect the operator's motion data in real time. The VR controllers collect six-dimensional pose data of the operator's wrist, including three-dimensional position coordinates and three-dimensional posture angles. The motion capture system collects joint angle data of the operator's upper arm and forearm. The sampling frequency is configured at 120 Hz to ensure sufficient temporal resolution.
[0033] The acquired raw data is preprocessed. Position data undergoes a low-pass filter to remove high-frequency noise, with the filter cutoff frequency set to 15 Hz. The wrist velocity vector is calculated based on two consecutive frames of position data. The calculation formula is: ,in for The position of the wrist at all times, The sampling period is [number]. Joint angular acceleration is calculated based on angular velocity data from three consecutive frames. The calculation formula is: The wrist pose data is represented as a homogeneous transformation matrix. It contains rotation matrices and translation vectors.
[0034] The preprocessed wrist velocity data, joint angular acceleration data, and pose trajectory data were time-series concatenated to construct a multimodal temporal coding sequence. The historical time window length L was set to 30 sampling points, corresponding to a time span of 250 milliseconds. The formula for constructing the multimodal temporal coding sequence is as follows: .
[0035] Step S2: Predicting action intent.
[0036] The intent prediction module uses a spatiotemporal attention encoder to process multimodal temporally encoded sequences. The spatiotemporal attention encoder consists of a feature embedding layer, multiple self-attention encoding layers, and a prediction output layer.
[0037] The feature embedding layer maps data from different modalities to a unified hidden space. Wrist velocity data is mapped to a 256-dimensional hidden vector through a linear transformation. Joint angular acceleration data is mapped to a 256-dimensional hidden vector through a linear transformation. Pose trajectory data is mapped to a 256-dimensional hidden vector through flattening and a linear transformation. The three hidden vectors are concatenated to obtain a 768-dimensional fused feature vector.
[0038] The multi-layer self-attention coding layer employs a 4-layer Transformer encoder structure, with each layer containing 8 attention heads. The self-attention mechanism can capture long-range dependencies and cross-modal correlations in time-series data. Positional encoding uses a sinusoidal positional encoding method, adding positional information to each time step in the sequence.
[0039] The prediction output layer employs a multilayer perceptron structure, mapping the encoder output to predicted action commands within a preset future time domain. The prediction time domain is set to 6 time steps, corresponding to a prediction range of 50 milliseconds. The formula for calculating action intent prediction is as follows: ,in The value range is from 1 to 6.
[0040] In this embodiment, the spatiotemporal attention encoder can infer the operator's intention in advance even with a network latency of 80 to 150 milliseconds. When the prediction confidence is higher than a preset confidence threshold, the system directly issues a predicted action command to the robot for execution. The confidence threshold is set to 0.85. When the prediction confidence is lower than the confidence threshold, the system triggers an operator confirmation mechanism, prompting the operator to confirm again via vibration of the VR controller.
[0041] Step S3: Acquisition of tactile array data and calculation of contact force.
[0042] The tactile sensing module collects pressure distribution data from the robot's hand tactile array. Two tactile sensing units are configured at each fingertip, for a total of six tactile sensing units per hand. The tactile sensors are configured with a sampling frequency of 200 Hz to ensure real-time response to contact dynamics.
[0043] Pressure distribution data output by the tactile array The vector is 6-dimensional, with each component representing the pressure measurement value of the corresponding tactile sensing unit. The pressure distribution data is converted into contact force data based on a pre-calibrated model. The calibration model is established using a polynomial fitting method, and a six-dimensional force sensor is used as a reference during the calibration process. The formula for calculating the contact force data is as follows: .
[0044] Set the target contact force according to the current task requirements. For the glassware grasping task, the target contact force is set to 8 Newtons to ensure sufficient grasping stability while avoiding excessive force that could break the glassware. The tactile deviation between the contact force data and the target contact force is calculated using the following formula: .
[0045] Step S4: Compliance control and torque compensation.
[0046] The compliance control module generates torque compensation commands based on tactile deviation and compliance control laws. The formula for calculating the torque compensation commands is as follows: .
[0047] Scale-up matrix and differential gain matrix The system adaptively adjusts based on the contact state. It determines the contact state by monitoring changes in the pressure distribution of the tactile array.
[0048] Contact detection: When the pressure measurement value of any sensing unit in the tactile array rises from zero to above the contact threshold, it is determined as an initial contact event. The contact threshold is set to 0.5 N. After the initial contact event occurs, the system enters a buffer response mode, adjusting the proportional gain matrix... The value is reduced to 40% of the normal value, and the differential gain matrix is reduced. The value decreased to 50% of the normal value. After the buffered response mode lasted for 200 milliseconds, the gain matrix gradually returned to the normal value.
[0049] Slip detection: The system continuously monitors the location of the center of pressure distribution data. The formula for calculating the distribution center is as follows: ,in For the first Pressure measurement values of each sensing unit For the first The spatial coordinates of each sensing unit are determined. A slippage trend is identified when the moving speed of the distribution center exceeds a slippage threshold. The slippage threshold is set to 5 millimeters per second. Upon detecting a slippage trend, the system increases the proportional gain matrix. The value is increased to 150% of the normal value to improve gripping force and prevent the container from falling.
[0050] Overvoltage protection: When the pressure measurement value of any sensing unit in the haptic array exceeds the overvoltage threshold, the system immediately executes overvoltage protection. The overvoltage threshold is set to 25 N. Overvoltage protection actions include: reducing the proportional gain matrix. The overvoltage value is reduced to 20% of the normal value; if the overvoltage duration exceeds 100 milliseconds, a rapid release action is performed and an alarm is triggered.
[0051] In this embodiment, the proportional gain matrix in normal mode Set as a diagonal matrix, with diagonal elements valued at 50 Nm per Newton. Differential gain matrix. Set it as a diagonal matrix with diagonal elements of 5 Nm / s.
[0052] Step S5: Full-body coordination and control.
[0053] The whole-body coordination module constructs a whole-body coordination control model based on zero-space projection. The wheeled humanoid robot has redundant degrees of freedom: 7 degrees of freedom for each arm, 3 degrees of freedom for the chassis, and 1 degree of freedom for the waist, for a total of 18 controllable degrees of freedom. The end-effector pose control for the grasping task only requires 6 degrees of freedom, thus there are 12 redundant degrees of freedom that can be used to optimize other objectives.
[0054] The formula for calculating the velocity command of all joints is as follows: The first item To ensure the end effector moves at the desired speed, the second term Optimize the objective function using redundant degrees of freedom .
[0055] Optimize objective function Taking into account multiple optimization objectives: Center of gravity stability target : Calculate the projected position of the robot's center of mass on the horizontal plane The projection is constrained to fall within the supporting polygon formed by the chassis wheels. As the centroid projection approaches the boundary of the supporting polygon, the gradient... Guide joint movement to move the center of mass projection toward the center of the supporting polygon. The weighting coefficient for the center of mass stability objective is set to 1.0.
[0056] Torso upright target Calculate the torso's attitude angle deviation relative to the vertical direction, and constrain this deviation to be less than a preset attitude threshold. The preset attitude threshold is set to 15 degrees. When the torso tilt angle approaches the threshold, the gradient... Guide the movement of the lumbar joints and chassis to restore the trunk to an upright position. The weighting coefficient for the trunk uprightness target is set to 0.8.
[0057] Collision avoidance target Calculate the minimum distance between each link of the arms and the torso, ensuring this distance is greater than a preset safety distance threshold. The preset safety distance threshold is set to 50 mm. When the arms approach the torso, the gradient... Guide the arm joints to move the arms away from the torso. The weighting coefficient for collision avoidance targets is set to 0.6.
[0058] The overall objective function is a weighted sum of the sub-objectives: .
[0059] In this embodiment, the system performs whole-body coordination control calculations at a frequency of 200 Hz. The pseudo-inverse of the Jacobian matrix is calculated using the damped least squares method, with the damping coefficient set to 0.01 to avoid numerical instability near singular configurations.
[0060] Step S6: Online estimation of the physical properties of the crawled object.
[0061] The parameter estimation module estimates the physical property parameters of the grasped object online based on the pressure distribution data of the tactile array and the end force feedback data.
[0062] During the steady-state phase after the grasping contact is established, the system acquires pressure distribution data sequences from the tactile array and force feedback data sequences from the end-effector six-dimensional force sensor. The data acquisition duration is set to 500 milliseconds.
[0063] Establish a simulation contact force model The model is based on the physical attribute parameters of the crawled object. Predict theoretical contact forces. Physical property parameters include mass parameters. and elastic modulus parameters The simulated contact force model adopts a simplified model based on Hertzian contact theory, taking into account the elastic contact deformation between the finger and the surface of the vessel.
[0064] The least squares optimization method is used to estimate the physical property parameters. The objective function of the optimization problem is: The time integral is calculated. Iterative optimization is performed using gradient descent, with 50 iterations and a learning rate of 0.01. The online estimation formula for the physical property parameters is as follows: .
[0065] Based on the estimated mass parameters and elastic modulus parameters, the system adaptively adjusts the grasping control parameters.
[0066] Adaptive mass parameter settings: The preset mass threshold is set to 500 grams. When the estimated mass parameter... When the weight exceeds 500 grams, the system classifies it as a heavy container, increases the maximum gripping force threshold to 15 Newtons, and reduces the angular velocity limit for tipping operations to 10 degrees per second. When the estimated mass parameter... When the weight is less than or equal to 500 grams, the upper limit of the grasping force threshold remains at 10 Newtons, and the angular velocity limit for the tilting operation remains at 20 degrees per second.
[0067] Adaptive elastic modulus parameter: The preset elastic threshold is set to 50 gigapascals. When the estimated elastic modulus parameter... When the force modulus is less than 50 gigapascals, the system identifies it as a thin-walled or flexible vessel, reduces the force control sensitivity parameter to 60% of its normal value, and activates a low-force gripping mode, reducing the target contact force to 5 Newtons. When the estimated elastic modulus parameter is greater than or equal to 50 gigapascals, the force control sensitivity parameter and target contact force remain at their normal values.
[0068] Example 2 This embodiment uses the transfer of corrosive liquids from an agrochemical warehouse as an application scenario. This task requires a robot to perform the grasping, handling, and dumping operations of containers containing corrosive liquids in remote operation mode, involving the dispensing and transfer of highly corrosive reagents such as concentrated sulfuric acid and concentrated hydrochloric acid.
[0069] The robot platform uses the same wheeled humanoid robot as in Example 1, standing 165 cm tall with 7 degrees of freedom in each arm and a three-finger dexterity hand design in each hand. Considering the special requirements of handling corrosive liquids, the robot's hands and forearms are covered with a corrosion-resistant protective coating. The chassis uses a differential drive structure, limiting the maximum movement speed to 0.8 m / s to ensure stability when handling liquids. The sensor configuration includes: a dual-finger tactile array sensor, with 6 tactile sensing units in each hand, a measurement range of 0 to 25 N, and a resolution of 0.03 N; a wrist six-dimensional force sensor, measuring force from 0 to 100 N in the force direction and torque from 0 to 10 Nm in the torque direction; a head-mounted RGB-D camera with a resolution of 1280 x 720 pixels and a frame rate of 30 frames per second; and joint encoders providing position, velocity, and torque feedback.
[0070] The remote control equipment configuration is the same as in Example 1. The operator station is located in a safety control room 15 meters away from the work area, and establishes a communication connection with the robot via a 5G wireless network. The network latency ranges from 80 to 150 milliseconds, with an average latency of 110 milliseconds.
[0071] The work environment is a reagent storage area in an agrochemical warehouse. The shelves are 2.2 meters high, with four layers, each holding 8 to 12 reagent containers of different sizes. The containers come in four sizes: 250 ml, 500 ml, 1 L, and 2.5 L, and are made of high-density polyethylene or borosilicate glass. The illumination in the work area is 300 to 500 lux, with slight visual interference from chemical vapors.
[0072] The control method in this embodiment includes the following steps: Step S1: Operator motion data acquisition and preprocessing.
[0073] The operator's motion data is collected in real time using VR controllers and a motion capture system. Because corrosive liquid transfer tasks require higher operational precision, the sampling frequency is increased to 150 Hz to obtain more detailed motion information.
[0074] The collected raw data was preprocessed. Position data underwent low-pass filtering to remove high-frequency noise; the filter cutoff frequency was set to 10 Hz, lower than in Example 1, to obtain a smoother trajectory. The wrist velocity vector was calculated based on two consecutive frames of position data. The calculation formula is: Calculate joint angular acceleration based on angular velocity data from three consecutive frames. The calculation formula is: The wrist pose data is represented as a homogeneous transformation matrix. .
[0075] The preprocessed data is temporally concatenated to construct a multimodal temporal coding sequence. The historical time window length L is set to 50 sampling points, corresponding to a time span of approximately 333 milliseconds, which is increased compared to Example 1 to provide more historical information for action intent prediction. The formula for constructing the multimodal temporal coding sequence is as follows: .
[0076] Step S2: Predicting action intent.
[0077] The intent prediction module uses a spatiotemporal attention encoder to process the multimodal temporal encoded sequence. The structure of the spatiotemporal attention encoder is the same as in Embodiment 1, including a feature embedding layer, a multi-layer self-attention encoding layer, and a prediction output layer.
[0078] The feature embedding layer maps data from different modalities to a unified hidden space. Wrist velocity data is mapped to a 256-dimensional hidden vector through a linear transformation. Joint angular acceleration data is mapped to a 256-dimensional hidden vector through a linear transformation. Pose trajectory data is mapped to a 256-dimensional hidden vector through flattening and a linear transformation. The three hidden vectors are concatenated to obtain a 768-dimensional fused feature vector.
[0079] The multi-layer self-attention encoding layer employs a 6-layer Transformer encoder structure, with each layer containing 8 attention heads. This increases the number of layers by 2 compared to Example 1, to improve the ability to model complex operational intentions. Position encoding uses a sinusoidal position encoding method.
[0080] The prediction output layer employs a multilayer perceptron structure. The prediction time domain is set to 10 time steps, corresponding to a prediction range of approximately 67 milliseconds, which is extended compared to Implementation Example 1 to handle longer network latency. The calculation formula for action intent prediction is as follows: ,in The value range is from 1 to 10.
[0081] The prediction confidence threshold is set to 0.92, which is higher than in Example 1 to reduce the risk of misprediction. When the prediction confidence is higher than 0.92, the system directly issues the predicted action command to the robot for execution. When the prediction confidence is lower than 0.92, the system triggers an operator confirmation mechanism, prompting the operator for secondary confirmation through VR controller vibration and voice prompts. The confirmation waiting time is set to 500 milliseconds.
[0082] Step S3: Acquisition of tactile array data and calculation of contact force.
[0083] The tactile sensing module collects pressure distribution data from the robot's hand tactile array. The tactile sensor sampling frequency is configured to 250 Hz, which is higher than that of Embodiment 1 to ensure a faster response to contact dynamics.
[0084] Pressure distribution data output by the tactile array It is a 6-dimensional vector. The pressure distribution data is converted into contact force data according to a pre-calibrated calibration model. The formula for calculating the contact force data is as follows: .
[0085] Target contact force The force was set to 6 Newtons, lower than in Example 1, to reduce the contact force on the container and decrease the risk of container breakage or liquid spillage. The tactile deviation between the contact force data and the target contact force was calculated using the following formula: .
[0086] Step S4: Compliance control and torque compensation.
[0087] The compliance control module generates torque compensation commands based on tactile deviation and compliance control laws. The formula for calculating the torque compensation commands is as follows: .
[0088] Scale-gain matrix in normal mode Set as a diagonal matrix with diagonal elements of 35 Nm / N, lower than in Example 1, to provide a smoother force control response. Differential gain matrix. Set it as a diagonal matrix with diagonal elements of 3 Nm / s.
[0089] Contact detection: The contact threshold is set to 0.3 N, lower than in Example 1, to improve the sensitivity of contact detection. After the initial contact event occurs, the system enters a buffer response mode, adjusting the proportional gain matrix. The value is reduced to 30% of the normal value, and the differential gain matrix is reduced. The value decreased to 40% of the normal value. After the buffer response mode lasted for 300 milliseconds, the gain matrix gradually recovered to the normal value. The recovery process used linear interpolation and the recovery time was 200 milliseconds.
[0090] Slip detection: The system continuously monitors the location of the center of pressure distribution data. The formula for calculating the distribution center is as follows: The slip threshold was set to 3 mm / s, lower than in Example 1, to improve slip detection sensitivity. After detecting a slip trend, the system increased the proportional gain matrix. The value should be reduced to 180% of the normal value.
[0091] Overvoltage protection: The overvoltage threshold is set to 20 N, lower than in Example 1, to better protect thin-walled containers. Overvoltage protection actions include: reducing the proportional gain matrix. The overvoltage value is reduced to 15% of the normal value; if the overvoltage lasts for more than 80 milliseconds, a rapid release action is performed and an alarm is triggered.
[0092] Liquid splash prevention: The system incorporates an additional liquid splash prevention mechanism. When the amplitude of the high-frequency vibration component of the haptic array pressure distribution exceeds the vibration threshold, a risk of liquid sloshing is identified. The vibration threshold is set to 0.8 N. Upon detecting a risk of liquid sloshing, the system reduces the end effector's movement speed to 50% of its current speed and increases the smoothness of the movement trajectory.
[0093] Step S5: Full-body coordination and control.
[0094] The whole-body coordination module constructs a whole-body coordination control model based on null-space projection. The calculation formula for whole-body joint velocity commands is as follows: .
[0095] Optimize objective function The weights of each sub-objective are adjusted to suit the liquid transport task: Center of gravity stability target The weighting coefficient is set to 1.5, which is higher than in Example 1 to enhance stability when handling liquid containers. Active chassis compensation is triggered when the centroid projection distance from the support polygon boundary is less than the safety margin. The safety margin is set to 80 mm.
[0096] Torso upright target The weighting coefficient is set to 1.0, which is higher than that of Example 1. The preset posture threshold is tightened to 10 degrees to ensure the stability of the torso when handling liquid.
[0097] Collision avoidance target The weighting coefficient remains at 0.6. The preset safe distance threshold is set to 60 mm.
[0098] Liquid surface stability target This embodiment adds a liquid surface stability target, constraining the end effector's attitude angle change rate to be less than a preset angular velocity threshold. The preset angular velocity threshold is set to 5 degrees per second. The weighting coefficient for the liquid surface stability target is set to 0.8.
[0099] The overall objective function is: .
[0100] The system performs whole-body coordinated control calculations at a frequency of 250 Hz. The pseudo-inverse of the Jacobian matrix is calculated using the damped least squares method, with the damping coefficient set to 0.015.
[0101] Step S6: Online estimation of the physical properties of the crawled object.
[0102] The parameter estimation module estimates the physical property parameters of the grasped object online based on the pressure distribution data of the tactile array and the end force feedback data.
[0103] In the steady-state phase after the grasp contact is established, the system collects data sequences, and the data collection time is set to 800 milliseconds, which is longer than that in Example 1 to obtain more stable estimation results.
[0104] Establish a simulation contact force model Physical property parameters include the mass parameter m and the elastic modulus parameter. and liquid filling ratio parameters The liquid filling ratio parameter is used to estimate the liquid loading level in the container. Iterative optimization is performed using gradient descent, with 80 iterations and a learning rate of 0.008. The online estimation formula for the physical property parameters is as follows: .
[0105] Adaptive mass parameter settings: The preset mass threshold is set to 300 grams. When the estimated mass parameter... For containers weighing over 300 grams, the system classifies them as heavy or heavily filled, increasing the maximum gripping force threshold to 12 Newtons and reducing the angular velocity limit for tipping operations to 5 degrees per second. When the estimated mass parameters... If the weight exceeds 800 grams, the system determines it to be an overweight container and activates the dual-arm collaborative gripping mode.
[0106] Adaptive elastic modulus parameter: The preset elastic threshold is set to 40 gigapascals. When the estimated elastic modulus parameter... When the force is less than 40 gigapascals, the system determines that it is a thin-walled or flexible container, reduces the force control sensitivity parameter to 50% of the normal value, and enables the low-force gripping mode, reducing the target contact force to 4 Newtons.
[0107] Liquid fill ratio adaptive: when the estimated liquid fill ratio When the value is greater than 0.7, the system determines it to be a high-fill container, further reduces the angular velocity limit of the tipping operation to 3 degrees per second, and activates the splash-proof mode during the handling process, limiting the maximum end acceleration to 0.3 meters per square second.
[0108] Example 3 This embodiment illustrates the trajectory tracking performance of the method of the present invention in the presence of network latency, and verifies the effectiveness of the action intent prediction module through comparative experiments.
[0109] The experimental platform uses the same wheeled humanoid robot system as in Example 1. The experimental scenario simulates an agricultural chemical laboratory environment, with five glass beakers of different sizes (100 ml, 250 ml, 500 ml, 1000 ml, and 2000 ml) placed on the workbench. The experimental task is to pick up the beakers from the left side of the workbench and place them at the target location on the right side of the workbench, with a horizontal movement distance of 800 mm per task.
[0110] The remote operation terminal device configuration is the same as in Example 1. To simulate a real network environment, a controllable network delay was artificially introduced in the experiment. The network delay was implemented using a software simulator, allowing for precise control of the delay time, with a fluctuation range of less than 5 milliseconds.
[0111] See Figure 3 This embodiment conducts four sets of comparative experiments: The first group of experiments: trajectory tracking experiment with a fixed network latency of 120 milliseconds.
[0112] Experimental setup: The network latency was fixed at 120 milliseconds. The operator controlled the robot via a teleoperation system to perform the task of grasping and placing a 250 ml beaker. Each method was repeated 20 times, and the trajectory data of each experiment was recorded.
[0113] When using the existing direct position mapping method, the robot's end effector position directly tracks the operator's hand position with a delay of 120 milliseconds. Due to the lack of motion intent prediction, the robot trajectory exhibits significant lag when the operator performs rapid movements. When the operator suddenly changes direction, the robot trajectory shows overshoot and oscillation. The root mean square value of the trajectory error is 23.5 mm, the maximum trajectory error is 45.2 mm, and the root mean square value of the trajectory jitter amplitude is 8.2 mm. During the grasping phase, the trajectory jitter causes a large deviation in the initial contact position between the fingers and the beaker, with an average contact position deviation of 12.3 mm.
[0114] When using the method of this invention, the spatiotemporal attention encoder predicts the operator's future action intentions based on historical motion data. The prediction time domain is set to 6 time steps, corresponding to a prediction range of 50 milliseconds. The system sends the predicted action instructions to the robot in advance for execution, effectively compensating for 50 milliseconds of the 120 millisecond network latency. The robot trajectory closely matches the expected trajectory, with the root mean square value of the trajectory error reduced to 3.8 mm, the maximum trajectory error reduced to 8.5 mm, and the root mean square value of the trajectory jitter amplitude reduced to 1.1 mm. During the grasping phase, the average contact position deviation is reduced to 2.1 mm.
[0115] The second set of experiments: Trajectory tracking experiment with varying network latency.
[0116] Experimental setup: Network latency varied randomly between 80 and 200 milliseconds, with a period of 500 milliseconds. The operator performed the task of grasping and placing a 500 ml beaker, with each method repeated 15 times.
[0117] When using existing techniques, the stability of the robot trajectory further decreases due to the uncertainty of network latency. The root mean square value of the trajectory error is 31.2 mm, and the root mean square value of the trajectory jitter amplitude is 12.5 mm. At the moment of sudden delay change, the robot exhibits a significant velocity jump, with a maximum velocity jump amplitude of 0.35 m / s.
[0118] When using the method of this invention, the spatiotemporal attention encoder can adaptively adjust the prediction strategy based on historical trajectory characteristics. When an increase in network latency is detected, the system automatically increases the number of steps in the prediction time domain. The root mean square value of the trajectory error is 5.6 mm, and the root mean square value of the trajectory jitter amplitude is 1.8 mm. The maximum value of the velocity jump amplitude is reduced to 0.08 m / s.
[0119] The third group of experiments: statistical experiments on the success rate and damage rate of grasping.
[0120] Experimental setup: The network latency was fixed at 120 milliseconds. The operator sequentially performed the tasks of grasping, moving, and placing five different sizes of beakers, repeating each size 10 times, for a total of 100 experiments. The grasping success rate, placement success rate, and beaker breakage rate were recorded.
[0121] The criteria for successful grasping are: the robot successfully grasps the beaker and the beaker does not slip during the lifting process. The criteria for successful placement are: the beaker is placed in the target position and the positional deviation is less than 20 mm. The criteria for damage are: the beaker breaks or cracks during grasping, moving, or placing.
[0122] Using existing techniques, the grasping success rate was 78%, with failures mainly due to excessive initial contact position deviation caused by trajectory jitter and unstable grasping force control leading to beaker slippage. The placement success rate was 72%, with failures mainly due to trajectory delay causing placement position deviations exceeding the threshold. The beaker breakage rate was 12%, with breakage mainly caused by excessive force at the moment of contact leading to the cracking of thin-walled beakers.
[0123] When using the method of this invention, the success rate of grasping is increased to 93%, the success rate of placement is increased to 91%, and the beaker breakage rate is reduced to 2%. Both instances of breakage when using the method of this invention occurred during the operation of a 100 ml thin-walled beaker. The reason for this is that the wall thickness of this size beaker is only 1.2 mm, making it extremely sensitive to contact force.
[0124] The fourth group of experiments: operator subjective evaluation experiment.
[0125] Experimental setup: Ten operators were invited to perform the same beaker grasping task using both existing and the methods of this invention. Each operator performed the experiment five times using each method. After the experiment, the operators subjectively rated the smoothness of operation, perceived latency, and fatigue level, with a rating range of 1 to 10, where 10 represents the best result.
[0126] When using existing technologies, the average score for operational smoothness was 4.2, indicating that operators generally reported sluggish robot response and the need for frequent adjustments to compensate for the delay. The average score for perceived delay was 3.5, indicating that operators clearly perceived the time difference between the robot's actions and their own. The average score for fatigue was 4.8, indicating that operators reported the need for high concentration to cope with the uncertainty caused by the delay.
[0127] When using the method of this invention, the average score for operational smoothness is 8.1, indicating that operators report timely robot responses and an experience close to direct operation. The average score for perceived delay is 7.6, showing a significant reduction in operator perception of delay. The average score for fatigue level is 7.9, indicating that operators report a relaxed and natural operation process without the need for conscious adjustment of movements.
[0128] Based on the above experimental results, the method of the present invention reduces trajectory error by 84% and trajectory jitter by 87% in the presence of network latency, increases the grasping success rate from 78% to 93%, reduces the beaker breakage rate from 12% to 2%, and improves the subjective score of operation smoothness by 93%. The experimental results verify the effectiveness of the action intent prediction module and the tactile closed-loop force control module of the present invention.
[0129] In summary, the embodiments disclosed herein have at least the following technical effects: This invention achieves action intent prediction through multimodal temporal coding sequences and spatiotemporal attention encoders, effectively compensating for the network latency effects in teleoperation and maintaining the continuity and stability of robot trajectories.
[0130] This invention achieves adaptive force control response through a compliant control law based on a tactile array. It can form a buffer response at initial contact to avoid damage to the vessel, and increase the gripping force when a slippage tendency is detected to prevent the vessel from falling.
[0131] This invention achieves coordination of multiple optimization objectives through a whole-body coordination control model with zero-space projection, avoiding posture jitter caused by instability in joint redundancy solutions.
[0132] This invention improves the system's adaptability to different types of containers by online estimation of physical property parameters and adaptive adjustment of control parameters.
[0133] It is understood that the above embodiments are merely exemplary implementations used to illustrate the principles of the present invention, and the present invention is not limited thereto. For those skilled in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A method for remotely controlling and grasping a humanoid robot, characterized in that, include: The operator's wrist speed data, joint angular acceleration data, and pose trajectory data are collected, and a multimodal temporal coding sequence is constructed by temporal splicing. Feature extraction and action intention prediction are performed on the multimodal temporal coding sequence, and a predicted action command is output. The robot collects pressure distribution data from the tactile array of its hand, converts the pressure distribution data into contact force data according to the calibration model, calculates the tactile deviation between the contact force data and the target contact force, and processes the tactile deviation based on a compliant control law to generate a torque compensation command. A whole-body coordinated control model is constructed based on zero-space projection. An optimization objective function is defined, which includes center of gravity stability, trunk uprightness and collision avoidance. The upper limb task instructions are fused with the redundant degrees of freedom optimization objective corresponding to the optimization objective function, and the whole-body joint velocity instructions are output. as well as Based on the pressure distribution data and end force feedback data, the mass parameters and elastic modulus parameters of the grasped object are calculated, and the grasping force threshold and force control sensitivity parameters are adaptively adjusted according to the mass parameters and elastic modulus parameters.
2. The humanoid robot teleoperation grasping control method according to claim 1, characterized in that, The formula for constructing the multimodal temporal coding sequence is as follows: ; in, From time At the time The multimodal temporal coding sequence, The length of the historical time window. for The wrist speed data at that moment, for The joint angular acceleration data at time [time], for The pose trajectory data at the given time.
3. The humanoid robot teleoperation grasping control method according to claim 2, characterized in that, Feature extraction and action intent prediction are performed on the multimodal temporal encoded sequence using a spatiotemporal attention encoder, wherein the calculation formula for action intent prediction is: ; in, For the predicted future The predicted action instruction at the specified time, Let be the mapping function of the spatiotemporal attention encoder. To predict the number of time-domain steps, The value ranges from 1 to the preset maximum prediction step number.
4. The humanoid robot teleoperation grasping control method according to claim 1, characterized in that, The formula for calculating the contact force data is: ; The formula for calculating the tactile deviation is: ; in, for The contact force data at that moment, The mapping function of the calibration model. for The pressure distribution data at time t, for The tactile deviation at that moment, The target contact force is denoted as .
5. The humanoid robot teleoperation grasping control method according to claim 4, characterized in that, The calculation formula for the torque compensation command is as follows: ; in, for The torque compensation command at that moment, It is a proportional gain matrix. The differential gain matrix is... The rate of change of the tactile deviation is denoted as .
6. The humanoid robot teleoperation grasping control method according to claim 5, characterized in that, The compliant control law also includes adaptive adjustment of contact state: When an initial contact event is detected, the values of the proportional gain matrix and the differential gain matrix are reduced to form a buffered response; When a shift in the distribution center of the pressure distribution data is detected, it is determined to be a slippage trend, and the value of the proportional gain matrix is increased to improve the gripping force.
7. The humanoid robot teleoperation grasping control method according to claim 1, characterized in that, The formula for calculating the whole-body joint velocity command is as follows: ; in, This refers to the velocity command for all joints in the body. This is the pseudo-inverse of the Jacobian matrix. The end-effector velocity corresponding to the upper limb task command. It is the identity matrix. For Jacobian matrices, The gradient of the optimization objective function is... This is the current joint angle vector.
8. The humanoid robot teleoperation grasping control method according to claim 7, characterized in that, The optimization objective function includes: The objective for center-of-gravity stability is to ensure that the robot's center-of-gravity projection falls within the supporting polygon. For targets with upright torso posture, the deviation of torso posture angle is constrained to be less than a preset posture threshold. Collision avoidance targets are achieved by ensuring that the distance between the arms and torso is greater than a preset safe distance threshold.
9. The humanoid robot teleoperation grasping control method according to claim 4, characterized in that, The mass parameters and elastic modulus parameters of the grasped object are calculated using the following formula: ; in, For the estimated quality parameters, For the estimated elastic modulus parameter, Let be the parameter vector to be estimated. For parameter-based The simulated contact force model output.
10. The humanoid robot teleoperation grasping control method according to any one of claims 1 to 9, characterized in that, The adaptive adjustment includes: When the mass parameter is greater than the preset mass threshold, the upper limit of the gripping force threshold is increased and the angular velocity limit of the tilting operation is reduced; When the elastic modulus parameter is less than the preset elastic threshold, the force control sensitivity parameter is reduced and the low-force gripping mode is enabled.