Skill acquisition and migration method for dexterous hand operation
By installing a multimodal data acquisition device and the Diffusion model on the end effector of a dexterous hand, the problem of learning dexterous hand operation skills is solved, and efficient multimodal data acquisition and skill transfer are achieved, supporting complex operation tasks.
Patent Information
- Application Number
- CN202511219205.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, dexterity hand operation skills are difficult to learn, remote operation teaching systems have a single data modality, visual images are not intuitive, teaching success rate is low, and it is difficult to handle complex contact operation tasks.
A multimodal operation data acquisition device is installed on the dexterous hand end effector, including an FSR flexible pressure sensor array, a bending sensor, an AD acquisition module, a DSP processing module, and a hand-eye camera. This device acquires image data of palm contact force distribution, finger joint angles, and viewing angles. The data is then processed and fused using a Transformer feature extractor to establish a diffusion model for training, thereby achieving synchronous control of the robot.
It improves the efficiency and generalization of multimodal data acquisition for dexterous hand operation skills, supports complex operation tasks, reduces device complexity, and enables efficient migration of multi-finger collaborative operation.
Smart Images

Figure CN120901995A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of robot control, and relates to a skill acquisition and migration method for dexterous hand operation. BACKGROUND
[0002] A dexterous hand generally refers to a multi-fingered end effector similar in size and degree of freedom to a human hand, which is hereinafter referred to as a dexterous hand. Due to the characteristics of high degree of freedom and complex interactive operation of the dexterous hand, the dexterous hand operation skill learning is difficult. The breakthrough of imitation learning and reinforcement learning algorithm has greatly developed the operation skill learning of high degree of freedom robots. However, such methods need rich simulation data or human demonstration data to guide the learning process. The imitation learning method is widely used in robot operation tasks due to its high training efficiency and predictable training effect. In the skill learning of the dexterous hand, the human hand action is usually mapped to the dexterous hand in a teleoperation mode, so as to collect skill success samples under the guidance of human action and realize the extraction of human skills. However, the existing teleoperation demonstration system usually uses joint or end position data for mapping, and has problems of single data mode, non-intuitive visual image, low demonstration success rate and low efficiency. The patent
Dexterous hand adaptive grasping method based on multi-modal fusion imitation learning, CN202411287014.6, authorized
Humanoid robot impedance control method and system based on multi-modal signal fusion, CN202310154443.5, accepted
Arm-hand robot skill learning method based on residual layered reinforcement learning, CN202410493736.0, authorized
[0003] The technical problem solved by the application is to overcome the deficiencies of the prior art, and to provide a skill acquisition and migration method for dexterous hand operation, to realize human-machine separated human multi-modal operation data acquisition, to increase dexterous hand operation skill representation, and to improve the efficiency and generalization of multi-finger cooperative operation skill migration.
[0004] The technical solution of the application is:
[0005] A skill acquisition and transfer method for dexterous hand operation, comprising:
[0006] A multi-modal operation data acquisition device is installed on the multi-finger end effector; a human hand wears the multi-finger end effector; the multi-finger end effector is operated by a human being according to specific data; the action of the multi-finger end effector is data-acquired by the multi-modal operation data acquisition device to obtain the trajectory of the multi-finger end effector perspective image contact force distribution and contact force size F (t) and finger joint angle information;
[0007] A position-encoding Transformer is used as a feature extractor, the feature extractor is trained, and the contact force distribution and contact force size F (t) , finger joint angle trajectory of the multi-finger end effector perspective image data processing and information fusion are performed to obtain a feature vector fused with information;
[0008] A training model Diffusion is established, the feature vector fused with information is used as the input of the training model Diffusion, and the finger joint angle is used as the expected output for training; the trained training model Diffusion is implanted into a robot; and the trained training model Diffusion is used to realize the synchronous control of the multi-finger end effector over the robot.
[0009] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the multi-finger end effector is in a glove-like structure; the multi-modal operation data acquisition device comprises an FSR flexible pressure sensor array, a bending sensor, an AD acquisition module, a DSP processing module, a USB hub, and a hand-eye camera;
[0010] The FSR flexible pressure sensor array is located at the palm of the multi-finger end effector, and the FSR flexible pressure sensor array covers the fingertip, finger pulp, and palm area of the multi-finger end effector.
[0011] The bending sensor is arranged at the back of the multi-finger end effector and located at the positions of the five fingers.
[0012] The FSR flexible pressure sensor array and the AD acquisition module are connected with the AD acquisition module; the AD acquisition module is connected with the DSP processing module.
[0013] The AD acquisition module, the DSP processing module, the USB hub and the hand-eye camera are installed at the back of the multi-finger end effector; the DSP processing module and the hand-eye camera are connected with the external control computer through the USB hub.
[0014] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the output contact force distribution matrix output by the FSR flexible pressure sensor array is in the form of:
[0015]
[0016] In the formula, n is the number of columns of the matrix;
[0017] m is the number of rows of the matrix.
[0018] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the sensor resistor is arranged in the FSR flexible pressure sensor and the bending sensor; the resistance change amount of the sensor resistor is converted into a digital voltage by the AD acquisition module, and the digital voltage is sent to the DSP processing module.
[0019] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the hand-eye camera is a camera device installed on the multi-finger end effector, which calculates the trajectory of the multi-finger end effector while recording the visual angle image during the movement of the multi-finger end effector.
[0020] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the DSP processing module receives the digital voltage of the FSR flexible pressure sensor and the digital voltage of the bending sensor transmitted by the AD acquisition module; the digital voltage of the FSR flexible pressure sensor is converted into the contact force distribution and the contact force size F (t) during the movement of the multi-finger end effector; and the digital voltage of the bending sensor is converted into the finger joint angle
[0021] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the training method of the feature extractor is:
[0022] The real feature-action pair is used as a positive sample; the negative sample is randomly combined; a discrimination model is constructed to distinguish between positive and negative samples; the discrimination model structure is a three-layer MLP, the output uses a sigmoid activation function, and the output is a discrimination probability of [0, 1]; the mutual information lower bound of the feature vector and the action vector is calculated using the JS divergence approximation; in the training process, the feature extractor is updated after every 2 discriminator update steps; the training process increases the mutual information, and the training of the feature extractor is completed.
[0023] In the skill acquisition and transfer method for dexterous hand operation, the feature extractor extracts the contact force distribution and the contact force F (t) The method for data processing is:
[0024] The m*n matrix in the form of the distribution matrix of the output contact force is processed into a k*5 matrix, k=m*n; the k*5 matrix is unfolded into a 1-dimensional column vector in column priority, as the first column of the output matrix; the column vector is obtained according to the unfolding order of the coordinate position (x, y) of the k*5 matrix, as the second and third columns of the output matrix; the area weight corresponding to each element position of the k*5 matrix is written as a column vector, as the fourth column; the fingertip area is assigned w tip , the finger palm area is assigned w finger , the palm area is assigned w palm , the palm edge area is assigned w edge , the typical value w tip =1.0, w finger =0.7, w palm =0.5, w edge =0.2; the increment of the k*5 matrix is calculated using the current time force matrix value, and the last time force matrix value is subtracted, and is unfolded into a 1-dimensional column vector in column priority, as the fifth column of the output matrix; thus, a k-dimensional sequence is obtained, each element of the sequence contains five tuples, which are normalized force size, x coordinate, y coordinate, area weight and dynamic increment;
[0025] Each element in the k-dimensional sequence is mapped to d_model dimensions using a linear layer to obtain an embedding vector; the embedding vector is calculated using a sinusoidal position encoding method at each position of the sequence to obtain a position vector, and the position vector and the embedding vector are element-wise stacked to obtain an output.
[0026] The multi-head attention mechanism calculation method is used to calculate multiple groups of attention on the output vector of the last module, and then the result matrix of all heads is spliced, projected back to the original space using a linear layer, and the non-linear expression ability is enhanced using a feedforward network to realize the attention module. The attention module is stacked for 4-6 layers to capture deeper spatial-mechanical patterns, and finally all the features at different positions are combined into a vector with a dimension of d_model by global average pooling.
[0027] In the skill acquisition and transfer method for dexterous hand operation, the finger joint angle The method for data processing is:
[0028] The finger joint angle feature vector is calculated
[0029]
[0030] In the formula, qscale is the joint angle;
[0031] is the joint angle lower limit;
[0032] is the element-wise multiplication.
[0033] Trajectory of multi-fingered end-effector The method for data processing is:
[0034] is the hand pose matrix; is expressed as a homogeneous matrix:
[0035]
[0036] wherein R is the hand back spatial position vector;
[0037] U is the hand back spatial pose matrix;
[0038] According to the homogeneous matrix of , the hand back spatial pose vector
[0039]
[0040] Finger joint angle feature vector is a 20-dimensional vector; the hand back spatial pose vector is a 6-dimensional vector; the finger joint angle is a 20-dimensional vector;
[0041] The sine values of the finger joint angle vector and are spliced by row to form a 40-dimensional joint position feature vector;
[0042] The hand back spatial pose vector is repeated 5 times by row to obtain a 30-dimensional hand pose feature vector;
[0043] The 40-dimensional joint position feature vector and the 30-dimensional hand pose feature vector are spliced by row to form a 70-dimensional position information feature vector;
[0044] The perspective image is processed by the method for data processing:
[0045] The feature vector is extracted from the perspective image
[0046] In the above-mentioned skill acquisition and transfer method for dexterous hand operation, the information fusion method is:
[0047] The contact force distribution and the contact force size F of the d_model dimension are obtained (t) The feature vector, the 70-dimensional position information feature vector, and the feature vector The feature vectors are spliced into column vectors to obtain the complete feature vector of information fusion.
[0048] The beneficial effects of the present application compared with the prior art are:
[0049] (1) The present application integrates the tactile sensor array (covering the fingertips, finger pads, and palm) and the joint bending sensor into a unified signal interface, reducing the complexity of the device, solving the problem of relying only on a single force signal or ignoring the contact position in the prior art, and significantly improving the data extraction capability for complex contact operations (such as sliding and pinching);
[0050] (2) The present application extracts and splices the personalized features of the data through position coding attention feature extraction, image feature multi-modal preprocessing, and joint angle data decoupling processing, improving the feature representation capability of multi-modal data;
[0051] (3) The present application combines the flexible joint angle sensor (20-dimensional joint vector) and the vSLAM pose tracking (T hand ∈ SE(3)) of the hand-eye camera, realizing the full parameterization of the hand motion, and providing more accurate geometric constraints for dexterous hand motion transfer;
[0052] (4) The present application implicitly learns the joint distribution of multi-modal data such as joint angle, contact force, and pose through the forward diffusion-reverse denoising process of the diffusion model, avoiding the dependence on deterministic mapping in traditional imitation learning, and being able to generate diverse and reasonable motion sequences (such as grasping strategies for different shaped objects);
[0053] (5) The hardware system of the present application has the characteristics of low cost, high integration, and light weight. The resistance type sensor (bending sensor, tactile array) and the commercial USB camera are used, and the hardware cost is reduced through a unified AD acquisition and synchronization protocol (timestamp alignment), while supporting modular expansion (such as adding an inertial measurement unit);
[0054] (6) The operation skill transfer system and method of the present application have complex operation task expansion. Through the fusion of force and position data, the system can support fine operations (such as plugging and assembling) and dynamic interaction. The standardized data package (including tactile, joint, pose, and image) collected can be adapted to different dexterous hand hardware (such as Shadow Hand and Allegro Hand), and only the joint mapping parameters (q scale , ) need to be adjusted. This mode is universal on various robots. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 Skill acquisition and transfer process for dexterous hand operation of the present application;
[0056] Figure 2 Schematic diagram of the multi-modal operation data acquisition device of the present application. DETAILED DESCRIPTION
[0057] The present application will be further described below in conjunction with examples.
[0058] The present application provides a skill acquisition and transfer method for dexterous hand operation, which improves the collection efficiency by separating the human-machine remote operation in the data collection process; high-efficiency synchronous collection of multi-modal feature information such as hand posture, joint angle, contact force and hand-eye image is achieved through special collection equipment; based on the pre-trained Transformer model to extract contact force features, an improved Diffusion operation skill imitation learning model with multiple data fusion is proposed. Compared with the real-time remote control imitation learning method of the human in the loop, the skill acquisition and transfer process is decoupled, and the learning efficiency of the dexterous hand operation skill is effectively improved.
[0059] The skill acquisition and transfer method for dexterous hand operation, as shown in Figure 1 , specifically includes the following steps:
[0060] Install a multi-modal operation data acquisition device on a multi-finger end effector, as shown in Figure 2 ; a human hand wears a multi-finger end effector; a human operates the multi-finger end effector according to specific data; the action of the multi-finger end effector is collected by the multi-modal operation data acquisition device to obtain the trajectory perspective image , contact force distribution and contact force size F (t) , and finger joint angle information.
[0061] The multi-modal operation data acquisition step is the action collection step of the human operation process in the robot skill transfer, which needs to complete the specific data in the human operation process to provide data and instructions for subsequent feature extraction and skill transfer training. The multi-modal operation data acquisition device combines palm internal tactile sensor, finger joint angle sensor, and hand-eye camera hardware to obtain palm internal contact force distribution, finger joint angle, hand-eye camera image, and hand motion pose trajectory data.
[0062] The multi-finger end effector is a glove-like structure; the multi-modal operation data acquisition device includes an FSR flexible pressure sensor array, a bending sensor, an AD acquisition module, a DSP processing module, a USB hub, and a hand-eye camera.
[0063] The FSR flexible pressure sensor array is located at the palm of the multi-finger end effector, and covers the finger tips, finger palms and palm centers of the multi-finger end effector.
[0064] The bending sensor is arranged at the back of the multi-finger end effector and is located at the positions of the five fingers.
[0065] The FSR flexible pressure sensor array and the AD acquisition module are connected with the AD acquisition module; the AD acquisition module is connected with the DSP processing module.
[0066] The AD acquisition module, the DSP processing module, the USB hub and the hand-eye camera are installed at the back of the multi-finger end effector; the DSP processing module and the hand-eye camera are connected with the external control computer through the USB hub.
[0067] The output contact force distribution matrix output by the FSR flexible pressure sensor array is in the form of:
[0068]
[0069] In the formula, n is the number of columns of the matrix;
[0070] m is the number of rows of the matrix.
[0071] The sensor resistor is arranged in the FSR flexible pressure sensor and the bending sensor; the resistance change amount of the sensor resistor is converted into a digital voltage by the AD acquisition module, and the digital voltage is sent to the DSP processing module.
[0072] The hand-eye camera is a camera device installed on the multi-finger end effector, which calculates the trajectory of the multi-finger end effector during the operation of the multi-finger end effector, and records the visual angle image at the same time.
[0073] The DSP processing module receives the digital voltage of the FSR flexible pressure sensor and the digital voltage of the bending sensor transmitted by the AD acquisition module; the digital voltage of the FSR flexible pressure sensor is converted into the contact force distribution and the contact force size F (t) during the operation of the multi-finger end effector; the digital voltage of the bending sensor is converted into the finger joint angle
[0074] The multi-modal operation data acquisition device features: a long strip-shaped bending sensor integrated on the back of the glove for measuring finger joint bending, a large-area array tactile sensor integrated on the palm side of the glove for covering the fingertip, finger pulp and palm area, and a fisheye lens USB camera installed on the palm back towards the web space for obtaining hand-eye video stream during operation. The joint bending sensor and the tactile sensor array both use resistance sensors, and the voltage signal is collected by an AD acquisition device to convert into contact force and finger joint angle information during hand operation.
[0075] During the acquisition process, the operator wears the glove, the USB interface of the hand back DSP processing module is connected with the USB port of the hand-eye camera through a USB hub, the USB interface output from the hub is connected with a control computer, and a program is run in the control computer to process and store the corresponding data. The control computer processes the hand-eye camera image data and obtains the camera pose according to the feature points tracked in the image, records the timestamps of each data, and then uniformly packages and stores them. The data collected by the system includes: finger joint angle (joint angle when picking up), contact force distribution and contact force size (large-area tactile sensor), hand spatial pose (hand-eye camera trajectory tracking), and hand-eye view image (hand-eye camera).
[0076] A position-encoding Transformer is used as a feature extractor, and the feature extractor is trained.
[0077] The training method of the feature extractor is:
[0078] Real feature-action pairs are used as positive samples, and random combinations are used as negative samples; a discrimination model is constructed to distinguish between positive and negative samples; the discrimination model structure is a three-layer MLP, the output uses a sigmoid activation function, and the output is a discrimination probability [0, 1]; the mutual information lower bound of the feature vector and the action vector is calculated using the JS divergence approximation; during the training process, the feature extractor is updated after every 2 discriminator update steps; the training process increases the mutual information, and the training of the feature extractor is completed.
[0079] The contact force distribution and contact force size F (t) , the finger joint angle Trajectory of the multi-finger end effector View image Data processing and information fusion are performed to obtain a complete feature vector.
[0080] The contact force distribution and contact force size F (t) The data processing method is:
[0081] The m*n matrix in the form of distribution matrix of output contact force is processed into a k*5 matrix, k=m*n; the k*5 matrix is unfolded into a 1-dimensional column vector in column priority, as the first column of the output matrix; the coordinate position (x, y) of the k*5 matrix is obtained according to the unfolding order to obtain a column vector, as the second and third columns of the output matrix; the area weight corresponding to each element position of the k*5 matrix is written as a column vector, as the fourth column; the fingertip region is assigned w tip , the finger palm region is assigned w finger , the palm region is assigned w palm , the palm edge region is assigned w edge , the typical value w tip =1.0, w finger =0.7, w palm =0.5, w edge =0.2; the increment of the k*5 matrix is calculated using the current moment matrix value and subtracting the last moment matrix value, and is unfolded into a 1-dimensional column vector in column priority, as the fifth column of the output matrix; thus, a k-dimensional sequence is obtained, each element of the sequence containing a five-tuple, respectively, the normalized force size, the x coordinate, the y coordinate, the area weight, and the dynamic increment.
[0082] Each element in the k-dimensional sequence is mapped to d_model dimensions using a linear layer to obtain an embedding vector; the embedding vector is calculated using a sinusoidal position encoding method at each position of the sequence to obtain a position vector, and the position vector and the embedding vector are element-wise superimposed as the output.
[0083] A multi-head attention mechanism calculation method is used to calculate multiple groups of attention on the output vector of the last module, then all the head calculation result matrices are spliced, projected back to the original space using a linear layer, and the non-linear expression ability is enhanced using a feedforward network to realize the attention module; the attention module is stacked for 4-6 layers to realize the feature extractor to capture deeper spatial-mechanical patterns, and finally all the features at all positions are merged into a vector with a dimension of d_model through global average pooling.
[0084] The method for processing the finger joint angle is as follows:
[0085] The finger joint angle feature vector is calculated as
[0086]
[0087] where q scale is the joint travel.
[0088] is the lower limit of the joint angle;
[0089] is the element-wise multiplication.
[0090] Trajectory of multi-fingered end effector The method for processing data is:
[0091] The is expressed as a homogeneous matrix:
[0092]
[0093] In the formula, R is a hand back spatial position vector.
[0094] U is a hand back spatial pose matrix.
[0095] According to the homogeneous matrix of , the hand back spatial pose vector
[0096]
[0097] The finger joint angle feature vector is a 20-dimensional vector; the hand back spatial pose vector is a 6-dimensional vector; the finger joint angle is a 20-dimensional vector.
[0098] The finger joint angle vector and The sine values are spliced by rows to form a 40-dimensional joint position feature vector.
[0099] The hand back spatial pose vector is repeated 5 times by rows to obtain a 30-dimensional hand pose feature vector;
[0100] The 40-dimensional joint position feature vector and the 30-dimensional head pose feature vector are spliced by rows to form a 70-dimensional position information feature vector.
[0101] The perspective image The method for processing data is:
[0102] The feature vector is extracted from the perspective image
[0103] The method for information fusion is:
[0104] The d_model-dimensional contact force distribution and the contact force size F (t) The feature vector, the 70-dimensional position information feature vector, and the feature vector are spliced by column vectors to obtain a complete feature vector of information fusion.
[0105] A training model called Diffusion is established, using the feature vector with complete information fusion as input to the training model Diffusion, and the finger joint angle as input. As the desired output, the model is trained; the trained Diffusion model is then implanted into the robot; and the multi-finger end effector is used to achieve synchronous control of the robot through the trained Diffusion model.
[0106] Based on the similarity in mechanical structure and body parameters between the human hand and a five-fingered dexterous hand, the joint position and contact force information obtained by the acquisition device during the operation process is processed. During the training phase, contact distribution data, finger joint data, hand spatial pose data, and hand-eye image data are synchronously used as the perception state. The joint angle position command obtained from the bending sensor is used as the desired action. The multimodal data distribution in the latent space is learned through a forward diffusion-backward denoising process of Diffusion. During the inference phase, real-time perception data is fed into the trained model to calculate the action sequence, controlling the robot's movement to complete skill reproduction.
[0107] 1. Diffusion model training. Forward diffusion process: For action sequence a... 1:N Gaussian noise is added gradually. The noise data for step i is:
[0108]
[0109] in α i This is the noise dispatch coefficient.
[0110] Reverse denoising process: Training the denoising network ∈ θ Predicted noise:
[0111]
[0112] Input is noise action a i State s, diffusion step i. Output is the predicted noise ∈ θ .
[0113] 2. Skill transfer and reproduction. Obtain the current state s. t (Touch, joints, pose, image features) are used to iteratively denoise and generate motion through a trained diffusion model.
[0114] 1) Initial noise sampling action
[0115] 2) Iterative calculation
[0116]
[0117] in σ iNoise variance.
[0118] 3) Output the final action a0.
[0119] 3) Robot control: convert action a0 into target joint angles:
[0120] Send to dexterous hand and robot arm controller for execution.
[0121] The present application reduces the complexity of the device by integrating the tactile sensor array (covering the fingertips, fingerpads, and palm) and the joint flexion sensor into a unified signal interface, solves the problem of relying only on a single force signal or ignoring the contact position in the prior art, and significantly improves the data extraction capability for complex contact operations (such as sliding and pinching).
[0122] The present application extracts and splices personalized features from the data by performing position coding attention feature extraction, image feature multi-modal preprocessing, and joint angle data decyclicization on the contact distribution data, thereby improving the feature representation capability of multi-modal data.
[0123] The present application combines a flexible joint angle sensor (20-dimensional joint vector) and a vSLAM pose tracking (T hand ∈SE(3)) of a hand-eye camera to achieve full parameterization of human hand motion and provide more accurate geometric constraints for dexterous hand action transfer.
[0124] The present application implicitly learns the joint distribution of multi-modal data such as joint angles, contact forces, and poses through the forward diffusion-reverse denoising process of the diffusion model, thereby avoiding the dependence of traditional imitation learning on deterministic mapping and enabling the generation of diverse and reasonable action sequences (such as grasping strategies for different shaped objects).
[0125] The hardware system described in the present application has the characteristics of low cost, high integration, and light weight. By using resistive sensors (flex sensors, tactile arrays) and commercial USB cameras, the hardware cost is reduced through a unified AD acquisition and synchronization protocol (timestamp alignment), while supporting modular expansion (such as adding an inertial measurement unit).
[0126] The fusion operation skill transfer system and method have complex operation task expandability. Through force-position fusion data, the system can support fine operations (such as plugging and assembling) and dynamic interactions. The collected standardized data packets (including tactile, joint, pose, and image) can be adapted to different dexterous hand hardware (such as Shadow Hand and Allegro Hand), and only the joint mapping parameters (q scale , ) need to be adjusted.
[0127] Although the present application has been disclosed with reference to the preferred embodiments, it is not intended to limit the present application, and any person skilled in the art can make possible changes and modifications to the technical solutions of the present application using the disclosed methods and technical contents without departing from the spirit and scope of the present application. Therefore, any simple modification, equivalent change and modification made to the above embodiments according to the technical essence of the present application without departing from the technical solutions of the present application shall fall within the protection scope of the technical solutions of the present application.
Claims
1. A skill acquisition and transfer method for dexterous hand manipulation, characterized by: The application relates to a multi-modal operation data acquisition device for a multi-finger end effector. A multi-finger end effector is worn on a human hand. The multi-finger end effector is operated according to specific data. The multi-finger end effector is in the form of a glove. The trajectory of the multi-fingered end effector is obtained by collecting data of actions of the multi-fingered end effector through a multi-modal operation data collection device Viewing angle image Contact force distribution and contact force size F (t) And finger joint angle Information; Adopt the transformer with position coding as the feature extractor, train the feature extractor; through the feature extractor, the contact force distribution and the contact force size F (t) , finger joint angle Trajectory of multi-finger end effector Viewing angle image Carry out data processing and information fusion to obtain a complete feature vector of information fusion; A training model Diffusion is established, the complete feature vector of information fusion is taken as the input of the training model Diffusion, and the finger joint angle As the expected output, training is carried out; the trained training model Diffusion is implanted into the robot; and the synchronization control of the multi-finger end effector on the robot is realized through the trained training model Diffusion. 2.The skill acquisition and transfer method for dexterous hand operation according to claim 1, wherein: The multi-modal operation data acquisition device comprises an FSR flexible pressure sensor array, a bending sensor, an AD acquisition module, a DSP processing module, a USB hub and a hand-eye camera. The FSR flexible pressure sensor array is arranged on the palm of the multi-finger end effector and covers the fingertips, finger pads and palm center of the multi-finger end effector. The bending sensor is arranged on the back of the multi-finger end effector and is located at the positions of the five fingers. The FSR flexible pressure sensor array and the AD acquisition module are connected with the AD acquisition module. The AD acquisition module, the DSP processing module, the USB hub and the hand-eye camera are arranged on the back of the multi-finger end effector.
3. The skill acquisition and transfer method for dexterous hand operation according to claim 2, wherein: The DSP processing module and the hand-eye camera are connected with an external control computer through the USB hub. The output contact force distribution matrix output by the FSR flexible pressure sensor array is in the form of: In the formula, n is the number of columns of the matrix.
4. The method for skill acquisition and transfer for dexterous hand operation according to claim 2, wherein: m is the number of rows of the matrix.
5. The method for skill acquisition and transfer for dexterous hand operation according to claim 2, wherein: The hand-eye camera is a camera device installed on the multi-finger end effector, which calculates the trajectory of the multi-finger end effector during the action of the multi-finger end effector At the same time, the recording of the perspective image is realized.
6. The method for skill acquisition and transfer for dexterous hand operation according to claim 2, wherein: The DSP processing module receives the digital voltage of the FSR flexible pressure sensor and the digital voltage of the bending sensor from the AD acquisition module; converts the digital voltage of the FSR flexible pressure sensor into the contact force distribution and the contact force size F during the action of the multi-finger end effector (t) ; and converts the digital voltage of the bending sensor into the finger joint angle during the operation of the human hand 7. The method for skill acquisition and transfer for dexterous hand operation according to claim 3, wherein: Sensor resistors are arranged in the FSR flexible pressure sensor and the bending sensor. The resistance change amount of the sensor resistors is converted into a digital voltage by the AD acquisition module, and the digital voltage is sent to the DSP processing module. The training method of the feature extractor is as follows:
8. The method for skill acquisition and transfer for dexterous hand operation according to claim 7, wherein: The feature extractor is applied to the contact force distribution and the contact force size F (t) The method for processing data is as follows: The m*n matrix in the form of distribution matrix of output contact force is processed into a k*5 matrix, k=m*n; the k*5 matrix is unfolded into a 1-dimensional column vector in column priority, as the first column of the output matrix; the coordinate position (x, y) of the k*5 matrix is obtained according to the unfolding order to obtain a column vector, as the second and third columns of the output matrix; the area weight corresponding to each element position of the k*5 matrix is written as a column vector, as the fourth column; the fingertip area is assigned w tip , the finger pad area is assigned w finger , the palm area is assigned w palm , the palm edge area is assigned w edge , the typical value is w tip =1.0, w finger =0.7, w palm =0.5, w edge =0.2; the increment of the k*5 matrix is unfolded into a 1-dimensional column vector in column priority using the current moment matrix value and subtracting the last moment matrix value, as the fifth column of the output matrix; thus a k-dimensional sequence is obtained, each element of the sequence contains a five-tuple, respectively, the normalized force size, x coordinate, y coordinate, area weight, and dynamic increment; Real feature-action pairs are used as positive samples, and negative samples are randomly combined. A discrimination model is constructed to distinguish positive and negative samples. The discrimination model structure is a three-layer MLP, the output uses a sigmoid activation function, and the output is a discrimination probability of [0, 1].
9. The method for skill acquisition and transfer for dexterous hand operation according to claim 8, wherein: To the angle of the finger joint The method for data processing is: Computing finger joint angle feature vectors In the formula, q scale is the joint travel; is the lower joint angle limit; The mutual information lower bound of the feature vector and the action vector is calculated using the JS divergence approximation. Trajectory for multi-fingered end effector The method for processing data is: will be described below. is expressed as a homogeneous matrix: In the training process, the feature extractor is updated after every 2 discriminator update steps. The training process increases the mutual information and completes the training of the feature extractor. According to the homogeneous matrix, the hand back spatial pose vector finger joint angle feature vector is a 20 dimensional vector; back of hand spatial pose vector is a 6 dimensional vector; finger joint angle is a 20 dimensional vector; concatenating the sine values of the finger joint angle vectors and to form a 40-dimensional joint position feature vector; Hand back spatial pose vector Repeat 5 times by row, get 30-dimensional hand pose feature vector; Each element in the k-dimensional sequence is mapped to a d_model-dimensional embedding vector using a linear layer. Viewing angle image The method for data processing is: extracting a feature vector from a perspective image using an image encoder ViT 10. The method for skill acquisition and transfer for dexterous hand operation according to claim 9, wherein: The position vector is calculated using the sine position encoding method at each position of the embedding vector. The position vector and the embedding vector are element-wise superimposed as the output. The multi-head attention mechanism is used to calculate multiple attention groups for the output vector of the previous module. Then, all the head calculation result matrices are spliced, projected back to the original space using a linear layer, and enhanced using a feedforward network to improve the non-linear expression ability. The attention module is stacked for 4-6 layers to capture deeper spatial-mechanical patterns. The features of all positions are merged into a vector with a dimension of d_model through global average pooling. is an element-wise multiplication. In the formula, R is a hand back spatial position vector. U is a hand back spatial pose matrix. A 40-dimensional joint position feature vector and a 30-dimensional head pose feature vector are spliced in rows to form a 70-dimensional position information feature vector. The information fusion method is as follows: The contact force distribution of d_model dimensions and the contact force size F (t) The feature vector, the 70-dimensional position information feature vector, and the feature vector The feature vectors are spliced into column vectors to obtain the complete feature vector of information fusion.
Citation Information
Patent Citations
Robot humanoid impedance control method and system based on multimodal signal fusion
CN116300436B
A skill learning method for arm-hand robots based on residual hierarchical reinforcement learning
CN118410705B
Adaptive grasping method for dexterous hands based on multimodal fusion imitation learning
CN118769260B