A Dexterous Hand Telemanipulation Method and System Based on Ultrasonic Imaging

CN120347752BActive Publication Date: 2026-08-14HARBIN INST OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0006]本发明为了解决现有灵巧手控制在信号采集技术中的一些常见问题,这些问题包括信号精度不足、噪声干扰较多、适应性差、个体差异较大,以及对人体可能造成的侵入性破坏等

Benefits of technology

[0065] This invention aims to address the problems of signal acquisition accuracy, noise interference, poor adaptability, large individual differences, and invasiveness associated with existing dexterity hand techniques. The system achieves precise control of the dexterity hand by acquiring ultrasound animations of arm muscles and hand kinematic data streams, combined with imitation learning and IRL methods. Ultrasound animations provide real-time motion signals by monitoring forearm muscle movement, while hand kinematic data streams acquire spatial motion information at 21 key points through data gloves or gesture recognition technology. This data is used to train an intelligent agent, learning by imitating expert demonstrations to optimize dexterity hand control strategies. This system not only improves learning efficiency but also avoids the invasiveness of traditional electromyography (EMG) signal acquisition. Ultimately, the trained model can achieve complex control of the dexterity hand, particularly suitable for amputees, providing a more personalized and efficient dexterity hand control solution with broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347752B_ABST
    Figure CN120347752B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for dexterous hand telemanipulation based on ultrasound imaging. It pertains to the fields of ultrasound imaging research and dexterous hand telemanipulation technology. The invention performs ultrasound imaging on the human arm, acquiring ultrasound animations of the arm muscles and recording the kinematic data stream of the hand. Based on the recorded data stream, a neural network is built, and the reward function of the model is learned through inverse reinforcement learning. The learned reward function is then used for forward reinforcement learning to learn the optimal strategy of the model. The optimized model is deployed to a server, and the ultrasound system and the dexterous hand are connected to the server to achieve precise telemanipulation control of the dexterous hand. This invention not only improves the learning efficiency of human hand movement information but also avoids the invasiveness of traditional electromyography (EMG) signal acquisition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of ultrasound imaging research and dexterous hand teleoperation technology for robots, and is a dexterous hand teleoperation method and system based on ultrasound imaging. Background Technology

[0002] Dexterous hand telemanipulation is an important research direction in the field of robotics in recent years. It combines advanced technologies such as robotics, brain-computer interfaces, electromyography (EMG) signal control, and artificial intelligence, aiming to provide amputees with a more natural and efficient dexterous hand manipulation experience. Compared with traditional prostheses, dexterous hand prostheses can simulate more complex hand movements, helping users complete fine motor tasks such as grasping and moving objects. The introduction of telemanipulation technology allows patients to control their dexterous hand through direct neural or EMG signals, achieving higher precision and more personalized motor control, thereby improving their quality of life. Currently, research and application of dexterous hand telemanipulation mainly focus on sensory feedback systems, control strategies, and the precision of actuators.

[0003] The technical approach to dexterous hand teleoperation mainly comprises two core components: signal acquisition and processing, and the implementation of a fine control system. In terms of signal acquisition and processing, electromyography (EMG) and brain-computer interface (BCI) technologies are currently widely used. EMG captures muscle activity in the residual limb to achieve basic control of the dexterous hand; this method has a mature application background, is simple to operate, and does not require complex surgical intervention. BCI technology, on the other hand, directly decodes neural signals from the brain, translating the user's thought commands into dexterous hand movements. This method offers greater flexibility and naturalness, especially for patients with complete amputations, providing a control experience close to that of a normal hand. Through these technologies, the dexterous hand can perform complex movements, such as independent control of multiple fingers, force adjustment, and precise grasping. Simultaneously, a feedback system is also an indispensable part of dexterous hand teleoperation. Because dexterous hand operation requires a high degree of coordination with the body's sensory system, integrating tactile and force sensors can provide users with real-time feedback information during operation, helping them perceive the state of their dexterous hand and interaction with objects, thereby improving the accuracy and comfort of operation.

[0004] Despite the progress made in dexterous hand telemanipulation technology, several technical challenges remain. First, the accuracy and real-time performance of signal decoding remain critical issues. Whether using electromyography (EMG) signals or brain-computer interfaces, signal acquisition is often subject to noise interference, leading to low decoding accuracy and consequently affecting the dexterous hand's operational performance. This is especially true when performing subtle movements, where even minor errors can cause failure or inaccuracy. Even with advanced algorithms, signal stability and real-time performance are difficult to guarantee, particularly when rapidly switching control modes or performing complex tasks; latency and instability become obstacles to technological advancement. Furthermore, individual differences in signals are also a concern. Each patient's EMG or EEG characteristics differ, requiring the system to be able to personalize and learn. While some adaptive algorithms and machine learning methods have made progress in this area, enabling the dexterous hand to quickly adapt to individual differences within a short period remains a problem worthy of further research.

[0005] Current dexterous hand telemanipulation technology based on ultrasound sensors has demonstrated advantages such as non-invasive detection and deep muscle dynamic capture. However, at the motion analysis level, existing solutions mostly rely on machine learning models to regress joint angles from ultrasound signals, resulting in bottlenecks such as high dependence on model training data and significant real-time motion response delays. Unlike traditional methods, this solution creatively integrates the subtraction processing of images obtained from the literature "A Flexible Ultrasonic Sensor and Its Arterial Blood Pressure Detection Method" (CN112869773A) that shows the depth changes of the anterior and posterior boundaries of human arterial vessels over time. This invention uses this method to obtain the dynamic deformation difference value between adjacent frames of the muscle cross-section ultrasound image, and simultaneously selects the difference data from the first N frames after muscle activation to construct a feature analysis window. This invention can analyze the refined movement intentions of finger joints in real time based on the difference data between flexor and extensor muscle groups. Summary of the Invention

[0006] This invention addresses several common problems in signal acquisition technology for dexterous hand control, including insufficient signal accuracy, significant noise interference, poor adaptability, large individual differences, and potential invasiveness. To resolve these issues, this invention proposes a method and system for dexterous hand telemanipulation based on ultrasound imaging. This system enables control of the dexterous hand by acquiring ultrasound animations of arm muscles and data streams of prosthetic hand movements. Through imitation learning, a trained model is deployed within the dexterous hand system, thereby achieving precise control of the dexterous hand.

[0007] This invention provides the following technical solutions:

[0008] A method and system for dexterous hand telemanipulation based on ultrasound imaging, the method comprising the following steps:

[0009] Step 1: Perform ultrasound imaging on the human arm, acquire ultrasound animation of the arm muscles, and record the kinematic data stream of the hand;

[0010] Step 2: Based on the recorded data stream, build a neural network, design a reward function model, and learn the model's reward function through inverse reinforcement learning.

[0011] Step 3: Perform positive reinforcement learning using the learned reward function to learn the optimal policy of the model;

[0012] Step 4: Deploy the optimized model to the server, connect the ultrasound system and the dexterous hand to the server, and achieve precise control of the dexterous hand.

[0013] Preferably, step 1 specifically includes:

[0014] The ultrasound imaging of the arm muscles was taken at the midpoint of the forearm, and the imaging information needed to clearly show the superficial and deep flexor muscles of the forearm, the thumb flexor, the ulna, the radius, and the muscle and bone structures.

[0015] Preferably, in terms of acquiring hand kinematic data streams, in addition to acquiring finger movement information through data gloves, RGB / RGB-D sensors can also be used for gesture recognition. The spatial motion information of key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point of the wrist, and 4 key points of each finger of the thumb, index finger, middle finger, ring finger and little finger.

[0016] Each frame of data collected constitutes a state space containing 21 parameters, comprehensively recording the movement state of each joint and fingertip of the hand.

[0017] Preferably, step 2 specifically comprises:

[0018] The first N frames of arm ultrasound images are used to perform inter-frame difference calculations, and the results of the first few differences are then weighted and summed to extract muscle movement change features. The calculation formula is as follows:

[0019]

[0020] Where (x, y) represents the pixel coordinates, I t (x,y) represents the current frame (at time t), I t-i (x,y) represents the first i frames, N represents the number of frames for which differential calculation needs to be performed, and in this invention, w represents the hyperparameters that need to be adjusted. i Represents the weighting coefficients, satisfying D t (x,y) represents the final weighted difference result at the current time (t).

[0021] By calculating weighted difference data and hand movement data, a TransFormer network conforming to its data dimensions is constructed. The input data is D. t-i (x,y), i∈[0,M], where each D is an input feature and M is a hyperparameter;

[0022] The reward function is optimized using IRL. The goal of IRL is to maximize the reward function obtained from the expert demonstration dataset, i.e., the collected expert action data. Within the IRL framework, the specific reward function is learned and optimized by maximizing the Q-value difference between expert actions and other possible actions. The optimization formula is as follows:

[0023]

[0024] Where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S represents the state space, s represents each independent state in the state space, A represents the action space, and a represents each independent action in the action space.

[0025] The reward function is designed as follows:

[0026] R = R sync +R stable +R eff

[0027] Among them, R sync R represents the reward for muscle action timing calibration. stable R represents the reward for muscle synergistic stability. eff This indicates a composite reward for action performance.

[0028] Specifically:

[0029]

[0030] Where D ultra This represents the inter-frame difference sequence of ultrasound, indicating the deformation rate of different muscles. (D) hand This represents the hand joint angular velocity sequence, indicating the motion velocity of each joint. DTW(·) represents the dynamic time warping algorithm, used to align the delay differences between two sets of time series data. σ represents the time calibration tolerance parameter.

[0031] This reward quantifies the physiological delay matching between muscle activation and joint movement using a Dynamic Time Warping (DTW) algorithm. For example, when an expert grasps an object, the differential ultrasound signal (D) of the flexor digitorum superficialis muscles... ultra It will occur earlier than the finger joint flexion movement (D) handThis is an inherent delay characteristic of neuromuscular transmission in the human body. The reward function, through exponential decay, penalizes strategies that do not conform to this physiological law, ensuring the temporal consistency of dexterity hand movements with the human kinetic chain.

[0032]

[0033] Where ΔU t =U t -U t-1 Let represent the difference matrix between adjacent ultrasound frames, and represent the instantaneous change in muscle activation intensity. W represents the muscle synergy weight matrix, ||·|| F This represents the Frobenius norm of the matrix, used to measure the overall magnitude of change in the matrix. ∈ indicates a very small quantity (taken as 10e-5) to prevent the denominator from being zero.

[0034] This reward item assesses the synergistic stability of antagonistic muscle groups using the matrix F norm. For example, in a normal fist-clenching motion, flexor activation should be accompanied by extensor inhibition (ΔU). t (Corresponding elements have opposite signs) If both are enhanced simultaneously, it will lead to stiff movements. The weight matrix W amplifies the abnormal synergistic signals of antagonistic muscle pairs, and through normalization processing, quantifies the overall coordination of the muscle group into a reward value of 0-1.

[0035]

[0036] Among them, v h v g a represents the current velocity vector of the hand's end effector and the target direction vector. h Cov(D) represents the angular acceleration vector of the hand joints. ultra Var(D) represents the determinant of the covariance matrix of the inter-frame differences of all muscle channels. ultra ) represents the sum of variances of the inter-frame differences of all muscle channels, and α, β, and γ represent adaptively adjusted parameters that decay exponentially with increasing training rounds.

[0037] This composite reward optimizes action efficiency from three dimensions:

[0038] Trajectory Direction Reward (First Item): Guide the hand to move in the target direction using the cosine similarity of the velocity vector.

[0039] Exercise-economic reward (second item): Suppressing high-frequency jitter, consistent with the principle of minimizing muscle energy (e.g., rapid jitter leads to ||a h || 2 (Too high).

[0040] Muscle group coordination reward (third item): Covariance and variance ratio measure the synchronicity of muscle activation. For example, the gripping action requires high covariance (synchronous contraction) of the finger flexor muscles, while there should be negative covariance (antagonistic effect) between the flexor and extensor muscles.

[0041] By maximizing the reward function in the expert demonstration dataset, the agent can learn the expert's action patterns and then imitate the expert's behavior.

[0042] The training model uses a lightweight convolutional neural network to process ultrasound image features and combines a temporal model to capture the dynamic temporal correlation of muscle movement.

[0043] Preferably, step 3 specifically comprises:

[0044] Optimize the learning strategy using maximum causal entropy regularization:

[0045]

[0046] Here, H(x) represents the information entropy of the random variable x.

[0047] Preferably, IRL(π) E Maximize the reward difference between expert strategies and reinforcement learning strategies by learning a reward function;

[0048] RL(R) represents positive reinforcement learning based on entropy regularization. Entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the agent's policy is, and the closer it is to the expert's decision.

[0049] The parameters are updated using the policy gradient, which is formulated as follows:

[0050]

[0051] Where θ represents the network parameters, J represents the policy optimization objective function, and α represents the entropy regularization coefficient (a hyperparameter that balances reward and uncertainty, dynamically controlling the degree of policy determinism by adaptively adjusting α).

[0052] The parameters are updated using gradient ascent via the policy gradient formula, which is:

[0053]

[0054] Where η represents the learning rate.

[0055] Preferably, step 4 specifically comprises:

[0056] The ultrasonic signal is preprocessed by hardware for noise reduction and motion compensation before being input into the neural network. Finally, it passes through a fully connected layer to output corresponding hand movement commands. The PD controller is used to precisely control the movement of the dexterous hand. The control latency is optimized end-to-end through edge computing devices, and the inference time is compressed to less than 5 milliseconds.

[0057] A dexterous hand teleoperation system based on ultrasound imaging, the system comprising:

[0058] The data acquisition module performs ultrasound imaging on the human arm, acquires ultrasound animations of the arm muscles, and records the kinematic data stream of the hand.

[0059] The model building module builds a neural network based on the recorded data stream and learns the model's reward function through inverse reinforcement learning.

[0060] An optimization module learns the optimal policy of the model by performing positive reinforcement learning through the learned reward function;

[0061] The control module deploys the optimized model to the server and connects the ultrasound system and the dexterous hand to the server to achieve precise control of the dexterous hand.

[0062] A computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a dexterous hand telemanipulation method based on ultrasound imaging.

[0063] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a dexterous hand teleoperation method based on ultrasound imaging.

[0064] The present invention has the following beneficial effects:

[0065] This invention aims to address the problems of signal acquisition accuracy, noise interference, poor adaptability, large individual differences, and invasiveness associated with existing dexterity hand techniques. The system achieves precise control of the dexterity hand by acquiring ultrasound animations of arm muscles and hand kinematic data streams, combined with imitation learning and IRL methods. Ultrasound animations provide real-time motion signals by monitoring forearm muscle movement, while hand kinematic data streams acquire spatial motion information at 21 key points through data gloves or gesture recognition technology. This data is used to train an intelligent agent, learning by imitating expert demonstrations to optimize dexterity hand control strategies. This system not only improves learning efficiency but also avoids the invasiveness of traditional electromyography (EMG) signal acquisition. Ultimately, the trained model can achieve complex control of the dexterity hand, particularly suitable for amputees, providing a more personalized and efficient dexterity hand control solution with broad application prospects. Attached Figure Description

[0066] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0067] Figure 1 The diagram shown is a schematic representation of the method flow of the present invention.

[0068] Figure 2 The flowchart shown is a process for learning a specific reward function using expert data, as described in this invention.

[0069] Figure 3 The diagram shown is a flowchart of the data acquisition, reasoning, and control processes of this invention. Detailed Implementation

[0070] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0071] The present invention will be described in detail below with reference to specific embodiments. Specific Implementation Example 1:

[0073] according to Figures 1 to 3 As shown, the specific optimized technical solution adopted by the present invention to solve the above-mentioned technical problems is: The present invention relates to a dexterous hand teleoperation method and system based on ultrasound imaging.

[0074] This invention provides a method and system for dexterous hand telemanipulation based on ultrasound imaging, the method comprising the following steps:

[0075] The method includes the following steps:

[0076] Step 1: Perform ultrasound imaging on the human arm, acquire ultrasound animation of the arm muscles, and record the kinematic data stream of the hand;

[0077] Step 2: Based on the recorded data stream, build a neural network and learn the model's reward function through inverse reinforcement learning.

[0078] Step 3: Perform positive reinforcement learning using the learned reward function to learn the optimal policy of the model;

[0079] Step 4: Deploy the optimized model to the server, connect the ultrasound system and the dexterous hand to the server, and achieve precise control of the dexterous hand. Specific Implementation Example 2:

[0081] The only difference between Embodiment 2 and Embodiment 1 of this application is that:

[0082] Step 1 is as follows:

[0083] The ultrasound imaging of the arm muscles was taken at the midpoint of the forearm, and the imaging information needed to clearly show the superficial and deep flexor muscles of the forearm, the thumb flexor, the ulna, the radius, and the muscle and bone structures. Specific Implementation Example 3:

[0085] The only difference between Embodiment 3 and Embodiment 2 of this application is that:

[0086] In terms of acquiring hand kinematics data streams, in addition to obtaining finger movement information through data gloves, RGB / RGB-D sensors can also be used for gesture recognition. The spatial motion information of key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point of the wrist, and 4 key points of each finger of the thumb, index finger, middle finger, ring finger and little finger.

[0087] Each frame of data collected constitutes a state space containing 21 parameters, comprehensively recording the movement state of each joint and fingertip of the hand. Specific Implementation Example 4:

[0089] The only difference between Embodiment 4 and Embodiment 3 of this application is that:

[0090] Step 2 specifically involves:

[0091] The first N frames of arm ultrasound images are used to perform inter-frame difference calculations, and the results of the first few differences are then weighted and summed to extract muscle movement change features. The calculation formula is as follows:

[0092]

[0093] Where (x, y) represents the pixel coordinates, I t (x,y) represents the current frame (at time t), I t-i (x,y) represents the first i frames, N represents the number of frames for which differential calculation needs to be performed, and in this invention, w represents the hyperparameters that need to be adjusted. i Represents the weighting coefficients, satisfying D t (x,y) represents the final weighted difference result at the current time (t).

[0094] By calculating weighted difference data and hand movement data, a TransFormer network conforming to its data dimensions is constructed. The input data is D. t-i (x,y), i∈[0,M], where each D is an input feature and M is a hyperparameter;

[0095] The reward function is optimized using IRL. The goal of IRL is to maximize the reward function obtained from the expert demonstration dataset, i.e., the collected expert action data. Within the IRL framework, the specific reward function is learned and optimized by maximizing the Q-value difference between expert actions and other possible actions. The optimization formula is as follows:

[0096]

[0097] Where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S represents the state space, s represents each independent state in the state space, A represents the action space, and a represents each independent action in the action space.

[0098] The reward function is designed as follows:

[0099] R = R sync +R stable +R eff

[0100] Among them, R sync R represents the reward for muscle action timing calibration. stable R represents the reward for muscle synergistic stability. eff This indicates a composite reward for action performance.

[0101] Specifically:

[0102]

[0103] Where D ultra This represents the inter-frame difference sequence of ultrasound, indicating the deformation rate of different muscles. (D) hand This represents the hand joint angular velocity sequence, indicating the motion velocity of each joint. DTW(·) represents the dynamic time warping algorithm, used to align the delay differences between two sets of time series data. σ represents the time calibration tolerance parameter.

[0104] This reward quantifies the physiological delay matching between muscle activation and joint movement using a Dynamic Time Warping (DTW) algorithm. For example, when an expert grasps an object, the differential ultrasound signal (D) of the flexor digitorum superficialis muscles... ultra It will occur earlier than the finger joint flexion movement (D) hand This is an inherent delay characteristic of neuromuscular transmission in the human body. The reward function, through exponential decay, penalizes strategies that do not conform to this physiological law, ensuring the temporal consistency of dexterity hand movements with the human kinetic chain.

[0105]

[0106] Where ΔU t =U t -U t-1 Let represent the difference matrix between adjacent ultrasound frames, and represent the instantaneous change in muscle activation intensity. W represents the muscle synergy weight matrix, ||·|| F This represents the Frobenius norm of the matrix, used to measure the overall magnitude of change in the matrix. ∈ indicates a very small quantity (taken as 10e-5) to prevent the denominator from being zero.

[0107] This reward item assesses the synergistic stability of antagonistic muscle groups using the matrix F norm. For example, in a normal fist-clenching motion, flexor activation should be accompanied by extensor inhibition (ΔU). t (Corresponding elements have opposite signs) If both are enhanced simultaneously, it will lead to stiff movements. The weight matrix W amplifies the abnormal synergistic signals of antagonistic muscle pairs, and through normalization processing, quantifies the overall coordination of the muscle group into a reward value of 0-1.

[0108]

[0109] Among them, v h v g a represents the current velocity vector of the hand's end effector and the target direction vector. h Cov(D) represents the angular acceleration vector of the hand joints. ultra Var(D) represents the determinant of the covariance matrix of the inter-frame differences of all muscle channels. ultra ) represents the sum of variances of the inter-frame differences of all muscle channels, and α, β, and γ represent adaptively adjusted parameters that decay exponentially with increasing training rounds.

[0110] This composite reward optimizes action efficiency from three dimensions:

[0111] Trajectory Direction Reward (First Item): Guide the hand to move in the target direction using the cosine similarity of the velocity vector.

[0112] Exercise-economic reward (second item): Suppressing high-frequency jitter, consistent with the principle of minimizing muscle energy (e.g., rapid jitter leads to ||a h || 2 (Too high).

[0113] Muscle group coordination reward (third item): Covariance and variance ratio measure the synchronicity of muscle activation. For example, the gripping action requires high covariance (synchronous contraction) of the finger flexor muscles, while there should be negative covariance (antagonistic effect) between the flexor and extensor muscles.

[0114] By maximizing the reward function in the expert demonstration dataset, the agent can learn the expert's action patterns and then imitate the expert's behavior.

[0115] The training model uses a lightweight convolutional neural network to process ultrasound image features and combines a temporal model to capture the dynamic temporal correlation of muscle movement. Specific Implementation Example 5:

[0117] The difference between Embodiment 5 and Embodiment 4 of the present invention lies only in:

[0118] Step 3 specifically involves:

[0119] Optimize the learning strategy using maximum causal entropy regularization:

[0120]

[0121] Here, H(x) represents the information entropy of the random variable x. Specific Implementation Example Six:

[0123] The difference between Embodiment Six and Embodiment Five of the present invention lies only in:

[0124] IRL(π E Maximize the reward difference between expert strategies and reinforcement learning strategies by learning a reward function;

[0125] RL(R) represents positive reinforcement learning based on entropy regularization. Entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the agent's policy is, and the closer it is to the expert's decision.

[0126] The parameters are updated using the policy gradient, which is formulated as follows:

[0127]

[0128] Where θ represents the network parameters, J represents the policy optimization objective function, and α represents the entropy regularization coefficient (a hyperparameter that balances reward and uncertainty, dynamically controlling the degree of policy determinism by adaptively adjusting α).

[0129] The parameters are updated using gradient ascent via the policy gradient formula, which is:

[0130]

[0131] Where η represents the learning rate. Specific Implementation Example 7:

[0133] The difference between Embodiment Seven and Embodiment Six of the present invention lies only in:

[0134] Step 4 is as follows:

[0135] The ultrasonic signal is preprocessed by hardware for noise reduction and motion compensation before being input into the neural network. Finally, it passes through a fully connected layer to output corresponding hand movement commands. The PD controller is used to precisely control the movement of the dexterous hand. The control latency is optimized end-to-end through edge computing devices, and the inference time is compressed to less than 5 milliseconds. Specific Implementation Example 8:

[0137] The difference between Embodiment 8 and Embodiment 7 of the present invention lies only in:

[0138] This invention provides a dexterous hand teleoperation system based on ultrasound imaging, characterized in that: the system comprises:

[0139] The data acquisition module performs ultrasound imaging on the human arm, acquires ultrasound animations of the arm muscles, and records the kinematic data stream of the hand.

[0140] The model building module builds a neural network based on the recorded data stream and learns the model's reward function through inverse reinforcement learning.

[0141] An optimization module learns the optimal policy of the model by performing positive reinforcement learning through the learned reward function;

[0142] The control module deploys the optimized model to the server and connects the ultrasound system and the dexterous hand to the server to achieve precise control of the dexterous hand.

[0143] This invention proposes an innovative dexterous hand teleoperation system based on ultrasound imaging. By acquiring ultrasound animations of arm muscles and kinematic data streams of the hand, and utilizing imitation learning and IRL methods, it achieves precise control of the dexterous hand. This system not only overcomes the shortcomings of existing signal acquisition technologies but also provides a more personalized and efficient dexterous hand control solution, showing broad application prospects.

[0144] The application of an ultrasound-based dexterous hand teleoperation system in medical rehabilitation can be summarized in the following technical details: The system uses non-invasive ultrasound imaging as its core, combined with deep learning and a real-time control framework, to provide a high-precision, low-latency dexterous hand control solution for upper limb amputees. In terms of hardware, the ultrasound imaging module uses a high-frequency linear array probe (center frequency 5-10MHz) to focus on the dynamics of the forearm flexor muscles (FDS / FDP / FPL) and bones (ulna, radius). It captures micro-deformations of muscles through B-mode ultrasound, with a frame rate of at least 120Hz and a resolution of 0.1mm. The probe size is adapted to the mid-forearm (diameter 2-3cm) to ensure signal stability and comfort. Ultrasound data is transmitted in real-time to the control terminal via a portable device (such as the GE Logiq E10). Simultaneously acquired hand kinematic data comes from the Manus Quantum Metagloves glove, which supports 6-DOF spatial positioning at 21 key points (sampling rate 120Hz) and is aligned with the ultrasound signal via ROS timestamps, with an error controlled within 5ms. Collection of training datasets, such as Figure 2 As shown, the ultrasound probe and data glove simultaneously collect ultrasound images and hand joint information, constructing a "state-action" dataset through time step alignment. The actuator is a dexterous hand equipped with 12 servo motors, covering a torque range of 0.1-5 N·m, and integrates a force feedback sensor (0.1 N resolution) at the fingertips to achieve fine tactile feedback.

[0145] At the software algorithm level, the system constructs a reward function based on IRL, optimizing policy differences by maximizing the reward value of the expert action dataset. During training, maximum causal entropy regularization is employed to encourage the agent to explore diverse action paths. The formula is designed as follows: first, the reward function R is learned through IRL(πE) to maximize the difference between the expert policy πE and the reinforcement policy π; second, forward reinforcement learning with entropy regularization is performed through RL(R), introducing a regularization term into the entropy H(π(·|s)) of the policy distribution to balance exploration and exploitation. The dataset contains a large number of random hand motion samples (the bending direction, speed, and angle of each finger are randomly generated), and each frame of data corresponds to the spatial coordinate parameters of 21 key points. The training model uses a lightweight convolutional neural network (such as MobileNet) to process ultrasound image features and combines a temporal model (such as Transformer) to capture the dynamic temporal correlation of muscle movements. To reduce the influence of individual differences, the system introduces a transfer learning mechanism. The pre-trained model is fine-tuned to adapt to different users' muscle structures, and an adaptive filtering algorithm (such as wavelet denoising) is used to eliminate bone and fat tissue interference in the ultrasound signal.

[0146] In the real-time control phase, the ultrasonic signal is preprocessed in hardware (noise reduction, motion compensation) and then input into the neural network. Finally, it passes through a fully connected layer to output corresponding hand movement commands, which are then precisely controlled by a PD controller. Control latency is optimized end-to-end using edge computing devices (such as NVIDIA Jetson AGX Orin), reducing inference time to less than 5ms. The data acquisition, inference, and control flowchart is shown below. Figure 3 As shown, the system uses a dual-thread architecture to synchronously process data acquisition and model inference, ensuring the real-time and continuous nature of action commands. Furthermore, this technology can be extended to industrial collaboration scenarios, using high-precision force feedback to control robotic arms to complete precision assembly tasks, or to achieve natural gesture interaction control in the smart home field. Specific Implementation Example Nine:

[0148] The difference between Embodiment Nine and Embodiment Eight of the present invention lies only in:

[0149] The present invention provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement a dexterous hand teleoperation method based on ultrasound imaging. Specific Implementation Example 10:

[0151] The only difference between Embodiment 10 and Embodiment 9 of the present invention is that:

[0152] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement a dexterous hand teleoperation method based on ultrasound imaging. Specific Implementation Example Eleven:

[0154] The only difference between Embodiment Eleven and Embodiment Ten of this invention is that:

[0155] Specifically, the model analyzes ultrasound animations of arm muscles to determine and control the movements of the dexterous hand. The flowchart is as follows: Figure 1 As shown.

[0156] This invention first requires ultrasound imaging of the human arm to acquire ultrasound animations of the arm muscles and simultaneously record the kinematic data stream of the hand. The acquisition location for the arm muscle ultrasound animations is selected near the midpoint of the forearm. Muscle movement information in this area can accurately reflect the movement of the arm, and the arm width at this location is suitable for the size of most ultrasound probes, facilitating probe placement and operation. To ensure the validity of the acquired data, the ultrasound imaging modes include not only the common B-mode but also 3D, 4D, and other modes. Crucially, the imaging information must clearly display the muscle and skeletal structures of the forearm, including the superficial flexor muscles (FDS), deep flexor muscles (FDP), flexor pollicis (FPL), ulna (ulna), and radius (radius).

[0157] In acquiring hand kinematics data streams, besides obtaining finger movement information through data gloves, RGB / RGB-D sensors can also be used for gesture recognition. The spatial motion information of key points includes the coordinates of 21 key points on the hand, specifically: one key point on the wrist, and four key points for each of the thumb, index finger, middle finger, ring finger, and little finger. To ensure the acquired dataset has good representativeness and generalization ability, each frame of hand motion data undergoes random bending and stretching at different directions, angles, and speeds. Ultimately, each frame of acquired data constitutes a state space containing 21 parameters, comprehensively recording the motion states of each joint and fingertip of the hand.

[0158] To enhance the model's generalization and adaptability, the collected dataset needs to cover a variety of hand movement models, ensuring that the movements of each finger can be recorded and learned under different circumstances. The goal is to enable the trained model to make precise control responses when faced with hand movements from different people.

[0159] This invention employs imitation learning as the primary learning method. Imitation learning allows the agent to quickly learn how to perform tasks by observing expert demonstrations. Unlike traditional reinforcement learning, imitation learning does not require exploring the action space from scratch; instead, it accelerates the training process by learning expert behavior. Therefore, based on existing expert system demonstrations, imitation learning can significantly improve the learning efficiency of the algorithm.

[0160] During training, this invention uses IRL to optimize the model's learning performance. IRL is a method that infers a reward function by observing expert behavior. In this method, the system does not directly know the reward function but infers it by observing the expert's action trajectories, thereby obtaining the optimal policy. In this invention, the core idea of ​​IRL is to maximize policy optimization by selecting a reward function R, so that the agent minimizes the difference from the optimal policy πE at each step. Specifically, the goal of IRL is to maximize the reward function obtained from the expert demonstration dataset (i.e., the collected expert action data).

[0161] Within the IRL framework, the first step is to learn a suitable reward function, the formula of which is:

[0162]

[0163] Where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S is the state space, s is each independent state, A is the action space, and a is each independent action.

[0164] By maximizing the reward function in the expert demonstration dataset, the agent can learn the expert's action patterns and then imitate their behavior. The next step is to further optimize the learning strategy using maximum causal entropy regularization. The formula is:

[0165]

[0166] Where H(x) is the information entropy of the random variable x.

[0167] The algorithm consists of two steps. The first step is IRL(πE), which learns a reward function to maximize the reward difference between the expert policy and the reinforcement learning policy. The second step is RL(R), which performs positive reinforcement learning based on entropy regularization. Here, the entropy H(π(·|s)) is the entropy function of the policy distribution in a given state. The higher the entropy, the more uncertain the agent's policy is, and the closer it is to the expert's decision.

[0168] Through this process, the intelligent agent can make decisions consistent with expert behavior when faced with each frame of ultrasound animation input, thereby controlling the movements of the dexterous hand. This method does not rely on traditional electromyography signals or other invasive signal acquisition methods; it can infer complex hand movements solely through non-invasive ultrasound animation, thus providing a more precise and efficient dexterous hand control solution for amputees.

[0169] The ultrasound imaging dexterous hand teleoperation system of this invention has many significant advantages. First, ultrasound imaging technology is a non-invasive signal acquisition method, which avoids invasive interference to the human body and reduces the physical burden on patients compared to traditional electromyography (EMG) signal acquisition. Second, through the combination of imitation learning and IRL, the system can effectively improve learning efficiency, enabling the dexterous hand to accurately mimic complex hand movements. By learning from a large amount of hand movement data, the system can adapt to individual differences among patients and provide highly personalized dexterous hand control solutions.

[0170] This system is particularly suitable for patients with hand amputations. By deploying a trained model into the dexterous hand, patients can achieve flexible control of the dexterous hand through ultrasound animation of the arm muscles, enabling them to perform complex hand movements such as grasping and pinching, thereby improving their quality of life.

[0171] The above description is merely a preferred embodiment of a dexterous hand teleoperation method and system based on ultrasound imaging. The scope of protection for such a method and system is not limited to the above embodiments; all technical solutions falling within this conceptual framework are within the protection scope of this invention. It should be noted that for those skilled in the art, any improvements and variations made without departing from the principles of this invention should also be considered within the protection scope of this invention.

Claims

1. A dexterous hand telemanipulation method based on ultrasound imaging, characterized in that: The method includes the following steps: Step 1: Perform ultrasound imaging on the human arm, acquire ultrasound animation of the arm muscles, and record the kinematic data stream of the hand; Step 2: Based on the recorded data stream, build a neural network and learn the model's reward function through inverse reinforcement learning. Step 2 specifically involves: The first N frames of arm ultrasound images are used to perform inter-frame difference calculations, and the results of the first few differences are then weighted and summed to extract muscle movement change features. The calculation formula is as follows: in,( x , y ) represents pixel coordinates. I t ( x , y ) indicates the current t Time frame, I t-i ( x , y ) indicates the preceding i frame, N This indicates the number of frames that require differential calculation. w i Represents the weighting coefficients, satisfying , D t ( x , y ) indicates the current t The final weighted difference result at each moment; By calculating weighted difference data and hand movement data, a TransFormer network conforming to its data dimensions is constructed. The input data is... D t-i ( x , y ), i [0, M ], each D As an input feature, M For hyperparameters; The reward function is optimized using IRL, which aims to maximize the reward function obtained from the expert demonstration dataset, i.e., the collected expert action data. Within the IRL framework, this is achieved by maximizing the reward function between expert actions and other possible actions. Q Based on the difference in values, we can learn to optimize the specific reward function. The optimization formula is as follows: in, R Represents the reward function, π Representation strategy, Q Indicates an estimate of the reward. E Indicates optimal. S Representing the state space, s This represents each independent state in the state space. A Represents the action space. a This represents each independent action in the action space; The reward function is designed as follows: in, R sync This indicates a reward for muscle movement timing calibration. R stable This indicates a reward for muscle synergistic stability. R eff This indicates a composite reward for action performance; Specifically: in, D ultra This represents the ultrasound frame difference sequence, indicating the deformation rate of different muscles; D hand This represents the angular velocity sequence of the hand joints, indicating the motion velocity of each joint, DTW ( () indicates the dynamic time warping algorithm, used to align the delay differences between two sets of time series data. σ This indicates the time calibration tolerance parameter; This reward uses the Dynamic Time Warping (DTW) algorithm to quantify the physiological delay matching between muscle activation and joint movement. When the expert grasps an object, the ultrasound differential signal of the superficial flexor digitorum muscles... D ultra Earlier than finger joint bending movements D hand The reward function, through exponential decay, penalizes strategies that do not conform to this physiological law, ensuring the temporal consistency between dexterity hand movements and the human kinetic chain. Where, Δ U t = U t - U t-1 The difference matrix represents the difference between adjacent ultrasound frames, and the instantaneous change in muscle activation intensity is represented by the matrix. W Represents the muscle synergy weight matrix, ∥ · ∥ F The Frobenius norm of a matrix is ​​used to measure the overall magnitude of change in the matrix. To represent a very small quantity, take 10e-5 to prevent the denominator from being zero; The reward criteria assess the synergistic stability of antagonistic muscle groups using the F-norm of the matrix. In a normal fist-clenching motion, flexor activation should be accompanied by extensor inhibition, Δ U t The corresponding elements in the matrix have opposite signs; if both are enhanced simultaneously, it will lead to stiff movements; weight matrix W Amplify the abnormal synergistic signals of antagonistic muscle pairs, and quantify the overall coordination of the muscle group into a reward value of 0-1 through normalization processing: in, v h , v g This represents the current velocity vector of the hand's end effector and the target direction vector. This is the vector of the maximum angular acceleration of the hand joint. a h Cov( represents the angular acceleration vector of the hand joints) D ultra ) represents the determinant of the covariance matrix of the inter-frame differences of all muscle channels, Var( D ultra () represents the sum of variances of the inter-frame differences for all muscle channels. α , β , γ This indicates that the parameters are adaptively adjusted and decay exponentially as the number of training epochs increases; By maximizing the reward function in the expert demonstration dataset, the agent can learn the expert's action patterns and then imitate the expert's behavior. The training model uses a lightweight convolutional neural network to process ultrasound image features and combines a temporal model to capture the dynamic temporal correlation of muscle movement. Step 3: Perform positive reinforcement learning using the learned reward function to learn the optimal policy of the model; Step 3 specifically involves: Optimize the learning strategy using maximum causal entropy regularization: in, H ( x ) represents a random variable x Information entropy Indicates the state s Next action a , IRL represents the average reward an expert receives when performing a series of actions according to their strategy πE. π E RL (Reinforcement Learning) maximizes the reward difference between expert policies and reinforcement learning policies by learning a reward function. R This indicates that positive reinforcement learning is performed based on entropy regularization, where entropy... H ( π (·| s )) is the entropy function of the policy distribution under a given state; Step 4: Deploy the optimized model to the server, connect the ultrasound system and the dexterous hand to the server, and achieve precise control of the dexterous hand.

2. The method according to claim 1, characterized in that: Step 1 specifically involves: The ultrasound imaging of the arm muscles was taken at the midpoint of the forearm, and the imaging information needed to clearly show the superficial and deep flexor muscles of the forearm, the thumb flexor, the ulna, the radius, and the muscle and bone structures.

3. The method according to claim 2, characterized in that: In terms of acquiring hand kinematics data streams, in addition to obtaining finger movement information through data gloves, RGB / RGB-D sensors can also be used for gesture recognition. The spatial motion information of key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point of the wrist, and 4 key points of each finger of the thumb, index finger, middle finger, ring finger and little finger. Each frame of data collected constitutes a state space containing 21 parameters, comprehensively recording the movement state of each joint and fingertip of the hand.

4. The method according to claim 3, characterized in that: AND L( π E Maximize the reward difference between expert strategies and reinforcement learning strategies by learning a reward function; RL( R This indicates that positive reinforcement learning is performed based on entropy regularization, where entropy... H ( π (·| s The entropy function is the policy distribution under a given state. The higher the entropy, the more uncertain the agent's policy is, and the closer it is to the expert's decision. The parameters are updated using the policy gradient, which is formulated as follows: in, θ Represents network parameters, J This represents the policy optimization objective function. α The entropy regularization coefficient represents a hyperparameter that balances reward and uncertainty, and is adaptively adjusted. α To determine the degree of determinism of the dynamic control strategy; The parameters are updated using gradient ascent via the policy gradient formula, which is: in, η Indicates the learning rate. The parameter is θ strategy, Indicates the parameter θ Find the gradient. a t Indicates time t The action performed s t Indicates time t state, Indicates the state s t Choose action under conditions a t .

5. The method according to claim 4, characterized in that: Step 4 specifically involves: The ultrasonic signal is preprocessed by hardware for noise reduction and motion compensation before being input into the neural network. Finally, it passes through a fully connected layer to output corresponding hand movement commands. The PD controller is used to precisely control the movement of the dexterous hand. The control latency is optimized end-to-end through edge computing devices, and the inference time is compressed to less than 5 milliseconds.

6. A dexterous hand teleoperation system based on ultrasound imaging, the system operating based on the dexterous hand teleoperation method based on ultrasound imaging according to claim 1, characterized in that: The system includes: The data acquisition module performs ultrasound imaging on the human arm, acquires ultrasound animations of the arm muscles, and records the kinematic data stream of the hand. The model building module builds a neural network based on the recorded data stream and learns the model's reward function through inverse reinforcement learning. An optimization module learns the optimal policy of the model by performing positive reinforcement learning through the learned reward function; The control module deploys the optimized model to the server and connects the ultrasound system and the dexterous hand to the server to achieve precise control of the dexterous hand.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method as claimed in any one of claims 1-5.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the method of any one of claims 1-5.

Citation Information

Patent Citations

  • Flexible ultrasonic sensor and arterial blood pressure detection method thereof

    CN112869773A

  • Robot autonomous ultrasonic scanning skill strategy generation method and device and storage medium

    CN114155940A

  • Remote operation space manipulator trajectory planning method based on deep reinforcement learning

    CN119115953A