Dexterous hand teleoperation method and system based on ultrasonic imaging

Through ultrasound imaging and reverse reinforcement learning, the accuracy and adaptability of signal acquisition in remote operation of smart hands is solved, and the precise control of smart hands is achieved, which is suitable for efficient operation of hand amputation patients.

CN120347752AActive Publication Date: 2025-07-22HARBIN INST OF TECH
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510699067.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-07-22
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

In the existing smart hand remote operation technology, the insufficient accuracy of signal acquisition, many noise interference, poor adaptability, large individual differences and invasive damage to the human body, affecting the operating performance and control accuracy of smart hand.

Method used

Using ultrasound imaging-based methods, ultrasound imaging of the arm muscles is collected, ultrasound animation and kinematic data flow is designed, and reward functions are designed to achieve accurate control of the dexterous hands.

Benefits of technology

It improves the accuracy and adaptability of signal acquisition, reduces noise interference, avoids the invasiveness of traditional electromyography signal acquisition, and provides a personalized and efficient and clever hand control solution, which is especially suitable for patients with hand amputation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120347752A_ABST
    Figure CN120347752A_ABST
Patent Text Reader

Abstract

The invention relates to a dexterous hand teleoperation method and system based on ultrasonic imaging. The invention relates to the field of ultrasonic imaging research and the technical field of dexterous hand teleoperation, and the method comprises the following steps: carrying out ultrasonic imaging on a human arm, collecting an ultrasonic motion graph of arm muscles, and recording a kinematics data stream of a hand; building a neural network based on the recorded data flow, and learning a reward function of the model through a reverse reinforcement learning method; forward reinforcement learning is carried out through the learned reward function, and the optimal strategy of the model is learned; and deploying the optimized model to a server, and connecting the ultrasonic system and the dexterous hand to the server to realize accurate teleoperation control of the dexterous hand. According to the method, the learning efficiency of the human hand movement information can be improved, and the invasive problem of traditional electromyographic signal acquisition can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of ultrasonic imaging research and the field of robotic dexterous hand teleoperation technology, and is a method and system for dexterous hand teleoperation based on ultrasonic imaging. Background Art

[0002] The dexterous hand teleoperation technology is an important research direction in the field of robotics in recent years. It combines advanced technologies such as robotics, brain-computer interface, electromyogram signal control, and artificial intelligence, aiming to provide a more natural and efficient dexterous hand operation experience for amputees. Compared with traditional prosthetics, dexterous hand prosthetics can simulate more complex hand movements and help users complete fine operation tasks, such as grasping and carrying objects. The introduction of teleoperation technology enables patients to control the dexterous hand through direct neural signals or electromyogram signals, thereby achieving higher-precision and more personalized motion control, and then improving the quality of life. Currently, the technical research and application of dexterous hand teleoperation mainly focus on aspects such as the perception feedback system, control strategies, and the accuracy of actuators.

[0003] The technical route of dexterous hand teleoperation mainly includes two core parts: signal acquisition and processing, and the implementation of a fine control system. In terms of signal acquisition and processing, electromyogram signals and brain-computer interface technologies are currently commonly used. Electromyogram signals achieve the basic control of the dexterous hand by capturing the muscle activities of the residual limb. This method has a relatively mature application background, is easy to operate, and does not require complex surgical intervention. The brain-computer interface technology, on the other hand, directly decodes the neural signals of the brain and converts the user's thought instructions into dexterous hand movements. This method has higher flexibility and naturalness. Especially for completely amputated patients, it can provide a control experience close to normal hand operations. Through these technologies, the dexterous hand can perform complex actions, such as independent control of multiple fingers, force adjustment, and precise grasping actions. At the same time, the feedback system is also an indispensable part of dexterous hand teleoperation. Since the operation of the dexterous hand needs to be highly coordinated with the human sensory system, integrating tactile and force sensors can provide real-time feedback information for users during the operation, helping them perceive the state of the dexterous hand and the interaction with objects, thereby improving the accuracy and comfort of the operation.

[0004] Although dexterous hand teleoperation technology has made some progress, it still faces several technical challenges. First, the accuracy and real-time performance of signal decoding are still key issues. Whether it is electromyographic signals or brain-computer interfaces, the signal acquisition process is often interfered by noise, resulting in low decoding accuracy, which in turn affects the operation performance of the dexterous hand. Especially when performing subtle movements, any slight error may lead to operation failure or inaccuracy. Even with advanced algorithms, the stability and real-time performance of the signal are still difficult to guarantee, especially when switching control modes quickly or performing complex tasks. Delays and instability have become obstacles to technological progress. In addition, individual differences in signals are also a problem. Each patient has different electromyographic signals or brain wave characteristics, which requires the system to be able to perform personalized adjustments and learning. Although some adaptive algorithms and machine learning methods have made progress in this field, how to enable dexterous hands to quickly adapt to individual differences in a shorter period of time is still a problem worthy of further study.

[0005] The current dexterous hand remote control technology based on ultrasonic sensors has demonstrated advantages such as non-invasive detection and deep muscle dynamic capture. However, in terms of motion analysis, existing solutions mostly rely on machine learning models to regress joint angles from ultrasonic signals, which has bottlenecks such as high dependence on model training data and significant delays in real-time motion response. Different from the traditional model, this solution creatively integrates the reference document A flexible ultrasonic sensor and its arterial blood pressure detection method (CN112869773A) to perform difference processing on the depth change over time of the anterior wall boundary and the posterior wall boundary of the human artery. The present invention obtains the dynamic deformation differential value between adjacent frames of the muscle cross-section ultrasonic image through this method, and at the same time selects the first N frames of differential data after the muscle activation trigger to construct a feature analysis window. The present invention can analyze the refined movement intention of the finger joints in real time based on the differential data of the flexor group and the extensor group. Summary of the invention

[0006] The present invention aims to solve some common problems in existing dexterous hand control in signal acquisition technology, including insufficient signal accuracy, more noise interference, poor adaptability, large individual differences, and possible invasive damage to the human body. To this end, the present invention proposes a dexterous hand remote operation method and system based on ultrasonic imaging, which can realize the control of the dexterous hand by collecting ultrasonic dynamic images of arm muscles and data streams of prosthetic hand movements. The system deploys the trained model into the dexterous hand system through the method of imitation learning, thereby realizing precise control of the dexterous hand.

[0007] The present invention provides the following technical solutions:

[0008] A dexterous hand teleoperation method and system based on ultrasonic imaging, the method comprising the following steps:

[0009] Step 1: Perform ultrasonic imaging on the human arm, collect ultrasonic kinetic images of the arm muscles, and record the kinematic data stream of the hand.

[0010] Step 2: Based on the recorded data stream, build a neural network, design a reward function model, and learn the reward function of the model through the inverse reinforcement learning method.

[0011] Step 3: Perform forward reinforcement learning through the learned reward function to learn the optimal strategy of the model.

[0012] Step 4: Deploy the optimized model to the server, connect the ultrasonic system and the dexterous hand to the server, and achieve precise control of the dexterous hand.

[0013] Preferably, Step 1 is specifically as follows:

[0014] The acquisition position of the ultrasonic kinetic image of the arm muscles is selected at the 1 / 2 position of the forearm, and the imaging information needs to clearly show the superficial flexor muscles, deep flexor muscles, flexor pollicis, ulna, radius muscles and bone structures of the forearm.

[0015] Preferably, in terms of the acquisition of the hand kinematic data stream, in addition to obtaining the motion information of the fingers through a data glove, an RGB / RGB-D sensor can also be used for gesture recognition. The spatial motion information of the key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point for the wrist, and 4 key points for each of the thumb, index finger, middle finger, ring finger and little finger.

[0016] Each frame of the collected data constitutes a state space containing 21 parameters, comprehensively recording the motion states of each joint and fingertip of the hand.

[0017] Preferably, Step 2 is specifically as follows:

[0018] Take the first N frames of arm ultrasonic image data for inter-frame difference calculation, and then perform weighted summation on the previous difference results to extract the muscle motion change characteristics. The calculation formula is:

[0019]

[0020] where (x,y) represents the pixel coordinate value, I t (x,y) represents the current (t moment) frame, I t-i (x,y) represents the previous i frames, N represents the number of frames required for difference calculation, which is a hyperparameter to be adjusted in this invention, w i represents the weight coefficient, satisfying D t (x,y) represents the final weighted difference result at the current (t moment).

[0021] Build a Transformer network that conforms to the data dimension based on the calculated weighted difference data and hand motion data. The input data is D t-i (x, y), i ∈ [0, M], where each D is an input feature and M is a hyperparameter;

[0022] The reward function is optimized by IRL. The goal of IRL is to maximize the reward function obtained from the expert demonstration dataset, that is, the collected expert action data. In the IRL framework, by maximizing the Q-value difference between expert actions and other possible actions, the specific reward function is learned and optimized. The optimization formula is as follows:

[0023]

[0024] where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S represents the state space, s represents each independent state in the state space, A represents the action space, and a represents each independent action in the action space;

[0025] The form of the reward function is designed as follows:

[0026] R = R sync + R stable + R eff

[0027] where, R sync represents the muscle action timing calibration reward, R stable represents the muscle coordination stability reward, R eff represents the action efficiency composite reward.

[0028] Specifically:

[0029]

[0030] where D ultra represents the ultrasonic inter-frame difference sequence, indicating the deformation speed of different muscles. D hand represents the hand joint angular velocity sequence, indicating the movement speed of each joint. DTW(·) represents the dynamic time warping algorithm, which is used to align the delay differences of two groups of time series data. σ represents the time calibration tolerance parameter.

[0031] This reward term quantifies the physiological delay matching degree between muscle activation and joint movement through the dynamic time warping (DTW) algorithm. For example, when an expert grasps an object, the ultrasonic differential signal of the flexor digitorum superficialis muscle (D ultra ) will be earlier than the finger joint bending action (D hand) This is the inherent delay characteristic of human neuromuscular conduction. The reward function imposes penalties on strategies that do not conform to this physiological law in the form of exponential decay, ensuring the temporal consistency between the dexterous hand movements and the human movement chain.

[0032]

[0033] Where ΔU t = U t - U t-1 represents the difference matrix between adjacent ultrasound frames, indicating the instantaneous change in muscle activation intensity. W represents the muscle synergy weight matrix, and ||·|| F represents the matrix Frobenius norm, which is used to measure the overall change amplitude of the matrix. ∈ represents a very small quantity (taking 10e - 5) to prevent the denominator from being zero.

[0034] This reward term evaluates the co - stability of antagonist muscle groups through the matrix F - norm. For example, in a normal fist - clenching action, flexor activation should be accompanied by extensor inhibition (the corresponding element signs in ΔU t are opposite). If both increase simultaneously, it will lead to a rigid movement. The weight matrix W amplifies the co - activation abnormal signals of antagonist muscle pairs, and quantifies the overall coordination of muscle groups into a reward value ranging from 0 to 1 through normalization processing.

[0035]

[0036] Where, v h and v g represent the current end - effector velocity vector of the hand and the target direction vector, a h represents the hand joint angular acceleration vector, Cov(D ultra ) represents the determinant of the covariance matrix of the inter - frame differences of all muscle channels, and Var(D ultra ) represents the sum of variances of the inter - frame differences of all muscle channels. α, β, and γ represent adaptively adjusted parameters, which decay exponentially as the number of training rounds increases.

[0037] This composite reward optimizes the action efficiency from three dimensions:

[0038] Trajectory direction reward (the first term): Guides the hand to move towards the target direction through the cosine similarity of the velocity vectors.

[0039] Motion economy reward (the second term): Inhibits high - frequency tremors, conforming to the principle of minimizing muscle energy (e.g., rapid jitter will cause ||a h || 2 to be too high).

[0040] Muscle group coordination reward (the third item): Covariance and variance ratio are used to measure the synchronization of muscle activation. For example, grasping actions require high covariance (synchronous contraction) of the finger flexor muscle groups, while negative covariance (antagonistic effect) should be presented between the flexor and extensor muscles.

[0041] By maximizing the reward function in the expert demonstration dataset, the agent can learn the action patterns of the expert and then imitate the expert's behavior;

[0042] The training model uses a lightweight convolutional neural network to process ultrasonic image features and combines a temporal model to capture the dynamic temporal correlations of muscle movements.

[0043] Preferably, step 3 is specifically:

[0044] Optimize the learning strategy through maximum causal entropy regularization:

[0045]

[0046] Among them, H(x) represents the information entropy of the random variable x.

[0047] Preferably, IRL(π E ) maximizes the reward difference between the expert policy and the reinforcement learning policy by learning the reward function;

[0048] RL(R) represents forward reinforcement learning based on entropy regularization. The entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the agent's policy is and the closer it is to the expert's decision.

[0049] Use the policy gradient to update the parameters. The policy gradient formula is:

[0050]

[0051] Among them, θ represents the network parameters, J represents the policy optimization objective function, and α represents the entropy regularization coefficient (a hyperparameter that balances the reward and uncertainty. The certainty degree of the policy is dynamically controlled by adaptively adjusting α).

[0052] Update the parameters through gradient ascent using the policy gradient formula. The parameter update formula is:

[0053]

[0054] Among them, η represents the learning rate.

[0055] Preferably, step 4 is specifically:

[0056] The ultrasonic signal is input into the neural network after being preprocessed by hardware for noise reduction and motion compensation, and finally passes through the fully connected layer to output the corresponding hand motion instructions. A PD controller is used to precisely control the motion of the dexterous hand, and the control delay is optimized end-to-end through an edge computing device, with the inference time compressed to within 5 milliseconds.

[0057] A dexterous hand teleoperation system based on ultrasonic imaging, the system comprising:

[0058] A data acquisition module, which performs ultrasonic imaging on the human arm, acquires ultrasonic kinetic images of the arm muscles, and records the kinematic data stream of the hand;

[0059] A model building module, which builds a neural network based on the recorded data stream and learns the reward function of the model through the inverse reinforcement learning method;

[0060] An optimization module, which performs positive reinforcement learning through the learned reward function to learn the optimal strategy of the model;

[0061] A control module, which deploys the optimized model to the server, connects the ultrasonic system and the dexterous hand to the server, and realizes the precise control of the dexterous hand.

[0062] A computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement a dexterous hand teleoperation method based on ultrasonic imaging.

[0063] A computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements a dexterous hand teleoperation method based on ultrasonic imaging when executing the computer program.

[0064] The present invention has the following beneficial effects:

[0065] The present invention aims to solve the problems of existing dexterous hands in signal acquisition accuracy, noise interference, poor adaptability, large individual differences, and invasiveness to the human body. The system realizes the precise control of the dexterous hand by collecting ultrasonic kinetic images of the arm muscles and the kinematic data stream of the hand, and combining the imitation learning and IRL methods. The ultrasonic kinetic image provides real-time motion signals by monitoring the motion of the forearm muscles, while the kinematic data stream of the hand collects the spatial motion information of 21 key points through a data glove or gesture recognition technology. These data are used to train the agent, which learns by imitating the actions of experts to optimize the control strategy of the dexterous hand. The system can not only improve the learning efficiency but also avoid the invasive problems of traditional electromyogram signal acquisition. Finally, the trained model can realize the complex control of the dexterous hand, especially suitable for hand amputees, providing a more personalized and efficient dexterous hand control solution, and having a wide application prospect. Description of the Drawings

[0066] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0067] Figure 1 It is shown as a schematic flowchart of the method of the present invention;

[0068] Figure 2 It is shown as a specific reward function flowchart of the present invention through expert data learning;

[0069] Figure 3 It is shown as a data acquisition, inference, and control flowchart of the present invention. Specific embodiments

[0070] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the drawings. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0071] The following describes the present invention in detail in conjunction with specific embodiments. Specific Embodiment 1:

[0073] According to Figures 1 to 3 As shown, the specific optimization technical solution adopted by the present invention to solve the above technical problems is: The present invention relates to a dexterous hand teleoperation method and system based on ultrasonic imaging.

[0074] The present invention provides a dexterous hand teleoperation method and system based on ultrasonic imaging. The method includes the following steps:

[0075] The method includes the following steps:

[0076] Step 1: Perform ultrasonic imaging on the human arm, collect the ultrasonic dynamic images of the arm muscles, and record the kinematic data stream of the hand;

[0077] Step 2: Based on the recorded data stream, build a neural network, and learn the reward function of the model through the inverse reinforcement learning method;

[0078] Step 3: Perform forward reinforcement learning through the learned reward function to learn the optimal strategy of the model;

[0079] Step 4: Deploy the optimized model to the server, connect the ultrasonic system and the dexterous hand to the server, and achieve precise control of the dexterous hand. Specific Embodiment 2:

[0081] The difference between Embodiment 2 and Embodiment 1 of this application is only that:

[0082] Step 1 is specifically:

[0083] The acquisition position of the ultrasonic echogram of the arm muscle is selected at the 1 / 2 position of the forearm, and the imaging information needs to clearly show the superficial flexor muscles, deep flexor muscles, flexor pollicis, ulna, radius muscles and bone structures of the forearm. Specific Embodiment 3:

[0085] The difference between Embodiment 3 and Embodiment 2 of this application is only that:

[0086] In the acquisition of the hand kinematic data stream, in addition to obtaining the finger movement information through the data glove, an RGB / RGB-D sensor can also be used for gesture recognition. The spatial movement information of the key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point for the wrist, and 4 key points for each of the thumb, index finger, middle finger, ring finger and little finger;

[0087] Each frame of data collected constitutes a state space containing 21 parameters, comprehensively recording the movement states of each joint and fingertip of the hand. Specific Embodiment 4:

[0089] The difference between Embodiment 4 and Embodiment 3 of this application is only that:

[0090] Step 2 is specifically:

[0091] Take the first N frames of arm ultrasonic image data for inter-frame difference calculation, and then perform weighted summation on the previous difference results to extract the muscle movement change characteristics. The calculation formula is:

[0092]

[0093] where (x,y) represents the pixel coordinate value, I t (x,y) represents the current (t moment) frame, I t-i (x,y) represents the previous i frame, N represents the number of frames for which difference calculation needs to be performed, and is a hyperparameter to be adjusted in this invention, w i represents the weight coefficient, satisfying D t (x,y) represents the final weighted difference result at the current (t moment).

[0094] Construct a Transformer network that conforms to the data dimension based on the calculated weighted difference data and hand motion data. The input data is D t-i (x, y), i ∈ [0, M], where each D is an input feature and M is a hyperparameter;

[0095] The reward function is optimized through IRL. The goal of IRL is to maximize the reward function obtained from the expert demonstration dataset, that is, the collected expert action data. In the IRL framework, by maximizing the Q-value difference between the expert actions and other possible actions, the specific reward function is learned and optimized. The optimization formula is as follows:

[0096]

[0097] where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S represents the state space, s represents each independent state in the state space, A represents the action space, and a represents each independent action in the action space;

[0098] The form of the reward function is designed as follows:

[0099] R = R sync + R stable + R eff

[0100] where, R sync represents the muscle action timing calibration reward, R stable represents the muscle coordination stability reward, R eff represents the action efficiency composite reward.

[0101] Specifically:

[0102]

[0103] where D ultra represents the ultrasonic inter-frame difference sequence, indicating the deformation speed of different muscles. D hand represents the hand joint angular velocity sequence, indicating the movement speed of each joint. DTW(·) represents the dynamic time warping algorithm, which is used to align the delay differences of two groups of time series data. σ represents the time calibration tolerance parameter.

[0104] This reward term quantifies the physiological delay matching degree between muscle activation and joint movement through the dynamic time warping (DTW) algorithm. For example, when an expert grasps an object, the ultrasonic differential signal of the flexor digitorum superficialis (D ultra ) will be earlier than the finger joint bending action (D hand ), which is an inherent delay characteristic of human neuromuscular conduction. The reward function uses an exponential decay form to impose penalties on strategies that do not conform to this physiological law, ensuring the temporal consistency between the dexterous hand actions and the human motion chain.

[0105]

[0106] where ΔU t = U t - U t-1 represents the difference matrix of adjacent ultrasound frames and represents the instantaneous change in muscle activation intensity. W represents the muscle co - activation weight matrix, and ||·|| F represents the matrix Frobenius norm, which is used to measure the overall change amplitude of the matrix. ∈ represents a very small quantity (taking 10e - 5) to prevent the denominator from being zero.

[0107] This reward term evaluates the co - activation stability of antagonist muscle groups through the matrix F - norm. For example, in a normal fist - clenching movement, flexor activation should be accompanied by extensor inhibition (the signs of the corresponding elements in ΔU t are opposite). If both increase simultaneously, it will lead to a rigid movement. The weight matrix W amplifies the co - activation abnormality signal of antagonist muscle pairs, and quantifies the overall coordination of muscle groups into a reward value ranging from 0 to 1 through normalization.

[0108]

[0109] where, v h and v g represent the current hand end - effector velocity vector and the target direction vector, a h represents the hand joint angular acceleration vector, Cov(D ultra ) represents the determinant of the covariance matrix of the inter - frame differences of all muscle channels, and Var(D ultra ) represents the sum of variances of the inter - frame differences of all muscle channels. α, β, and γ represent adaptively adjusted parameters that decay exponentially as the number of training epochs increases.

[0110] This composite reward optimizes movement efficiency from three dimensions:

[0111] Trajectory direction reward (the first term): Guides the hand to move towards the target direction through the cosine similarity of the velocity vectors.

[0112] Movement economy reward (the second term): Suppresses high - frequency tremors, in line with the principle of minimizing muscle energy (such as rapid jitter will cause ||a h || 2 to be too high).

[0113] Muscle group coordination reward (the third term): The covariance - to - variance ratio measures the synchronization of muscle activation. For example, a grasping movement requires high covariance (synchronous contraction) of the finger flexor muscle groups, while there should be negative covariance (antagonistic effect) between the flexor and extensor muscles.

[0114] By maximizing the reward function in the expert demonstration dataset, the agent can learn the action patterns of the expert and then imitate the expert's behavior;

[0115] The training model uses a lightweight convolutional neural network to process ultrasonic image features and combines a temporal model to capture the dynamic temporal correlations of muscle movements. Specific Embodiment Five:

[0117] The difference between Embodiment Five and Embodiment Four of the present invention is only that:

[0118] Step 3 is specifically:

[0119] Optimize the learning strategy through maximum causal entropy regularization:

[0120]

[0121] where H(x) represents the information entropy of the random variable x. Specific Embodiment Six:

[0123] The difference between Embodiment Six and Embodiment Five of the present invention is only that:

[0124] IRL(π E ) maximizes the reward difference between the expert policy and the reinforcement learning policy by learning the reward function;

[0125] RL(R) represents forward reinforcement learning based on entropy regularization. The entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the agent's policy and the closer it is to the expert's decision.

[0126] Update the parameters using the policy gradient. The policy gradient formula is:

[0127]

[0128] where θ represents the network parameters, J represents the policy optimization objective function, and α represents the entropy regularization coefficient (a hyperparameter that balances the reward and uncertainty, and dynamically controls the certainty degree of the policy by adaptively adjusting α).

[0129] Update the parameters by performing gradient ascent through the policy gradient formula. The parameter update formula is:

[0130]

[0131] where η represents the learning rate. Specific Embodiment Seven:

[0133] The difference between Embodiment Seven and Embodiment Six of the present invention is only that:

[0134] Step 4 is specifically:

[0135] The ultrasonic signal is input into the neural network after being preprocessed by hardware for noise reduction and motion compensation, and finally the corresponding hand motion instructions are output through the fully connected layer. A PD controller is used to precisely control the motion of the dexterous hand, and the control delay is optimized end-to-end through an edge computing device, with the inference time compressed to within 5 milliseconds. Specific Embodiment Eight:

[0137] The difference between Embodiment Eight and Embodiment Seven of the present invention lies only in:

[0138] The present invention provides a dexterous hand teleoperation system based on ultrasonic imaging, characterized in that: the system includes:

[0139] A data acquisition module, which performs ultrasonic imaging on the human arm, acquires ultrasonic motion images of the arm muscles, and records the kinematic data stream of the hand;

[0140] A model building module, which builds a neural network based on the recorded data stream and learns the reward function of the model through the method of inverse reinforcement learning;

[0141] An optimization module, which performs positive reinforcement learning through the learned reward function to learn the optimal strategy of the model;

[0142] A control module, which deploys the optimized model to the server, connects the ultrasonic system and the dexterous hand to the server, and realizes the precise control of the dexterous hand.

[0143] The present invention proposes an innovative dexterous hand teleoperation system based on ultrasonic imaging. By acquiring ultrasonic motion images of the arm muscles and the kinematic data stream of the hand, and using the methods of imitation learning and IRL, the precise control of the dexterous hand is realized. This system can not only overcome the various deficiencies of the existing signal acquisition technologies, but also provide a more personalized and efficient dexterous hand control solution, with broad application prospects.

[0144] The application of the dexterous hand teleoperation system based on ultrasonic imaging in medical rehabilitation can be elaborated around the following technical details: The system takes non-invasive ultrasonic imaging as the core, combines deep imitation learning and real-time control framework, and provides a high-precision and low-latency dexterous hand control solution for upper limb amputees. In terms of hardware, the ultrasonic imaging module uses a high-frequency linear array probe (center frequency 5-10 MHz), focuses on the dynamics of the forearm flexor muscles (FDS / FDP / FPL) and bones (ulna, radius), captures the micro-deformation of muscles through the B-mode ultrasonic mode, with a frame rate of not less than 120 Hz and a resolution of 0.1 mm. The probe size is adapted to the middle section of the forearm (diameter 2-3 cm) to ensure signal stability and comfort. The ultrasonic data is transmitted to the control terminal in real time through a portable device (such as GE Logiq E10), and the simultaneously collected hand kinematic data comes from the Manus Quantum Metagloves glove, which supports 6-degree-of-freedom spatial positioning of 21 key points (sampling rate 120 Hz) and is aligned with the ultrasonic signal through the ROS timestamp, with the error controlled within 5 ms. The collection of the training dataset is as Figure 2 shown. The ultrasonic probe and the data glove collect ultrasonic images and hand joint information simultaneously, and construct a "state-action" dataset through time step alignment. The actuator is a dexterous hand equipped with 12 servos, with a torque range covering 0.1-5 N·m, and a force feedback sensor (resolution 0.1 N) is integrated at the fingertip to achieve fine tactile feedback.

[0145] At the software algorithm level, the system constructs a reward function based on IRL, and optimizes the policy difference by maximizing the reward value of the expert action dataset. During the training process, maximum causal entropy regularization is adopted to encourage the agent to explore diverse action paths. The formula is designed as follows: First, learn the reward function R through IRL(πE) to maximize the difference between the expert policy πE and the reinforcement policy π; Second, perform forward reinforcement learning with entropy regularization through RL(R), and introduce a regularization term for the entropy H(π(·|s)) of the policy distribution to balance exploration and exploitation. The dataset contains a large number of random hand movement samples (the bending direction, speed, and angle of each finger are randomly generated), and each frame of data corresponds to the spatial coordinate parameters of 21 key points. The training model uses a lightweight convolutional neural network (such as MobileNet) to process the ultrasonic image features, and combines a temporal model (such as Transformer) to capture the dynamic temporal correlation of muscle movements. To reduce the influence of individual differences, the system introduces a transfer learning mechanism, and the pre-trained model is fine-tuned to adapt to the muscle structures of different users, and an adaptive filtering algorithm (such as wavelet denoising) is used to eliminate the interference of bone and fat tissues in the ultrasonic signal.

[0146] In the real-time control loop, the ultrasonic signal is preprocessed by hardware (noise reduction, motion compensation) and then input into the neural network. Finally, it passes through the fully connected layer to output the corresponding hand motion commands, and a PD controller is used to precisely control the dexterous hand. The control delay is optimized end-to-end through an edge computing device (such as NVIDIA Jetson AGX Orin), and the inference time is compressed to within 5 ms. The data acquisition, inference, and control flow chart is as shown in Figure 3 shown. The system synchronously processes data acquisition and model inference through a dual-thread architecture to ensure the real-time and continuous nature of the motion commands. In addition, this technology can be extended to industrial collaboration scenarios to complete precise assembly tasks by controlling a robotic arm through high-precision force feedback, or to achieve natural gesture interaction control in the field of smart home. Specific Embodiment Nine:

[0148] The difference between the ninth embodiment of the present invention and the eighth embodiment is only that:

[0149] The present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement a dexterous hand teleoperation method based on ultrasonic imaging. Specific Embodiment Ten:

[0151] The difference between the tenth embodiment of the present invention and the ninth embodiment is only that:

[0152] The present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements a dexterous hand teleoperation method based on ultrasonic imaging. Specific Embodiment Eleven:

[0154] The difference between the eleventh embodiment of the present invention and the tenth embodiment is only that:

[0155] Specifically, the model analyzes the ultrasonic motion images of the arm muscles, and then judges and controls the actions of the dexterous hand. The flow chart is as shown in Figure 1 shown.

[0156] The present invention first requires ultrasonic imaging of the human arm to collect ultrasonic kinematic images of the arm muscles and simultaneously record the kinematic data stream of the hand. The acquisition position of the ultrasonic kinematic images of the arm muscles is selected near the midpoint of the forearm. The muscle movement information at this location can accurately reflect the movement of the arm, and the width of the arm at this position is suitable for the sizes of most ultrasonic probes, facilitating the placement and operation of the ultrasonic probes. To ensure the effectiveness of the collected data, the ultrasonic imaging mode not only includes the common B mode but also can adopt different modes such as 3D and 4D. However, importantly, the imaging information needs to clearly display muscle and bone structures such as the FDS (flexor digitorum superficialis), FDP (flexor digitorum profundus), FPL (flexor pollicis longus), Ulna, and Radius of the forearm.

[0157] In terms of the acquisition of the hand kinematic data stream, in addition to obtaining finger movement information through a data glove, RGB / RGB-D sensors can also be used for gesture recognition. The spatial movement information of key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point for the wrist and 4 key points for each of the thumb, index finger, middle finger, ring finger, and little finger. To ensure that the collected data set has good representativeness and generalization ability, each frame of data of the hand movement is randomly bent and stretched in different directions, angles, and speeds. Finally, each frame of the collected data constitutes a state space containing 21 parameters, comprehensively recording the movement states of each joint and fingertip of the hand.

[0158] To enhance the generalization ability and adaptability of the model, the collected data set needs to cover a variety of hand movement models to ensure that the movement of each finger can be recorded and learned under different circumstances. The purpose of doing this is to enable the trained model to still make accurate control responses when facing the hand movements of different people.

[0159] The present invention adopts imitation learning as the main learning method. Imitation learning enables the agent to quickly learn how to perform tasks by observing the operations demonstrated by experts. Different from traditional reinforcement learning, imitation learning does not need to explore the action space from scratch but accelerates the training process by learning expert behaviors. Therefore, based on the existing expert system demonstrations, imitation learning can greatly improve the learning efficiency of the algorithm.

[0160] During the training process, the present invention uses IRL to optimize the learning effect of the model. IRL is a method of inferring the reward function by observing expert behavior. In this method, the system does not directly know the reward function, but infers the reward function by observing the action trajectory of the expert, so as to obtain the optimal policy. In the present invention, the core idea of IRL is to maximize the optimization of the policy by selecting the reward function R, so that when the intelligent agent makes a choice at each step, it will try to minimize the difference from the optimal policy πE. Specifically, the goal of IRL is to maximize the reward function obtained from the expert demonstration dataset (i.e., the collected expert action data).

[0161] Under the IRL framework, first, a suitable reward function needs to be learned, and the formula of this reward function is:

[0162]

[0163] where R represents the reward function, π represents the policy, Q represents the estimate of the reward, E represents the optimal, S is the state space, s is each independent state, A is the action space, and a is each independent action.

[0164] By maximizing the reward function in the expert demonstration dataset, the intelligent agent can learn the action pattern of the expert and then imitate the behavior of the expert. The next step is to further optimize the learning strategy through maximum causal entropy regularization. Its formula is:

[0165]

[0166] where H(x) is the information entropy of the random variable x.

[0167] In this process, the algorithm is divided into two steps. The first step is IRL(πE), which maximizes the reward difference between the expert policy and the reinforcement learning policy by learning the reward function. The second step is RL(R), that is, forward reinforcement learning based on entropy regularization. Here, the entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the policy of the intelligent agent is and the closer it is to the decision of the expert.

[0168] Through this process, when facing the input of each frame of ultrasound echocardiogram, the intelligent agent can make decisions consistent with the behavior of the expert, and then control the actions of the dexterous hand. This method does not rely on traditional electromyogram signals or other invasive signal acquisition means, and can infer complex hand movements only through non-invasive ultrasound echocardiograms, thus providing a more accurate and efficient dexterous hand control scheme for hand amputee patients.

[0169] The ultrasonic imaging dexterous hand teleoperation system of the present invention has many remarkable advantages. First of all, ultrasonic imaging technology is a non-invasive signal acquisition method. Compared with traditional electromyogram signal acquisition, it can avoid invasive interference to the human body and reduce the physical burden on patients. Secondly, through the combination of imitation learning and IRL, the system can effectively improve the learning efficiency, enabling the dexterous hand to accurately imitate the complex movements of the human hand. By learning a large amount of hand movement data, the system can adapt to the individual differences of different patients and provide a highly personalized dexterous hand control solution.

[0170] This system is particularly suitable for hand amputee patients. By deploying the trained model into the dexterous hand, patients can achieve flexible control of the dexterous hand through the ultrasonic echogram of the arm muscles, and can perform complex hand movements such as grasping and pinching, thereby improving the quality of life of patients.

[0171] The above is only a preferred embodiment of a dexterous hand teleoperation method and system based on ultrasonic imaging. The protection scope of a dexterous hand teleoperation method and system based on ultrasonic imaging is not limited to the above embodiments. All technical solutions within this concept belong to the protection scope of the present invention. It should be pointed out that for those skilled in the art, several improvements and changes made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.

Claims

1. A dexterous hand teleoperation method based on ultrasonic imaging, characterized in that: The method includes the following steps: Step 1: Perform ultrasonic imaging on a human arm, collect ultrasonic kinematic images of the arm muscles, and record the kinematic data stream of the hand; Step 2: Based on the recorded data stream, build a neural network and learn the reward function of the model through the inverse reinforcement learning method; Step 3: Perform forward reinforcement learning through the learned reward function to learn the optimal strategy of the model; Step 4: Deploy the optimized model to the server, connect the ultrasonic system and the dexterous hand to the server, and achieve precise control of the dexterous hand.

2. The method according to claim 1, wherein: The specific content of Step 1 is as follows: The acquisition position of the ultrasonic kinematic image of the arm muscles is selected at the 1 / 2 position of the forearm, and the imaging information needs to clearly display the superficial flexor muscles, deep flexor muscles, flexor pollicis, ulna, radius muscles and bone structures of the forearm.

3. The method according to claim 2, wherein: In terms of the acquisition of the hand kinematic data stream, in addition to obtaining the motion information of the fingers through a data glove, an RGB / RGB-D sensor can also be used for gesture recognition. The spatial motion information of the key points includes the coordinate information of 21 key points of the hand, specifically including: 1 key point for the wrist, and 4 key points for each of the thumb, index finger, middle finger, ring finger and little finger; Each frame of the collected data constitutes a state space containing 21 parameters, comprehensively recording the motion states of each joint and fingertip of the hand.

4. The method according to claim 3, characterized in that: The specific content of Step 2 is as follows: Take the first N frames of arm ultrasonic image data for inter-frame difference calculation, and then perform weighted summation on the results of the previous differences to extract the muscle motion change characteristics. The calculation formula is: Among them, (x, y) represents the pixel coordinate value, I t (x, y) represents the frame at the current moment t, I t-i (x, y) represents the previous i frames, N represents the number of frames for which differential calculation needs to be performed, and is a hyperparameter to be adjusted in this invention, w i represents the weight coefficient, satisfying D t (x, y) represents the final weighted differential result at the current moment t; Construct a Transformer network that conforms to the data dimension based on the calculated weighted difference data and hand movement data. The input data is D t-i (x, y), i ∈ [0, M], where each D is an input feature and M is a hyperparameter; The reward function is optimized through IRL. The goal of IRL is to maximize the reward function obtained from the expert demonstration dataset, that is, the collected expert action data. In the IRL framework, by maximizing the Q-value difference between the expert action and other possible actions, the specific reward function is learned and optimized. The optimization formula is as follows: Where, R represents the reward function, π represents the strategy, Q represents the estimation of the reward, E represents the optimal, S represents the state space, s represents each independent state in the state space, A represents the action space, and a represents each independent action in the action space; The form of the reward function is designed as follows: R = R sync + R stable + R eff Among them, R sync represents the muscle movement timing calibration reward, R stable represents the muscle coordination stability reward, R eff represents the action efficiency composite reward; Specifically: Among them, D ultra represents the ultrasonic inter-frame difference sequence, indicating the deformation speed of different muscles; D hand represents the hand joint angular velocity sequence, indicating the movement speed of each joint. DTW(·) represents the dynamic time warping algorithm, which is used to align the delay differences of two sets of time series data. σ represents the time calibration tolerance parameter; This reward item quantifies the physiological delay matching degree between muscle activation and joint movement through the dynamic time warping (DTW) algorithm. When an expert grasps an object, the ultrasonic differential signal D of the flexor digitorum superficialis ultra is earlier than the finger joint bending motion D hand , and the reward function imposes a penalty on strategies that do not conform to this physiological law in the form of exponential decay to ensure the timing consistency between the dexterous hand movement and the human motion chain: where, ΔU t = U t - U t-1 represents the difference matrix between adjacent ultrasound frames, representing the instantaneous change in muscle activation intensity, W represents the muscle co - activation weight matrix, ||·|| F represents the matrix Frobenius norm, used to measure the overall change amplitude of the matrix, ∈ represents a very small quantity, taking 10e - 5 to prevent the denominator from being zero; The reward item evaluates the co - stability of antagonist muscle groups through the matrix F - norm. During normal fist - clenching movements, flexor activation should be accompanied by extensor inhibition, and the signs of the corresponding elements in ΔU t are opposite. If both increase simultaneously, it will lead to rigid movements; the weight matrix W amplifies the co - activation abnormal signals of antagonist muscle pairs, and quantifies the overall coordination of muscle groups into a reward value between 0 and 1 through normalization processing: where, v h and v g represent the current end - effector velocity vector and the target direction vector, a h represents the hand joint angular acceleration vector, Cov(D ultra ) represents the determinant of the covariance matrix of the inter - frame differences of all muscle channels, Var(D ultra ) represents the sum of variances of the inter - frame differences of all muscle channels, and α, β, γ represent the adaptively adjusted parameters, which decay in an exponential trend as the number of training rounds increases; By maximizing the reward function in the expert demonstration dataset, the agent can learn the action pattern of the expert and then imitate the behavior of the expert; The training model uses a lightweight convolutional neural network to process ultrasonic image features and combines a temporal model to capture the dynamic temporal correlation of muscle motion.

5. The method according to claim 4, characterized in that: The specific content of Step 3 is as follows: Optimize the learning strategy through maximum causal entropy regularization: Where, H(x) represents the information entropy of the random variable x.

6. The method according to claim 5, wherein: IRL(π E ) maximizes the reward difference between the expert policy and the reinforcement learning policy by learning the reward function; RL(R) represents forward reinforcement learning based on entropy regularization. The entropy H(π(·|s)) is the entropy function of the policy distribution under a given state. The higher the entropy, the more uncertain the agent's strategy and the closer it is to the expert's decision; Use the policy gradient to update the parameters. The policy gradient formula is: Among them, θ represents the network parameters, J represents the policy optimization objective function, α represents the entropy regularization coefficient, a hyperparameter that balances the reward and uncertainty, and the degree of determinacy of the policy is dynamically controlled by adaptively adjusting α; Gradient ascent is performed through the policy gradient formula to update the parameters, and the parameter update formula is: Among them, η represents the learning rate.

7. The method according to claim 6, characterized in that: The specific step 4 is as follows: The ultrasonic signal is input into the neural network after being preprocessed by hardware for noise reduction and motion compensation, and finally the corresponding hand motion instructions are output through the fully connected layer. The PD controller is used to precisely control the motion of the dexterous hand, and the control delay is optimized end-to-end through the edge computing device, and the inference time is compressed to within 5 milliseconds.

8. A dexterous hand teleoperation system based on ultrasonic imaging, characterized in that: The system includes: A data acquisition module that performs ultrasonic imaging on the human arm, acquires ultrasonic kinematic images of the arm muscles, and records the kinematic data stream of the hand; A model building module that builds a neural network based on the recorded data stream and learns the reward function of the model through the inverse reinforcement learning method; An optimization module that performs forward reinforcement learning through the learned reward function to learn the optimal policy of the model; A control module that deploys the optimized model to the server, connects the ultrasonic system and the dexterous hand to the server, and realizes the precise control of the dexterous hand.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the method as claimed in claims 1-7.

10. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the method as claimed in claims 1-7.

Citation Information

Patent Citations

  • Flexible ultrasonic sensor and arterial blood pressure detection method thereof

    CN112869773A

  • Robot autonomous ultrasonic scanning skill strategy generation method and device and storage medium

    CN114155940A

  • Remote operation space manipulator trajectory planning method based on deep reinforcement learning

    CN119115953A

  • Dexterous hand teleoperation method, device and system based on data glove

    CN121223791A

  • Method and system for autonomous therapy

    US20220387118A1