Hand rehabilitation synchronous training method and system based on CNN and LSTM network

By combining CNN, LSTM and self-attention mechanism motion prediction model, the problem of single training mode of existing hand rehabilitation robot is solved, and synchronous coordination and personalized training of affected hand and healthy hand are achieved.

CN120392477APending Publication Date: 2025-08-01CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510455350.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing hand rehabilitation robot training mode is single, lacking time information and motion trajectory optimization, resulting in poor control delay and synchronization, which cannot meet personalized needs.

Method used

Convolutional neural network (CNN) and long and short-term memory network (LSTM) are used to combine self-attention mechanism to build a motion prediction model, and predict synchronous training movements of the affected hand through the bending of the healthy hand and wrist movement data.

Benefits of technology

It improves the prediction accuracy and stability of hand rehabilitation training, reduces exercise delay, realizes synchronous coordination between the affected hand and the healthy hand, and enhances the personalized training effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120392477A_ABST
    Figure CN120392477A_ABST
Patent Text Reader

Abstract

The invention provides a hand rehabilitation synchronous training method and system based on a CNN and an LSTM network, and relates to the field of intelligent control, and the method comprises the steps: obtaining finger bending data of a healthy side hand and wrist motion data of the healthy side hand; constructing a motion prediction model through a convolutional neural network, a long-short term memory network and a self-attention mechanism, and training the motion prediction model; inputting the finger bending data and the wrist motion data into the trained motion prediction model to obtain a hand motion result; the affected-side hand driving device is controlled through the hand movement result, and synchronous rehabilitation training of the affected-side hand is achieved. According to the technical scheme, the fluency and naturalness of control can be effectively improved, the motion delay is remarkably reduced, and the synchronization performance of the system is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent control, and in particular, to a hand rehabilitation synchronous training method and system based on CNN and LSTM networks. Background Art

[0002] Most of the existing hand rehabilitation robots are passive training devices with a single training mode, which are difficult to meet the personalized needs of different patients. Although the traditional direct mapping control method (that is, directly mapping the bending sensor value to the bending degree of the pneumatic driving glove) is simple to implement, it has the following problems:

[0003] 1. Lack of time information: It is impossible to predict the next action based on the historical motion trend, resulting in poor control delay and synchronization.

[0004] 2. Lack of motion trajectory optimization: The direct mapping method may cause the movement of the affected hand to be rigid and uncoordinated, lacking a compensation mechanism for the natural motion trajectory.

[0005] The existing deep learning-based control methods can achieve end-to-end feature learning, can learn the continuous motion patterns of the patient's hand, rather than isolated single-frame states, and can also predict the upcoming actions of the healthy hand based on current and historical information, so as to drive the affected hand to complete the corresponding movement in advance. However, there are still problems of large motion delay and inaccurate action synchronization. Summary of the Invention

[0006] The purpose of the present invention is to provide a hand rehabilitation synchronous training method and system based on CNN and LSTM networks to solve the problem that the existing technology cannot perform finger synchronous training in real time and accurately.

[0007] The above object of this application is achieved by the following technical solutions:

[0008] S1: Obtain the finger bending data of the healthy hand and the wrist movement data of the healthy hand;

[0009] S2: Construct a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism and train it;

[0010] S3: Input the finger bending data and the wrist movement data into the trained motion prediction model to obtain the hand movement result;

[0011] S4: Control the driving device of the affected hand through the hand movement result to realize the synchronous rehabilitation training of the affected hand.

[0012] Optionally, step S1 includes:

[0013] The finger bending data is the finger bending data of 5 fingers of the healthy hand, and the finger bending data includes: finger bending angle data and timing information;

[0014] The wrist movement data includes: wrist attitude angle timing data, angular velocity timing data, and acceleration timing data.

[0015] Optionally, step S2 includes:

[0016] The motion prediction model includes: a finger feature extraction module, a wrist feature extraction module, a feature fusion module, and a fully connected layer;

[0017] The finger feature extraction module is connected to the feature fusion module; the wrist feature extraction module is connected to the feature fusion module;

[0018] The feature fusion module is connected to the fully connected layer.

[0019] Optionally, the finger feature extraction module includes: a CNN convolutional unit and a first LSTM unit;

[0020] Through the CNN convolutional unit, the spatial features of the finger angle change of the finger bending data are extracted;

[0021] The spatial features are input into the first LSTM unit to obtain the first dynamic features of the finger movement changing with time.

[0022] Optionally, the wrist feature extraction module uses an LSTM long short-term memory network;

[0023] Through the wrist feature extraction module, the second dynamic features of the overall movement of the wrist movement data are extracted.

[0024] Optionally, the feature fusion module uses a self-attention mechanism to fuse the first dynamic features and the second dynamic features;

[0025] The fused features pass through the fully connected layer to generate the final finger movement prediction value, that is, the hand movement result.

[0026] A hand rehabilitation synchronous training system based on a CNN and LSTM network, the system includes: a data acquisition module, a processing module, a driving module, a voice interaction module, and a display module;

[0027] The data acquisition module includes: a bending sensor and a nine-axis attitude sensor;

[0028] The bending sensor is used to collect the finger bending data of the healthy hand;

[0029] The nine-axis attitude sensor is used to collect the wrist movement data of the healthy hand; the wrist movement data includes attitude angle, angular velocity, and acceleration;

[0030] The processing module is used to construct and train a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism;

[0031] The processing module is used to input finger bending data and wrist movement data into the trained motion prediction model to obtain the hand movement result;

[0032] The driving module is used to control the driving setting of the affected hand through the hand movement result to realize the synchronous rehabilitation training of the affected hand.

[0033] Optionally, the voice interaction module adopts a speech recognition model based on the Transformer architecture;

[0034] The voice interaction module is used to conduct language training on the patient to promote the collaborative exercise between the brain and the fingers through speech.

[0035] Optionally, the driving device of the affected hand of the driving module is wrapped with a flexible rehabilitation glove.

[0036] The beneficial effects brought by the technical solution provided by this application are:

[0037] 1. Combining the convolutional neural network CNN and the long short-term memory network LSTM to process the data of the bending sensor. Introducing a CNN module at the front end of the model to extract the spatial structure features between fingers, and then further modeling its dynamic evolution in the time dimension through LSTM. This combination makes full use of the complementary advantages of the two types of networks in spatial and temporal modeling, thus significantly improving the accuracy and stability of prediction. It solves the technical problem that the traditional method relies on a single LSTM structure, which can only model the dependence relationship in the time dimension and ignores the collaborative characteristics between fingers in space.

[0038] 2. Introducing the self-attention mechanism, which enhances the model's ability to extract and fuse key information. When dealing with long sequence data, traditional LSTM is prone to problems such as gradient dissipation, resulting in the model being difficult to focus on the key node information in the early stage. The self-attention mechanism has the ability to model global dependence relationships. It can not only enhance the model's attention to key time segments, but also establish an effective dynamic association between local finger information and overall hand information, automatically identify the important corresponding relationships between the two types of information, and achieve the deep fusion of the overall hand movement and local action features, thereby further improving the model's understanding ability and prediction accuracy for complex hand movements. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The following will further illustrate this application in conjunction with the drawings. In the drawings:

[0040] Figure 1It is the flowchart in the embodiment of the present application;

[0041] Figure 2 It is the network structure diagram in the embodiment of the present application;

[0042] Figure 3 It is the comparison chart of predicted response curves in the embodiment of the present application. Detailed implementation manners

[0043] For a clearer understanding of the technical features, objectives, and effects of the present application, the detailed implementation manners of the present application will now be described in detail with reference to the accompanying drawings.

[0044] The embodiment of the present application provides a hand rehabilitation synchronous training method based on CNN and LSTM networks.

[0045] Please refer to Figure 1 , Figure 1 It is the flowchart of a hand rehabilitation synchronous training method based on CNN and LSTM networks in the embodiment of the present application, including:

[0046] S1: Obtain the finger bending data of the healthy hand and the wrist movement data of the healthy hand;

[0047] S2: Construct a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism and train it;

[0048] S3: Input the finger bending data and the wrist movement data into the trained motion prediction model to obtain the hand movement result;

[0049] S4: Control the driving device of the affected hand through the hand movement result to achieve synchronous rehabilitation training of the affected hand.

[0050] Step S1 includes:

[0051] The finger bending data is the finger bending data of 5 fingers of the healthy hand, and the finger bending data includes: finger bending angle data and timing information;

[0052] The wrist movement data includes: wrist attitude angle timing data, angular velocity timing data, and acceleration timing data.

[0053] Step S2 includes:

[0054] The motion prediction model includes: a finger feature extraction module, a wrist feature extraction module, a feature fusion module, and a fully connected layer;

[0055] The finger feature extraction module is connected to the feature fusion module; the wrist feature extraction module is connected to the feature fusion module; <\

[0056] The feature fusion module is connected to the fully connected layer.

[0057] The finger feature extraction module includes: a CNN convolutional unit and a first LSTM unit;

[0058] The CNN convolutional unit extracts the spatial features of the finger angle changes in the finger bending data;

[0059] The spatial features are input into the first LSTM unit to obtain the first dynamic features of the finger movement changing over time.

[0060] As an embodiment, the movement of the fingers is not completely independent and there is a synergistic effect. For example, in actions such as grasping and pinching, multiple fingers will bend or stretch simultaneously. Therefore, analyzing the data of each finger separately may not be able to fully capture the overall movement pattern. In addition, the movement of the fingers is a continuous time series, and relying only on single-frame data cannot predict the movement trend at the next moment. As Figure 2 shown, the CNN+LSTM combination method is used to process the finger bending data to simultaneously extract spatial features and time dependencies, improving the prediction accuracy.

[0061] The wrist feature extraction module uses an LSTM long short-term memory network;

[0062] The wrist feature extraction module extracts the second dynamic features of the overall movement of the wrist movement data.

[0063] As an embodiment, the nine-axis attitude sensor is mainly used to measure the overall movement state of the hand, including sequential data such as attitude angles (Roll, Pitch, Yaw), angular velocity, and acceleration. These data can reflect the overall movement trend of the hand, such as the rotation, tilt, or movement of the palm. Since these signals are time series data and the current hand state is affected by the historical state, we use the LSTM network to process this sequential information.

[0064] The feature fusion module uses a self-attention mechanism to fuse the first dynamic features and the second dynamic features;

[0065] The fused features pass through a fully connected layer to generate the final finger movement prediction value, that is, the hand movement result.

[0066] As an embodiment, in order to achieve the efficient fusion of the features extracted by the bending sensor (finger local information) and the nine-axis attitude sensor (hand overall information), a self-attention mechanism (Self-Attention) is introduced into the network structure for weighted fusion.

[0067] Represent the finger features extracted by the bending sensor as Query, which represents the independent movement patterns of each finger; while the wrist features extracted by the attitude sensor are represented as Key and Value, which are used to describe the overall movement trend of the hand. The core calculation of the self-attention mechanism is as follows:

[0068]

[0069] Among them, Q, K, and V are the Query, Key, and Value matrices respectively, and d k is the scaling factor of the vector dimension, which is used to stabilize the gradient. This formula calculates the similarity between Query and Key to obtain a set of attention weights, and applies them to Value to obtain the fused feature representation.

[0070] Through the self-attention mechanism, the network can dynamically judge the importance of the two types of information according to the current input. For example, in a scenario where finger movement is relatively autonomous and there is no significant coupling with the wrist, the network is more inclined to rely on the local features extracted by the "finger feature extraction module" CNN; while in a situation where finger movement is significantly affected by wrist drive, the network automatically enhances its attention to the overall dynamic features extracted by the "wrist feature extraction module". This flexible weight allocation ability enables the model to have stronger generalization ability and can adapt to a wider range of complex hand movement patterns.

[0071] As an embodiment, the bending sensor data processing module (CNN and LSTM networks): extracts the spatial and temporal features of the bending sensor to capture the correlation between fingers. The nine-axis attitude sensor data processing module (LSTM network): processes the timing information of the nine-axis attitude sensor (JY901S) to predict the movement trend. The fusion layer (Self-Attention + Fully Connected): fuses the features extracted by CNN and LSTM to generate the final movement prediction result.

[0072] A hand rehabilitation synchronous training system based on CNN and LSTM networks, the system includes: a data acquisition module, a processing module, a driving module, a voice interaction module, and a display module;

[0073] The data acquisition module includes: a bending sensor and a nine-axis attitude sensor;

[0074] The bending sensor is used to collect the finger bending data of the healthy hand;

[0075] The nine-axis attitude sensor is used to collect the wrist movement data of the healthy hand; the wrist movement data includes attitude angle, angular velocity, and acceleration;

[0076] The processing module is used to construct and train a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism;

[0077] The processing module is used to input finger bending data and wrist movement data into the trained motion prediction model to obtain the hand movement result;

[0078] The driving module is used to control the driving setting of the affected hand through the hand movement result to realize the synchronous rehabilitation training of the affected hand.

[0079] As an embodiment, the system integrates advanced artificial intelligence technology to improve the intelligence and accuracy of rehabilitation training. Among them, the motion trajectory prediction model is based on a deep learning framework, combines a convolutional neural network (CNN) and a long short-term memory network (LSTM), and predicts the motion intention and synchronously drives the affected hand by collecting the data of the healthy hand in real time. Compared with traditional methods, this model can more accurately fit complex motion trajectories, reduce motion delay, and improve the naturalness and effectiveness of training.

[0080] The voice interaction module adopts a voice recognition model based on the Transformer architecture;

[0081] The voice interaction module is used to conduct language training for the patient to promote the collaborative exercise of the brain and fingers through speech.

[0082] As an embodiment, the system is equipped with an on-board large model, supports voice interaction functions, and the voice recognition model based on the Transformer architecture can accurately identify and execute the patient's instructions to realize functions such as training mode switching, parameter adjustment, and training control. At the same time, it has the ability of daily conversation, and the patient can have a simple conversation with the system, such as asking about the weather, playing music, telling stories, etc., to help the patient conduct language training and promote the recovery of brain consciousness. The system also supports voice broadcast, which can real-time feedback information such as training progress, finger angle, air pressure control status, and electromyogram signal characteristics, improving the human-computer interaction experience.

[0083] The periphery of the affected hand driving device of the driving module is wrapped with a flexible rehabilitation glove.

[0084] As an embodiment, the driving module of the present invention constructs a pneumatic single-system to drive the synchronous rehabilitation training of the affected hand, designs a flexible rehabilitation glove, and reduces the risk of secondary injury caused by using traditional mechanical rehabilitation devices. At the same time, the research and development and production costs of this system are relatively low, with high cost performance. With the help of a deep learning framework and combined with the long short-term memory network (LSTM), the system can accurately predict the motion trajectory and trend, enabling the affected-side controller to respond in advance to the changes in the motion data of the healthy hand, thereby effectively improving the coordination of bilateral movements. In addition, the introduction of an interactive display screen and a voice interaction system not only enhances the intelligence level of the system but also provides a more personalized rehabilitation experience for patients, significantly improving the effect of rehabilitation training and the patient's participation.

[0085] As an embodiment, in order to comprehensively evaluate the performance of the proposed algorithm in actual finger motion prediction tasks, a comparative experiment was designed to test the performance of three methods in terms of prediction accuracy, action recognition accuracy, and system response time, respectively.

[0086] Method 1. Direct mapping control method (without using deep learning, directly mapping the data of the healthy hand to the bending degree of the affected-side pneumatic driving glove);

[0087] Method 2. Ordinary LSTM prediction (the LSTM network processes the bending sensor data);

[0088] Method 3. The algorithm of the present invention.

[0089] In the initial stage of the experiment, simulation tests were first carried out using the Simulink module in MATLAB. The trained algorithm model of the present invention was loaded through the StatefulPredict module in Simulink, and the corresponding system modules were built to achieve input-output linkage. During the simulation process, the input of the model was a simulated sine wave signal (amplitude of 1, frequency of 0.5 Hz), and this waveform simulated the dynamic characteristics of the healthy glove during the grasping action. As Figure 3 shown in the response curves of different algorithms under this excitation signal. The results show that the algorithm of the present invention can continuously and stably predict the input signal, and the output curve closely follows the real input, and is significantly superior to the control group in terms of time delay and fitting accuracy. Compared with the direct data mapping method, the response delay is reduced by about 0.4 seconds; compared with the ordinary LSTM method, the response time is advanced by about 0.1 second, and the output bending angle is almost the same as the height of the original input curve.

[0090] As an example, comparative verification of real finger movement prediction was carried out on this system. During the test, the actual bending sensor data of the healthy-side glove was collected. Three methods were respectively used to predict and drive the affected-side pneumatic glove, and the prediction results of each group of tests were recorded and compared with the true values. The following performance indicators were statistically obtained:

[0091] Method Name Average Bending Angle Error (°) Maximum Bending Angle Error (°) Action Recognition Accuracy (%) Direct Mapping Control Method 8.85 19.9 76.6 Ordinary LSTM Prediction 5.29 13.1 84.9 Algorithm of the Present Invention 3.08 8.3 92.6

[0092] The experimental results show that the algorithm of the present invention performs best in three key indicators, significantly improving the prediction accuracy and response ability, and fully verifying its effectiveness and practical value in complex hand movement prediction tasks.

[0093] The above are only exemplary embodiments of the present disclosure and should not be used to limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure.

[0094] This application aims to cover any variations, uses or adaptations of the present disclosure that follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The description and examples are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.

Claims

1. A hand rehabilitation synchronous training method based on CNN and LSTM networks, characterized in that, The method includes the following steps: S1: Obtain the finger bending data of the healthy hand and the wrist movement data of the healthy hand; S2: Construct a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism and train it; S3: Input the finger bending data and the wrist movement data into the trained motion prediction model to obtain the hand movement result; S4: Control the driving device of the affected hand through the hand movement result to achieve synchronous rehabilitation training of the affected hand.

2. The hand rehabilitation synchronous training method based on CNN and LSTM networks according to claim 1, characterized in that Step S1 includes: The finger bending data is the finger bending data of 5 fingers of the healthy hand, and the finger bending data includes: finger bending angle data and timing information; The wrist movement data includes: wrist attitude angle timing data, angular velocity timing data, and acceleration timing data.

3. The hand rehabilitation synchronous training method based on CNN and LSTM networks according to claim 1, characterized in that Step S2 includes: The motion prediction model includes: a finger feature extraction module, a wrist feature extraction module, a feature fusion module, and a fully connected layer; The finger feature extraction module is connected to the feature fusion module; the wrist feature extraction module is connected to the feature fusion module; The feature fusion module is connected to the fully connected layer.

4. The hand rehabilitation synchronous training method based on CNN and LSTM networks according to claim 3, characterized in that, The finger feature extraction module includes: a CNN convolutional unit and a first LSTM unit; Through the CNN convolutional unit, extract the spatial features of the finger angle change of the finger bending data; Input the spatial features into the first LSTM unit to obtain the first dynamic features of the finger movement changing with time.

5. The hand rehabilitation synchronous training method based on CNN and LSTM networks according to claim 4, wherein The wrist feature extraction module uses an LSTM long short-term memory network; Through the wrist feature extraction module, extract the second dynamic features of the overall movement of the wrist movement data.

6. The hand rehabilitation synchronous training method based on CNN and LSTM networks according to claim 5, characterized in that, The feature fusion module uses a self-attention mechanism to fuse the first dynamic features and the second dynamic features; The fused features pass through the fully connected layer to generate the final finger movement prediction value, that is, the hand movement result.

7. A hand rehabilitation synchronous training system based on CNN and LSTM networks, which is used to implement a hand rehabilitation synchronous training method based on CNN and LSTM networks as described in any one of claims 1-6, characterized in that, The system includes: a data acquisition module, a processing module, a driving module, a voice interaction module, and a display module; The data acquisition module includes: a bending sensor and a nine-axis attitude sensor; The bending sensor is used to collect the finger bending data of the healthy hand; The nine-axis attitude sensor is used to collect the wrist movement data of the healthy hand; the wrist movement data includes attitude angle, angular velocity, and acceleration; The processing module is used to construct a motion prediction model through a convolutional neural network, a long short-term memory network, and a self-attention mechanism and train it; The processing module is used to input the finger bending data and the wrist movement data into the trained motion prediction model to obtain the hand movement result; The driving module is used to control the driving setting of the affected hand through the hand movement result to achieve synchronous rehabilitation training of the affected hand.

8. The hand rehabilitation synchronous training system based on CNN and LSTM networks according to claim 7, wherein The voice interaction module uses a speech recognition model based on the Transformer architecture; The voice interaction module is used to conduct language training on the patient to promote the collaborative exercise of the brain and fingers through speech.

9. The hand rehabilitation synchronous training system based on CNN and LSTM networks according to claim 7, characterized in that The driving device of the affected hand of the driving module is peripherally wrapped with a flexible rehabilitation glove.