Cerebral stroke upper limb rehabilitation training system and method based on dynamic reward feedback

By using multimodal data acquisition and fusion technology, combined with improved algorithms and models, we have achieved accurate decoding of motor intentions and personalized reward decision-making in the upper limb rehabilitation training system for stroke. This solves the problems of insufficient accuracy in motor intention recognition and insufficient adaptability of reward strategies in existing technologies, thereby improving the effectiveness and safety of rehabilitation training.

CN121601151APending Publication Date: 2026-03-03SHANGHAI SECOND REHABILITATION HOSPITAL (SHANGHAI BAOSHAN NO 1 STEEL HOSPITAL)

Patent Information

Application Number
CN202511784130.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-01
Publication Date
2026-03-03

Smart Images

  • Figure CN121601151A_ABST
    Figure CN121601151A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a cerebral apoplexy upper limb rehabilitation training system and method based on dynamic reward feedback. A data acquisition module is used for acquiring physiological signals, motion signals, emotional state data and training performance data; the motion intention decoding module is used for performing motion intention feature extraction and reliability evaluation on the physiological signals according to an improved MREE-Net + + fusion algorithm, performing weighted fusion after feature weights are adjusted according to a reliability result, obtaining a unified motion intention feature vector, performing motion intention decoding, and obtaining a motion intention intensity index and a motion intention vector; the dynamic reward decision module is used for processing the facial micro-expression data according to an improved VGG-Face model to obtain an emotion titer, and obtaining an emotion awakening degree according to the voice signal; a reward action is generated according to an improved deep reinforcement learning algorithm; and the adaptive training regulation and control module is used for generating a personalized virtual training scene according to the generative adversarial network and regulating and controlling the training intensity according to the fatigue index. The rehabilitation training effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of rehabilitation medicine technology, and relates to, but is not limited to, a system and method for upper limb rehabilitation training for stroke based on dynamic reward feedback. Background Technology

[0002] Stroke is a leading cause of long-term disability, causing upper limb motor dysfunction and severely impacting quality of life. Therefore, appropriate and effective rehabilitation training is crucial. Current rehabilitation training methods primarily utilize optical and inertial sensors to acquire real-time upper limb movement trajectories, creating virtual rehabilitation training environments using virtual reality technology. Intelligent algorithms are employed to identify and assess the patient's rehabilitation status and posture during training, and reward strategies are used to enhance motivation and participation. However, while existing methods have some effectiveness, they also have limitations. Existing brain-computer interface systems often rely on single-modal signals, making them susceptible to noise interference and limiting the accuracy of motor intention decoding, hindering accurate identification of patient intentions. Furthermore, current technologies often use fixed or simple probability-based rewards, lacking personalized dynamic adjustments and failing to adapt to the patient's real-time emotional state, making it difficult to maintain long-term training motivation. Moreover, current technologies struggle to monitor patient fatigue in real-time, leading to overtraining or undertraining, affecting rehabilitation outcomes and posing safety risks. Additionally, existing virtual rehabilitation training scenarios are often pre-set, unable to dynamically adjust based on the patient's specific functional impairments and real-time abilities, resulting in poor training targeting.

[0003] Therefore, there is an urgent need for a more comprehensive rehabilitation training system to address the problems existing in current rehabilitation training technologies, such as limited accuracy in motor intention recognition, insufficient adaptability of reward strategies, difficulty in matching training intensity with patient status in real time, and poor targeting of training content. The system should construct a complete closed loop of perception-decision-execution-feedback optimization to significantly improve the accuracy of motor intention decoding, patient training motivation, and rehabilitation training effectiveness. Summary of the Invention

[0004] This application provides a system and method for upper limb rehabilitation training after stroke based on dynamic reward feedback.

[0005] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a stroke upper limb rehabilitation training system based on dynamic reward feedback. The system includes a data acquisition module, a movement intention decoding module, a dynamic reward decision-making module, and an adaptive training control module, wherein: The data acquisition module is used to collect patients' physiological signals, motion signals, emotional state data, and training performance data. The physiological signals include electroencephalogram (EEG) signals, blood oxygen saturation signals, electromyography (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and speech signals. The motion intent decoding module is used to extract motion intent features and assess reliability from the physiological signals using an improved MTREE-Net++ fusion algorithm. Based on the multimodal reliability results, the module adjusts the multimodal motion intent feature weights, performs weighted fusion based on these weights to obtain a unified motion intent feature vector, and then decodes the motion intent based on this unified feature vector. The system includes a motion intent intensity index and a motion intent vector; a dynamic reward decision module, which processes the facial micro-expression data according to an improved VGG-Face model to obtain an emotional valence, wherein the improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer, and obtains an emotional arousal level based on the speech signal; a reward action is generated based on the emotional valence and the emotional arousal level using an improved deep reinforcement learning algorithm; and an adaptive training control module, which generates personalized virtual training scenarios based on the reward action using a generative adversarial network, calculates a fatigue index based on real-time physiological signals, and controls the training intensity according to the fatigue index to obtain a personalized training task.

[0006] The technical solution provided in this application collects patients' physiological signals, motion signals, emotional state data, and training performance data through a data acquisition module to achieve simultaneous acquisition and fusion of multi-dimensional and multi-modal data, providing a comprehensive and reliable data foundation for subsequent processing. In the motion intention decoding module, motion intention features are extracted and reliability is assessed based on the improved MTREE-Net++ fusion algorithm. The multi-modal motion intention feature weights are adjusted based on the multi-modal reliability results, and weighted fusion is performed according to these weights to obtain a unified motion intention feature vector. Motion intention decoding is then performed based on this unified feature vector to obtain a motion intention intensity index and a motion intention vector, thereby significantly improving the accuracy and robustness of motion intention decoding and effectively solving the decoding failure problem caused by instantaneous quality degradation of a single signal, providing a refined structured input for subsequent decision-making. In the dynamic reward decision module, facial micro-expression data is processed using an improved VGG-Face model to obtain… Regarding emotional valence, the improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. Emotional arousal is obtained based on speech signals, enabling accurate perception of the patient's emotional state and providing reliable emotional input for reward decisions. Based on emotional valence and emotional arousal, a reward action is generated using an improved deep reinforcement learning algorithm. An emotional decay factor is introduced to dynamically adjust the decay rate of the reward value, making the reward feedback more personalized and aligned with the patient's real-time psychological state, effectively maintaining and enhancing the patient's training motivation and participation. In the adaptive training control module, personalized virtual training scenarios are generated based on reward actions using a generative adversarial network to ensure that the training tasks are highly aligned with the patient's actual rehabilitation needs and life scenarios. A fatigue index is calculated based on real-time physiological signals, and the training intensity is adjusted according to the fatigue index to obtain personalized training tasks. This maximizes training efficiency and safety while avoiding excessive fatigue and injury to the patient, effectively balancing training intensity and the patient's physiological tolerance.

[0007] Optionally, the motion intent decoding module includes a signal purification unit, a feature fusion unit, and an intent decoding unit, wherein: the signal purification unit is used to adaptively denoise the physiological signal based on an improved wavelet threshold denoising algorithm, filter electrooculogram artifacts, and remove motion artifacts in conjunction with the motion signal; the feature fusion unit is used to extract multimodal motion intent features based on the physiological signal according to an improved MTREE-Net++ fusion algorithm, perform reliability assessment based on the multimodal motion intent features, dynamically adjust the weights of the multimodal motion intent features based on the multimodal reliability results, and perform weighted fusion of the multimodal motion intent features according to the weights of the multimodal motion intent features to obtain a unified motion intent feature vector; the intent decoding unit is used to calculate a motion intent intensity index based on the multimodal motion intent features and the unified motion intent feature vector, and perform motion intent decoding based on the motion intent intensity index and the unified motion intent feature vector to obtain a motion intent vector, wherein the motion intent vector includes action type, expected intensity, and execution confidence.

[0008] Optionally, the feature fusion unit includes a feature extraction subunit, a reliability assessment subunit, and a weight fusion subunit, wherein: the feature extraction subunit is used to calculate the degree of μ rhythm inhibition based on the EEG signal, calculate the rate of change of oxyhemoglobin concentration based on the blood oxygen signal, and calculate the root mean square value and rate of change of the electromyography (EMG) signal based on the EMG signal, and standardize the degree of μ rhythm inhibition, the rate of change of oxyhemoglobin concentration, the root mean square value and the rate of change of the EMG signal; the reliability assessment subunit is used to determine the quality of the physiological signal based on the signal-to-noise ratio, cross-modal correlation, and physiological parameter range, and when the quality of the physiological signal is lower than a preset threshold, it is considered an unreliable signal, wherein when the... When the μ-rhythm signal-to-noise ratio of the EEG signal is lower than 1.2 or the correlation coefficient with the blood oxygenation signal feature is lower than 0.6, the EEG signal is considered unreliable. When the ratio of the change in oxyhemoglobin concentration to the change in deoxyhemoglobin concentration of the blood oxygenation signal exceeds the physiological range of 0.8-1.2, the blood oxygenation signal is considered unreliable. When the root mean square rate of change of the electromyography (EMG) signal exceeds 1.5 times the initial baseline value, the EMG signal is considered unreliable. The weighted fusion subunit is used to dynamically adjust the weights of multimodal motion intention features based on the multimodal reliability results. The multimodal motion intention features are weighted and fused according to the weights to obtain a unified motion intention feature vector. The weight adjustment formula is shown below: ; In the formula, Indicates the first Adjusted weights for unreliable signals; Indicates the first Initial weights for unreliable signals; Indicates the first The current signal-to-noise ratio of an unreliable signal; Indicates the first The signal-to-noise ratio stability threshold for an unreliable signal; Indicates the first Adjusted weights of reliable signals; Indicates the first The initial weights of a reliable signal; This represents the total number of modes participating in the fusion.

[0009] Optionally, the dynamic reward decision module includes an emotion state perception unit and a reward decision unit, wherein: the emotion state perception unit is used to input the facial micro-expression data and the motion signal into an improved VGG-Face model, perform motion artifact correction on the facial micro-expression data based on the motion signal in a motion interference filtering layer, input the corrected facial image into a convolutional feature extraction network, enhance the local micro-expression features of the corner of the eyes and corner of the mouth through a micro-expression enhancement layer, and obtain the emotion valence through a fully connected layer; and obtain the emotional arousal by processing the speech signal through a regression model; the reward decision unit is used to construct a comprehensive state vector based on motion accuracy data, the emotion valence, the emotional arousal, the motion intention intensity index, and the motion intention vector, and process the comprehensive state vector based on an improved deep Q-network algorithm to generate a reward action, wherein the emotion decay factor in the deep Q-network is dynamically adjusted according to the emotion valence, when the emotion valence is less than -0.3, the emotion decay factor is 0.6, and when the emotion valence is greater than or equal to -0.3, the emotion decay factor is 0.9.

[0010] Optionally, the emotion state perception unit includes a motion interference filtering subunit, a micro-expression enhancement subunit, and a speech emotion perception subunit, wherein: the motion interference filtering subunit is used to calculate the head displacement compensation amount based on the motion signal, and to perform motion artifact correction on the facial micro-expression data through affine transformation; the micro-expression enhancement subunit is used to process the corrected facial micro-expression data through a convolutional feature extraction network, add the output features of the 4th convolutional layer to the input features of the 5th convolutional layer through residual connections, and perform 1×1 convolution dimensionality reduction and 3×3 convolution operations in the residual branch to enhance the local micro-expression features of the corners of the eyes and mouth, and map the enhanced features to emotional valence through a fully connected layer; the speech emotion perception subunit is used to perform frame segmentation, windowing, fast Fourier transform, and Mel frequency cepstral coefficient feature extraction on the speech signal, and input it into a regression model to obtain the emotional arousal level.

[0011] Optionally, the adaptive training control module includes a training scenario generation unit and a training intensity control unit, wherein: the training scenario generation unit is used to generate personalized virtual training scenarios based on the results of daily living activities assessment, shoulder joint range of motion, and movement intention intensity index, according to a generative adversarial network; the training intensity control unit is used to calculate a fatigue index based on real-time root mean square values ​​of electromyography signals, changes in deoxyhemoglobin concentration, and changes in oxyhemoglobin concentration, and to control the training intensity according to the fatigue index. When the fatigue index is greater than 0.7, the training intensity is reduced; when the fatigue index is less than 0.3, the training intensity is gradually increased; when the fatigue index is greater than or equal to 0.3 and less than or equal to 0.7, the current training intensity is maintained, thus obtaining a personalized training task. The formula for calculating the fatigue index is shown below: ; In the formula, Indicates fatigue index; Indicates muscle fatigue weight; This represents the root mean square value of the electromyographic signal; Indicates the central fatigue weight; This represents the change in deoxyhemoglobin concentration; This represents the change in oxyhemoglobin concentration.

[0012] Optionally, the reward action includes reward type, reward intensity, and reward timing. The reward type is determined based on the comparison between the emotional valence and a preset threshold. When the emotional valence is greater than the preset threshold, an achievement-type reward is triggered; otherwise, an encouragement-type reward is triggered. The reward intensity is determined based on the motion accuracy data and the emotional arousal level. The reward timing is within 300ms after the action is completed.

[0013] Secondly, embodiments of this application provide a method for upper limb rehabilitation training after stroke based on dynamic reward feedback. This method is applied to a system for upper limb rehabilitation training after stroke based on dynamic reward feedback. The system includes a data acquisition module, a motor intention decoding module, a dynamic reward decision-making module, and an adaptive training control module. The method includes: acquiring the patient's physiological signals, motor signals, emotional state data, and training performance data. The physiological signals include electroencephalogram (EEG) signals, blood oxygen saturation signals, electromyogram (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and speech signals. The method also involves extracting motor intention features and assessing reliability of the physiological signals using an improved MTREE-Net++ fusion algorithm, and adjusting the multimodal motor intention features based on the multimodal reliability results. The weights are weighted and fused according to the multimodal motion intention feature weights to obtain a unified motion intention feature vector. Motion intention is decoded based on the unified motion intention feature vector to obtain a motion intention intensity index and a motion intention vector. The facial micro-expression data is processed using an improved VGG-Face model to obtain an emotional valence. The improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. Emotional arousal is obtained based on the speech signal. A reward strategy is generated based on the emotional valence and the emotional arousal using an improved deep reinforcement learning algorithm. A personalized virtual training scenario is generated using a generative adversarial network. A fatigue index is calculated based on real-time physiological signals. The training intensity is adjusted based on the fatigue index to obtain a personalized training task.

[0014] Thirdly, embodiments of this application provide an electronic device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the steps in the above-described method for upper limb rehabilitation training for stroke based on dynamic reward feedback.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the above-described method for upper limb rehabilitation training for stroke based on dynamic reward feedback.

[0016] The beneficial effects of the technical solutions provided in this application include at least the following: This application provides a stroke upper limb rehabilitation training system and method based on dynamic reward feedback. The system collects patients' physiological signals, motor signals, emotional state data, and training performance data through a data acquisition module to achieve simultaneous acquisition and fusion of multi-dimensional and multi-modal data, providing a comprehensive and reliable data foundation for subsequent processing. In the motor intention decoding module, the physiological signals are used to extract motor intention features and assess reliability based on an improved MTREE-Net++ fusion algorithm. The multi-modal motor intention feature weights are adjusted based on the multi-modal reliability results, and weighted fusion is performed according to these weights to obtain a unified motor intention feature vector. Motor intention decoding is then performed based on this unified feature vector to obtain a motor intention intensity index and a motor intention vector, thereby significantly improving the accuracy and robustness of motor intention decoding and effectively solving the decoding failure problem caused by the instantaneous quality degradation of a single signal, providing a refined structured input for subsequent decision-making. In the dynamic reward decision-making module, an improved VGG-Face model is used... The system processes facial micro-expression data to obtain emotional valence. The improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. Emotional arousal is calculated based on speech signals, enabling accurate perception of the patient's emotional state and providing reliable emotional input for reward decisions. Based on emotional valence and emotional arousal, an improved deep reinforcement learning algorithm generates reward actions, and an emotional decay factor is introduced to dynamically adjust the decay rate of the reward value. This makes the reward feedback more personalized and aligned with the patient's real-time psychological state, effectively maintaining and enhancing the patient's training motivation and participation. In the adaptive training control module, personalized virtual training scenarios are generated based on reward actions using a generative adversarial network. This ensures that the training tasks are highly aligned with the patient's actual rehabilitation needs and life scenarios. A fatigue index is calculated based on real-time physiological signals, and the training intensity is adjusted accordingly to obtain personalized training tasks. This maximizes training efficiency and safety while avoiding excessive fatigue and injury, effectively balancing training intensity with the patient's physiological tolerance. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 A schematic diagram of a stroke upper limb rehabilitation training system based on dynamic reward feedback provided in an embodiment of this application; Figure 2 A schematic diagram of an improved VGG-Face model provided for an embodiment of this application; Figure 3 A flowchart illustrating a stroke upper limb rehabilitation training method based on dynamic reward feedback, provided as an embodiment of this application; Figure 4 This is a schematic diagram of the hardware entity of an electronic device provided in an embodiment of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0020] It should be noted that the terms "first, second, and third" used in the embodiments of this application are merely to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0021] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this application pertain. It should also be understood that terms such as those defined in general dictionaries should be understood to have a meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0022] The embodiments of this application will be further described below with reference to the accompanying drawings.

[0023] In view of the current problems in the research on upper limb rehabilitation training for stroke in the field of rehabilitation medicine, this application provides a system and method for upper limb rehabilitation training for stroke based on dynamic reward feedback.

[0024] The technical solution of this application is described below, starting with the system implementation of this application.

[0025] Please refer to Figure 1 The illustration shows a schematic diagram of a stroke upper limb rehabilitation training system based on dynamic reward feedback provided in an embodiment of this application, such as... Figure 1 As shown, the system includes a data acquisition module 01, a motion intention decoding module 02, a dynamic reward decision module 03, and an adaptive training control module 04. The data acquisition module 01 is used to collect the patient's physiological signals, motion signals, emotional state data, and training performance data. The physiological signals include electroencephalogram (EEG) signals, blood oxygen saturation signals, electromyography (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and speech signals. The motion intention decoding module 02 is used to extract motion intention features and assess reliability of the physiological signals using an improved MTREE-Net++ fusion algorithm. Based on the multimodal reliability results, it adjusts the multimodal motion intention feature weights, performs weighted fusion based on the multimodal motion intention feature weights to obtain a unified motion intention feature vector, and then determines the motion intention based on the unified motion intention feature vector. Decoding yields the motion intent intensity index and motion intent vector; the dynamic reward decision module 03 processes the facial micro-expression data according to the improved VGG-Face model to obtain the emotional valence, wherein the improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer, and obtains the emotional arousal level based on the speech signal; based on the emotional valence and the emotional arousal level, a reward strategy is generated according to an improved deep reinforcement learning algorithm; the adaptive training control module 04 generates personalized virtual training scenarios according to the generative adversarial network, calculates the fatigue index based on real-time physiological signals, and controls the training intensity according to the fatigue index to obtain personalized training tasks.

[0026] In this embodiment, the data acquisition module 01 is used to collect the patient's physiological signals, motor signals, emotional state data, and training performance data. Specifically, the physiological signals include electroencephalogram (EEG) signals, blood oxygenation signals, electromyography (EMG) signals, and heart rate data. The EEG signals are acquired through a 64-channel EEG cap, focusing on monitoring changes in the μ rhythm of the motor cortex to reflect the neural activity generated by the brain's motor intention. For example, when the μ rhythm shows power inhibition, it indicates that the patient has generated a motor intention. The blood oxygenation signals are acquired through an 8-channel functional near-infrared spectroscopy module, mainly monitoring changes in the concentration of oxyhemoglobin and deoxyhemoglobin to assess the activation state of the motor cortex. The EMG signals and heart rate data are acquired through an EMG wristband. The EMG signals can reflect the electrical activity of the patient's upper limb muscles during contraction to assess muscle strength and fatigue state. The heart rate data can monitor the patient's cardiovascular status during training as an auxiliary indicator for safety and fatigue assessment. Motion signals are obtained by collecting upper limb movement posture data through an inertial measurement unit (IMU), including upper limb movement position, angle, and acceleration. Their main uses are reconstructing the movement trajectory, calculating the error between the actual movement and the target trajectory, and assisting in removing head motion artifacts in facial expressions. Emotional state data includes facial micro-expression data and speech signals. Facial micro-expression data is acquired through a camera and is mainly used for subsequent emotion valence quantification, while speech signals are acquired through a microphone and are mainly used for subsequent emotion arousal quantification. Training performance data includes task completion rate, movement accuracy score, and action execution time, primarily determined based on upper limb movement trajectory data collected by the IMU and system clock timing. Reward strategies and training difficulty can be dynamically adjusted using training performance data. Furthermore, physiological signals, motion signals, emotional state data, and training performance data are synchronized and aligned using a unified timestamp to ensure consistency across the time dimension, providing a reliable data foundation for subsequent processing.

[0027] In this embodiment, the motion intent decoding module 02 includes a signal purification unit, a feature fusion unit, and an intent decoding unit. The signal purification unit adaptively denoises the physiological signal based on an improved wavelet threshold denoising algorithm, filtering electrooculography (EOG) artifacts and removing motion artifacts by combining the motion signal. Specifically, an adaptive filtering window of 50-200ms is designed. When a low-frequency signal abrupt change caused by EOG artifacts is detected, the filtering window is automatically reduced to 50ms to achieve accurate identification and filtering of EOG artifact ranges. When no artifacts are detected, the filtering window is extended to 200ms to ensure signal continuity. Simultaneously, motion artifact correction is performed on the fNIRS signal by combining upper limb motion posture data from the motion signal, removing motion artifacts caused by head shaking to ensure the quality of the blood oxygenation signal. Furthermore, the electromyography (EMG) signal, after bandpass filtering and baseline correction, is used as an auxiliary signal for subsequent processing.

[0028] In this embodiment, the feature fusion unit is used to extract multimodal motion intention features based on physiological signals according to the improved MTREE-Net++ fusion algorithm, perform reliability assessment based on the multimodal motion intention features, dynamically adjust the weights of the multimodal motion intention features based on the multimodal reliability results, and perform weighted fusion of the multimodal motion intention features according to the weights to obtain a unified motion intention feature vector. Specifically, in the feature extraction subunit, for the EEG, blood oxygenation, and electromyography signals in the purified physiological signals, features strongly correlated with motion intention, i.e., motion intention features, are extracted according to the physiological significance of each modality signal. Among them, for the EEG signal, transient features reflecting motion intention are extracted, and the 8-13Hz μ rhythm is analyzed by Fast Fourier Transform to determine the degree of μ rhythm power suppression 1 second before and at the time of motion intention generation; for the blood oxygenation signal, stability features reflecting motor cortex activation are extracted, and the purified oxyhemoglobin concentration is analyzed. After baseline correction of the change value signal, the rate of change of oxyhemoglobin concentration within the time period related to the movement intention is calculated; muscle execution state features are extracted from the electromyography (EMG) signal, and the root mean square (RMS) value and the RMS rate of change of the EMG signal within a 5-second sliding window are calculated. The degree of μ rhythm inhibition, the rate of change of oxyhemoglobin concentration, the RMS value and the RMS rate of change of the EMG signal are standardized and mapped to the 0-1 interval to eliminate dimensional differences, thereby obtaining the movement intention features corresponding to the EEG signal, blood oxygenation signal and EMG signal respectively. In the reliability assessment subunit, the quality of physiological signals is evaluated through a triple mechanism of signal-to-noise ratio, cross-modal correlation, and physiological parameter range. When the quality of physiological signals is lower than a preset threshold, they are considered unreliable signals. Specifically, when the μ-rhythm signal-to-noise ratio of the EEG signal is lower than 1.2 or the correlation coefficient with the blood oxygen signal characteristics is lower than 0.6, the EEG signal is considered unreliable. When the ratio of the change in oxyhemoglobin concentration to the change in deoxyhemoglobin concentration of the blood oxygen signal exceeds the normal physiological range of 0.8-1.2, the blood oxygen signal is considered unreliable. Using the root mean square value of the initial 5 minutes of training as a benchmark, when the real-time root mean square change rate of the electromyography signal exceeds 1.5 times the initial benchmark value, the electromyography signal is considered unreliable.

[0029] Furthermore, in the weighted fusion subunit, initial weights are set according to the physiological significance of each modality signal. The initial weight for EEG signals is 0.4 to directly reflect motor cortex activity, the initial weight for blood oxygenation signals is 0.3, and the initial weight for electromyography (EMG) signals is 0.3 to assist in muscle state assessment. The weights of multimodal motor intention features are dynamically adjusted based on multimodal reliability results. When a modality signal is determined to be unreliable, its weight is dynamically adjusted. The weight adjustment formula is shown below: ; In the formula, Indicates the first Adjusted weights for unreliable signals; Indicates the first Initial weights for unreliable signals; Indicates the first The current signal-to-noise ratio of an unreliable signal; Indicates the first The signal-to-noise ratio stability threshold for an unreliable signal; Indicates the first Adjusted weights of reliable signals; Indicates the first The initial weights of a reliable signal; This represents the total number of modalities participating in the fusion, N=3. Finally, the motion intention features of each modality are multiplied by their corresponding motion intention feature weights and then summed to obtain the unified motion intention feature vector.

[0030] In this embodiment, the motion intent is accurately quantified in the intent decoding unit. A motion intent intensity index is calculated based on multimodal motion intent features to quantify the intensity of the motion intent. This primarily focuses on activation of the motor cortex, excluding interference from muscle fatigue. The calculation formula for the motion intent intensity index is shown below: ; In the formula, Indicates the intensity index of the intent to move; This represents the change in oxyhemoglobin concentration; Weights representing the motor intent features of EEG signals; The motion intent feature weights represent the blood oxygenation signal. The motor intention intensity index is standardized to a range of 0-1, with a higher value indicating stronger motor intention, 1 representing the strongest motor intention, and 0 representing no motor intention. Motor intention is decoded based on the detailed information in the motor intention intensity index and the unified motor intention feature vector, resulting in a motor intention vector. This vector includes the action type, expected intensity, and execution confidence. The action type is determined by analyzing the spatial distribution of EEG signal characteristics and the activation areas of functional near-infrared spectroscopy signals, such as raising the arm, grasping, and relaxing. The expected intensity is categorized into low, medium, and high levels, determined by combining the root mean square value of the electromyography signal with the motor intention intensity index. The execution confidence is determined based on the reliability scores of each modality signal. Finally, the motor intention vector and motor intention intensity index are pushed to the dynamic reward decision module 03 at a frequency of 12.5Hz, with the decoding delay strictly controlled within 80ms to meet the time sensitivity requirements of neural feedback and match the optimal time window for the formation of brain neuroplasticity.

[0031] In this embodiment, the dynamic reward decision module 03 includes an emotion state perception unit and a reward decision unit. The emotion state perception unit is used to input the facial micro-expression data and the motion signals into the improved VGG-Face model. (Please refer to...) Figure 2 The illustration shows a schematic diagram of an improved VGG-Face model provided in this application embodiment. In the motion interference filtering subunit, the motion signal is calculated by the motion interference filtering layer through the kinematic model to obtain the head displacement compensation amount. The facial micro-expression data is corrected for motion artifacts by affine transformation to eliminate expression artifacts caused by head movement and return the key facial feature points to the reference position. For example, if the head is detected to rotate 5° to the left, the facial micro-expression image is corrected to rotate 5° to the right. The corrected facial micro-expression data is processed by a convolutional feature extraction network. The micro-expression enhancement layer is located between the 4th and 5th convolutional layers. In the micro-expression enhancement subunit, the output features of the 4th convolutional layer and the input features of the 5th convolutional layer are added through residual connections. In the residual branch, 1×1 convolution dimensionality reduction and 3×3 convolution operations are performed sequentially to enhance the local micro-expression features of the corner of the eye region [20%W, 30%H] to [30%W, 40%H] and the corner of the mouth region [40%W, 65%H] to [60%W, 75%H]. W and H are the width and height of the image, respectively. Redundant information in non-critical areas such as the forehead and cheeks is suppressed, thereby significantly improving the model's ability to capture micro-expression features. After improving the VGG-Face model, the micro-expression recognition accuracy reached 92.7%, and the recognition error caused by head movement was reduced by 42%. Finally, the motion signal was used as an auxiliary vector and incorporated into the features after the 5th convolution through a fully connected layer to further suppress the motion interference that was not completely eliminated. Finally, the enhanced features were mapped to the emotional efficacy value that continuously varies in the range of [-1,1]. In the speech emotion perception subunit, the continuous speech signal is first divided into frames of 20-30ms with 50% overlap between frames. After adding a Hanning window to each frame, the DC component is removed by high-pass filtering and pre-emphasis is performed. The spectral amplitude of each frame signal is obtained by fast Fourier transform. Mel spectrum is generated by a triangular filter bank with 26-40 Mel scales. After taking the logarithm, discrete cosine transform is performed to extract the first 12-13 Mel frequency cepstral coefficients. First-order and second-order differences are added to construct a complete feature vector. Feature dimensions that are strongly correlated with arousal are selected from the feature vector, and after dimensionality reduction, they are input into the trained regression model. The model output value is normalized to the [-1,1] interval and fine-tuned with clinical data to finally obtain the emotional arousal level.

[0032] In this embodiment, the reward decision unit generates a reward strategy based on an improved deep reinforcement learning algorithm. Specifically, a comprehensive state vector is constructed based on motion accuracy data, emotional valence, emotional arousal, motion intention intensity index, and motion intention vector. The comprehensive state vector is then processed using an improved deep Q-network algorithm to generate a three-dimensional reward action, including reward type, reward intensity, and reward timing. The reward type is determined based on a comparison between emotional valence and a preset threshold of 0.3. When the emotional valence is greater than the preset threshold of 0.3, it indicates that the patient is experiencing a positive emotion, triggering an achievement-type reward, including virtual badge display and complex multi-frequency tactile feedback. When the emotional valence is less than or equal to the preset threshold of 0.3, it indicates that the patient is experiencing a positive emotion, triggering an achievement-type reward, including virtual badge display and complex multi-frequency tactile feedback. When the patient is in a neutral or negative emotional state, an encouraging reward is triggered, using voice guidance (such as "Good job, keep it up") combined with basic vibrational feedback. By adapting the reward type to the patient's emotional state, the feedback is more targeted. The reward intensity is determined based on motor accuracy data and emotional arousal, and is positively correlated with these two factors, calculated through linear or nonlinear mapping. Event-related potential synchronization technology is used to accurately capture the completion point of the movement, controlling the timing of the reward to be triggered within 300ms after the action is completed, in order to match the optimal time window for the formation of brain neuroplasticity and strengthen the neural connections of motor learning.

[0033] It is important to note that an emotional decay factor is introduced into deep Q-networks to dynamically adjust the rate of reward value decay. Specifically, in the traditional Q-value update formula of deep Q-networks, the fixed discount factor is replaced with an emotional decay factor. This emotional decay factor is dynamically adjusted according to emotional valence, with a value range of 0.6-0.9. When the emotional valence is less than -0.3, i.e., when the patient is depressed, the emotional decay factor is reduced from the usual 0.9 to 0.6 to slow down the rate of reward value decay, making the reward effect of a successful action last longer. When the emotional valence is greater than or equal to -0.3, i.e., when the patient's emotions are stable or positive, the emotional decay factor rises back to 0.9 to maintain a normal decay rhythm, thereby effectively preventing patients from losing training motivation due to short-term frustration.

[0034] In optional embodiments, the three-dimensional reward decision is converted into specific execution instructions, including controlling the vibration parameters of the haptic feedback glove, the special effects level of the virtual reality scene, and personalized voice text. Simultaneously, the reward effect data after the patient receives reward feedback is monitored, including behavioral response lag and the magnitude of emotional valence changes. This reward effect data is then fed back to the motor intention decoding module for fine-tuning and optimization, forming a complete decision-feedback-optimization closed loop. The technical solution provided in this application significantly improves patient training motivation and rehabilitation outcomes through precise emotion perception and intelligent reward decision-making.

[0035] In this embodiment, the adaptive training control module 04 includes a training scenario generation unit and a training intensity control unit. The training scenario generation unit generates personalized virtual training scenarios based on the assessment results of daily living activities, shoulder joint range of motion, and the movement intention intensity index, using a generative adversarial network. Specifically, the assessment results of daily living activities mainly reflect the degree of functional impairment in specific daily life scenarios such as eating and dressing; the range of motion of the affected shoulder joint mainly quantifies the limitation of joint range of motion; and the movement intention intensity index mainly characterizes the strength of the patient's real-time movement intention. The assessment results of daily living activities, shoulder joint range of motion, and movement intention intensity index are encoded separately and then concatenated to generate a unified feature vector. This unified feature vector is input into the generator, and after feature expansion through multiple fully connected layers, a preliminary virtual scene spatial feature map is generated via a transposed convolutional layer. Taking the rehabilitation of eating function in stroke patients as an example, the preliminary virtual scene spatial feature map includes the outline structure of the dining table, tableware, and food. The generator integrates a spatial attention mechanism, which enhances the features of key areas through global average pooling and fully connected layers, suppresses irrelevant background information, and strengthens the scene areas most relevant to the patient's functional impairment. For example, it enhances the features of the tableware gripping area, food placement area, tableware outline, and food position in the eating scene, and suppresses the features of the wall decoration background. Finally, the feature map is processed by convolution and sigmoid activation function to obtain personalized virtual training scene images. For example, the scene includes a 200g lightweight spoon to adapt to the patient's initial muscle strength, a table height to adapt to the sitting posture to avoid overextension of the shoulder joint, and food placed within a 30° range of motion on the affected side to adapt to the range of motion of the shoulder joint. Furthermore, the personalized virtual training scene images output by the generator are evaluated twice by a discriminator. First, visual realism is assessed, i.e., whether it conforms to the characteristics of real-world daily life activities. Second, functional adaptability is assessed, such as whether a spoon weighing 200g is suitable for a severely impaired patient's ability, and whether the height of the dining table matches a 30° shoulder joint range of motion. After processing the personalized virtual training scene images, the discriminator outputs the scene's true probability and adaptability score. When both the scene's true probability and adaptability score reach preset thresholds, such as a true probability greater than 0.7 and an adaptability score greater than 0.6, the scene is judged as a qualified scene and delivered for use. The technical solution provided in this application, through adversarial processing between the generator and the discriminator, generates virtual training scenes and continuously optimizes them, ultimately obtaining highly targeted and reliable virtual rehabilitation training scenes.

[0036] In this embodiment, the training intensity control unit calculates a fatigue index based on real-time root mean square (RMS) values ​​of electromyography (EMG) signals, changes in deoxyhemoglobin concentration, and changes in oxyhemoglobin concentration, and adjusts the training intensity according to the fatigue index. Specifically, it calculates the real-time RMS value of EMG signals based on real-time EMG signals, calculates the real-time RMS value change rate using a 5-second sliding window, collects real-time changes in deoxyhemoglobin concentration and oxyhemoglobin concentration in the motor cortex, and calculates the ratio of the deoxyhemoglobin concentration change to the oxyhemoglobin concentration change. After standardizing the RMS values ​​of EMG signals and the ratio of the deoxyhemoglobin concentration change to the oxyhemoglobin concentration change, a weighted fusion is performed based on clinically optimized weights, such as a muscle fatigue weight of 0.6 and a central fatigue weight of 0.4, to obtain the fatigue index. The calculation formula for the fatigue index is shown below: ; In the formula, Indicates fatigue index; Indicates muscle fatigue weight; This represents the root mean square value of the electromyographic signal; Indicates the central fatigue weight; This represents the change in deoxyhemoglobin concentration; This represents the change in oxyhemoglobin concentration; This represents the ratio of the standardized change in deoxyhemoglobin concentration to the change in oxyhemoglobin concentration. Furthermore, a preset training intensity control strategy is implemented based on the fatigue index. When the fatigue index is greater than 0.7, it indicates that the patient is in a state of significant fatigue. In this case, an intervention mechanism is immediately triggered, such as inserting a 5-second virtual reality relaxation scene and temporarily reducing the complexity of the training task, such as shortening the movement trajectory length or reducing the weight of the virtual object. When the fatigue index is less than 0.3, it indicates that the patient's fatigue level is low. In this case, the training difficulty is gradually increased, such as increasing the weight of the virtual object or adjusting the placement of items to expand the range of motion, to continuously challenge and promote the patient's functional progress. When the fatigue index is greater than or equal to 0.3 and less than or equal to 0.7, the current training intensity is maintained, resulting in a final personalized training task. This ensures that the training process is always within the patient's physiological tolerance range, effectively balancing training stimulation and fatigue recovery, thereby maximizing rehabilitation training efficiency while avoiding excessive fatigue. The training scenario generation unit and the training intensity control unit work together to form an adaptive training execution engine. The training scenario generation unit generates training scenarios that are theoretically adapted to the patient's abilities, while the training intensity control unit dynamically fine-tunes some parameters in the scenario based on the real-time fatigue index, such as temporarily reducing the weight of virtual objects to cope with a sudden increase in fatigue. The data from the processing is fed back to the dynamic reward decision module in real time to continuously determine the reward strategy for the next round, thereby continuously optimizing the system's decision-making accuracy and forming a complete closed loop of rehabilitation training: perception-decision-execution-optimization.

[0037] In summary, the stroke upper limb rehabilitation training system based on dynamic reward feedback provided in this application collects patients' physiological signals, motor signals, emotional state data, and training performance data through a data acquisition module to achieve synchronous acquisition and fusion of multi-dimensional and multi-modal data, providing a comprehensive and reliable data foundation for subsequent processing. In the motor intention decoding module, the physiological signals are used to extract motor intention features and assess reliability based on an improved MTREE-Net++ fusion algorithm. The multi-modal motor intention feature weights are adjusted based on the multi-modal reliability results, and weighted fusion is performed according to the multi-modal motor intention feature weights to obtain a unified motor intention feature vector. Motor intention decoding is then performed based on this unified motor intention feature vector to obtain a motor intention intensity index and a motor intention vector, thereby significantly improving the accuracy and robustness of motor intention decoding and effectively solving the decoding failure problem caused by the instantaneous quality degradation of a single signal, providing a refined structured input for subsequent decision-making. In the dynamic reward decision-making module, an improved VGG-Face algorithm is used... The model processes facial micro-expression data to obtain emotional valence. The improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. Emotional arousal is obtained based on speech signals, enabling accurate perception of the patient's emotional state and providing reliable emotional input for reward decisions. Based on emotional valence and emotional arousal, a reward action is generated using an improved deep reinforcement learning algorithm. An emotional decay factor is introduced to dynamically adjust the decay rate of the reward value, making the reward feedback more personalized and aligned with the patient's real-time psychological state, effectively maintaining and enhancing the patient's training motivation and participation. In the adaptive training control module, personalized virtual training scenarios are generated based on reward actions using a generative adversarial network to ensure that the training tasks are highly aligned with the patient's actual rehabilitation needs and life scenarios. A fatigue index is calculated based on real-time physiological signals, and the training intensity is adjusted accordingly to obtain personalized training tasks. This maximizes training efficiency and safety while avoiding excessive fatigue and injury, effectively balancing training intensity with the patient's physiological tolerance.

[0038] The above is a description of the system embodiments of this application. Based on the foregoing embodiments, the method embodiments of this application are described below.

[0039] Please refer to Figure 3 The diagram illustrates a flowchart of a stroke upper limb rehabilitation training method based on dynamic reward feedback, provided in an embodiment of this application. This method is applied to, for example... Figure 1 This diagram illustrates a stroke upper limb rehabilitation training system based on dynamic reward feedback. For details not disclosed in the method embodiment, please refer to the system embodiment. The system includes a data acquisition module, a motor intention decoding module, a dynamic reward decision-making module, and an adaptive training control module. (The diagram is repeated in the original text.) Figure 3 As shown, the method includes the following steps S310 to S340.

[0040] Step S310: Collect the patient's physiological signals, motion signals, emotional state data, and training performance data. The physiological signals include electroencephalogram (EEG) signals, blood oxygenation signals, electromyography (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and voice signals.

[0041] Step S320: Extract motion intent features and evaluate reliability of the physiological signal according to the improved MTREE-Net++ fusion algorithm, adjust the multimodal motion intent feature weights based on the multimodal reliability results, perform weighted fusion according to the multimodal motion intent feature weights to obtain a unified motion intent feature vector, and perform motion intent decoding according to the unified motion intent feature vector to obtain the motion intent intensity index and motion intent vector.

[0042] Step S330: Process the facial micro-expression data according to the improved VGG-Face model to obtain the emotional valence, wherein the improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer, and obtains the emotional arousal level according to the speech signal; based on the emotional valence and the emotional arousal level, generate a reward strategy according to the improved deep reinforcement learning algorithm.

[0043] Step S340: Generate personalized virtual training scenarios based on generative adversarial networks, calculate fatigue index based on real-time physiological signals, and adjust training intensity according to the fatigue index to obtain personalized training tasks.

[0044] It should be noted that, in the embodiments of this application, if the above-mentioned method for upper limb rehabilitation training based on dynamic reward feedback for stroke is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a magnetic disk, or an optical disk. Thus, the embodiments of this application are not limited to any specific hardware and software combination.

[0045] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the steps in any of the above embodiments of a stroke upper limb rehabilitation training method based on dynamic reward feedback. Correspondingly, embodiments of this application also provide a computer program product, which, when executed by a processor of an electronic device, is used to implement the steps in any of the above embodiments of a stroke upper limb rehabilitation training method based on dynamic reward feedback.

[0046] Based on the same technical concept, this application provides an electronic device for implementing a stroke upper limb rehabilitation training method based on dynamic reward feedback as described in the above method embodiments. Figure 4 This is a hardware entity diagram of an electronic device provided in an embodiment of this application, such as... Figure 4 As shown, the electronic device 400 includes a memory 410 and a processor 420. The memory 410 stores a computer program that can run on the processor 420. When the processor 420 executes the program, it implements the steps in any of the embodiments of this application of a method for upper limb rehabilitation training for stroke based on dynamic reward feedback.

[0047] The memory 410 is configured to store instructions and applications executable by the processor 420, and can also cache data to be processed or already processed by the processor 420 and various modules in the electronic device (e.g., image data, audio data, voice communication data and video communication data), which can be implemented by flash memory or random access memory (RAM).

[0048] When the processor 420 executes the program, it implements the steps of a stroke upper limb rehabilitation training method based on dynamic reward feedback, as described above. The processor 420 typically controls the overall operation of the electronic device 400.

[0049] The aforementioned processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), Central Processing Unit (CPU), Controller, Microcontroller, and Microprocessor. It is understood that other electronic devices can also implement the functions of the aforementioned processor, and this application does not specifically limit the specific implementation.

[0050] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various electronic devices that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0051] It should be noted that the descriptions of the storage medium and device embodiments above are similar to the descriptions of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0052] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0053] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0054] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0055] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0056] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0057] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause the device automatic test line to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.

[0058] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0059] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.

[0060] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A stroke upper limb rehabilitation training system based on dynamic reward feedback, characterized in that, The system includes: The data acquisition module is used to collect the patient's physiological signals, motion signals, emotional state data, and training performance data. The physiological signals include electroencephalogram (EEG) signals, blood oxygenation signals, electromyogram (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and voice signals. The motion intent decoding module is used to extract motion intent features and evaluate reliability of the physiological signal according to the improved MTREE-Net++ fusion algorithm, adjust the multimodal motion intent feature weights based on the multimodal reliability results, perform weighted fusion according to the multimodal motion intent feature weights to obtain a unified motion intent feature vector, and perform motion intent decoding according to the unified motion intent feature vector to obtain a motion intent intensity index and a motion intent vector. A dynamic reward decision module is used to process the facial micro-expression data according to the improved VGG-Face model to obtain the emotional valence. The improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. The emotional arousal level is obtained according to the speech signal. Based on the emotional valence level and the emotional arousal level, a reward action is generated according to an improved deep reinforcement learning algorithm. An adaptive training control module is used to generate personalized virtual training scenarios based on the reward action and a generative adversarial network, calculate a fatigue index based on real-time physiological signals, and adjust the training intensity according to the fatigue index to obtain personalized training tasks.

2. The system according to claim 1, characterized in that, The motion intent decoding module includes a signal cleanup unit, a feature fusion unit, and an intent decoding unit, wherein: The signal purification unit is used to adaptively denoise physiological signals based on an improved wavelet threshold denoising algorithm, filter out electrooculogram artifacts, and remove motion artifacts in conjunction with the motion signal. The feature fusion unit is used to extract multimodal motion intention features based on the physiological signal according to the improved MTREE-Net++ fusion algorithm, perform reliability assessment based on the multimodal motion intention features, dynamically adjust the weights of the multimodal motion intention features based on the multimodal reliability results, and perform weighted fusion of the multimodal motion intention features according to the weights of the multimodal motion intention features to obtain a unified motion intention feature vector. The intent decoding unit is used to calculate a motion intent intensity index based on the multimodal motion intent features and the multimodal motion intent features, and to perform motion intent decoding based on the motion intent intensity index and the unified motion intent feature vector to obtain a motion intent vector, wherein the motion intent vector includes action type, expected intensity and execution confidence.

3. The system according to claim 2, characterized in that, The feature fusion unit includes a feature extraction subunit, a reliability evaluation subunit, and a weight fusion subunit, wherein: The feature extraction subunit is used to calculate the degree of μ rhythm inhibition based on the EEG signal, calculate the rate of change of oxyhemoglobin concentration based on the blood oxygen signal, calculate the root mean square value and the root mean square rate of change of the electromyography signal based on the electromyography signal, and standardize the degree of μ rhythm inhibition, the rate of change of oxyhemoglobin concentration, the root mean square value and the root mean square rate of change of the electromyography signal. The reliability assessment subunit is used to determine the quality of physiological signals based on signal-to-noise ratio, cross-modal correlation, and physiological parameter range. When the quality of the physiological signal is lower than a preset threshold, it is considered an unreliable signal. Specifically, when the μ-rhythm signal-to-noise ratio of the EEG signal is lower than 1.2 or the correlation coefficient with the blood oxygenation signal characteristics is lower than 0.6, the EEG signal is considered unreliable. When the ratio of the change in oxyhemoglobin concentration to the change in deoxyhemoglobin concentration of the blood oxygenation signal exceeds the physiological range of 0.8-1.2, the blood oxygenation signal is considered unreliable. When the root mean square rate of change of the electromyography signal exceeds 1.5 times the initial reference value, the electromyography signal is considered unreliable. The weighted fusion subunit is used to dynamically adjust the weights of multimodal motion intent features based on the multimodal reliability results. The multimodal motion intent features are weighted and fused according to these weights to obtain a unified motion intent feature vector. The weight adjustment formula is shown below: ; In the formula, Indicates the first Adjusted weights for unreliable signals; Indicates the first Initial weights for unreliable signals; Indicates the first The current signal-to-noise ratio of an unreliable signal; Indicates the first The signal-to-noise ratio stability threshold for an unreliable signal; Indicates the first Adjusted weights of reliable signals; Indicates the first The initial weights of a reliable signal; This represents the total number of modes participating in the fusion.

4. The system according to claim 1, characterized in that, The dynamic reward decision module includes an emotion state perception unit and a reward decision unit, wherein: The emotional state perception unit is used to input the facial micro-expression data and the motion signal into the improved VGG-Face model. In the motion interference filtering layer, motion artifact correction is performed on the facial micro-expression data based on the motion signal. The corrected facial image is then input into a convolutional feature extraction network. The micro-expression enhancement layer enhances the local micro-expression features of the corners of the eyes and mouth. The emotional valence is obtained by mapping through a fully connected layer. The emotional arousal level is obtained by processing the speech signal through a regression model. The reward decision unit is used to construct a comprehensive state vector based on motion accuracy data, the emotional valence, the emotional arousal, the motion intention intensity index, and the motion intention vector. The comprehensive state vector is then processed using an improved deep Q-network algorithm to generate a reward action. The emotional decay factor in the deep Q-network is dynamically adjusted according to the emotional valence. When the emotional valence is less than -0.3, the emotional decay factor is 0.6; when the emotional valence is greater than or equal to -0.3, the emotional decay factor is 0.

9.

5. The system according to claim 4, characterized in that, The emotion state perception unit includes a motion interference filtering subunit, a micro-expression enhancement subunit, and a speech emotion perception subunit, wherein: The motion interference filtering subunit is used to calculate the head displacement compensation amount based on the motion signal and to perform motion artifact correction on the facial micro-expression data through affine transformation. The micro-expression enhancement subunit is used to process the corrected facial micro-expression data through a convolutional feature extraction network. The output features of the 4th convolutional layer are added to the input features of the 5th convolutional layer through residual connections. In the residual branch, 1×1 convolution dimensionality reduction and 3×3 convolution operations are performed in sequence to enhance the local micro-expression features of the corner of the eye and corner of the mouth. The enhanced features are then mapped to emotional valence through a fully connected layer. The speech emotion perception subunit is used to perform frame segmentation, windowing, fast Fourier transform, and Mel frequency cepstral coefficient feature extraction on the speech signal, and input it into a regression model to obtain the emotional arousal level.

6. The system according to claim 1, characterized in that, The adaptive training control module includes a training scenario generation unit and a training intensity control unit, wherein: The training scenario generation unit is used to generate personalized virtual training scenarios based on the assessment results of daily life activities, shoulder joint range of motion, and movement intention intensity index, according to a generative adversarial network. The training intensity control unit calculates a fatigue index based on real-time root mean square values ​​of electromyography (EMG) signals, changes in deoxyhemoglobin concentration, and changes in oxyhemoglobin concentration. It then adjusts the training intensity according to this fatigue index: when the fatigue index is greater than 0.7, the training intensity is reduced; when the fatigue index is less than 0.3, the training intensity is gradually increased; and when the fatigue index is greater than or equal to 0.3 and less than or equal to 0.7, the current training intensity is maintained, resulting in a personalized training task. The formula for calculating the fatigue index is shown below: ; In the formula, Indicates fatigue index; Indicates muscle fatigue weight; This represents the root mean square value of the electromyographic signal; Indicates the central fatigue weight; This represents the change in deoxyhemoglobin concentration; This represents the change in oxyhemoglobin concentration.

7. The system according to claim 1, characterized in that, The reward action includes reward type, reward intensity, and reward timing. The reward type is determined based on the comparison between the emotional valence and a preset threshold. When the emotional valence is greater than the preset threshold, an achievement-type reward is triggered; otherwise, an encouragement-type reward is triggered. The reward intensity is determined based on the motion accuracy data and the emotional arousal level. The reward timing is within 300ms after the action is completed.

8. A method for upper limb rehabilitation training after stroke based on dynamic reward feedback, characterized in that, An upper limb rehabilitation training system for stroke based on dynamic reward feedback is provided. The system includes a data acquisition module, a motor intention decoding module, a dynamic reward decision-making module, and an adaptive training control module. The method includes: The patient's physiological signals, motor signals, emotional state data, and training performance data are collected. The physiological signals include electroencephalogram (EEG) signals, blood oxygenation signals, electromyogram (EMG) signals, and heart rate data. The emotional state data includes facial micro-expression data and voice signals. The physiological signal is subjected to motion intention feature extraction and reliability assessment using the improved MTREE-Net++ fusion algorithm. The multimodal motion intention feature weights are adjusted based on the multimodal reliability results. Weighted fusion is performed based on the multimodal motion intention feature weights to obtain a unified motion intention feature vector. Motion intention is decoded based on the unified motion intention feature vector to obtain the motion intention intensity index and motion intention vector. The facial micro-expression data is processed using an improved VGG-Face model to obtain an emotional valence. The improved VGG-Face model includes a motion interference filtering layer and a micro-expression enhancement layer. Emotional arousal is obtained based on the speech signal. Based on the emotional valence and the emotional arousal, a reward strategy is generated using an improved deep reinforcement learning algorithm. Personalized virtual training scenarios are generated using generative adversarial networks, fatigue indices are calculated based on real-time physiological signals, and training intensity is adjusted according to the fatigue indices to obtain personalized training tasks.

9. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method of claim 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 8.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method for medical care robot

    CN114724224A

  • Transform-based face recognition method

    CN115311725A

  • Multi-source data-driven cerebral apoplexy upper limb rehabilitation virtual-real interaction method and system

    CN118098604A

  • Intelligent analysis system and method for prenatal state of puerpera

    CN119008016A

  • AI-assisted limb rehabilitation system

    CN119580933A

Cited By

  • Rehabilitation training adjustment method and system based on multi-modal feature fusion and storage medium

    CN122091086A