Rehabilitation training data mining system and method based on deep reinforcement learning
By combining deep reinforcement learning with brain-computer interfaces and other technologies, personalized and precise rehabilitation training has been achieved, solving the problem of traditional rehabilitation training relying on experience-based judgment and improving the efficiency and effectiveness of rehabilitation training.
Patent Information
- Application Number
- CN202411691285.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional rehabilitation training methods rely on the experience and judgment of doctors or therapists, lacking precision and individualization, resulting in poor rehabilitation outcomes.
A rehabilitation training data mining system based on deep reinforcement learning is adopted. By combining brain-computer interface, electromyography interface, force sensor and other technologies, the system actively identifies the movement intention of the training object and uses deep reinforcement learning model to adaptively adjust rehabilitation training parameters to achieve personalized and precise rehabilitation training.
It enables real-time, adaptive parameter adjustment during rehabilitation training, improving training efficiency and patient participation, reducing the workload of doctors' manual assessment, and enhancing the accuracy and individualization of rehabilitation outcomes.
Smart Images

Figure CN120932907A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of rehabilitation medicine technology, and in particular to a rehabilitation training data mining system and method based on deep reinforcement learning. Background Technology
[0002] Rehabilitation medicine, as an important branch of modern medicine, is dedicated to helping patients restore physical function and improve their quality of life through various treatment methods. However, traditional rehabilitation training methods often rely on the experience and judgment of doctors or therapists, lacking precision and individualization, and exhibiting certain subjectivity and limitations. In recent years, with the rapid development of information technology, especially the application of artificial intelligence, the accuracy and efficiency of rehabilitation training have been significantly improved. Deep reinforcement learning, as an advanced machine learning method, has attracted widespread attention because it can autonomously learn environmental rules and make optimal decisions. Nevertheless, its application in rehabilitation training is still in its early stages. Therefore, how to effectively apply deep reinforcement learning to rehabilitation training has become one of the key research focuses. Summary of the Invention
[0003] This invention addresses the problem that traditional rehabilitation training methods often rely on the experience and judgment of doctors or therapists, lacking precision and individualization. It proposes a rehabilitation training data mining system and method based on deep reinforcement learning, aiming to improve the effectiveness of rehabilitation training and personal experience through intelligent means.
[0004] To achieve the above objectives, the following technical solution is proposed: A rehabilitation training data mining system based on deep reinforcement learning includes: The intelligent cognition and parameter mining module is equipped with a deep reinforcement learning rehabilitation training model and a rehabilitation database. The deep reinforcement learning rehabilitation training model is trained using the rehabilitation database. The evaluation and analysis module evaluates and analyzes various physiological data collected from patients during rehabilitation training to obtain initial training parameters. The initial training parameters are input into the trained deep reinforcement learning rehabilitation training model to obtain the optimal rehabilitation training parameters. The human-computer interaction module includes a control module, which receives the optimal rehabilitation training parameters and generates control signals to the audiovisual feedback module and the rehabilitation training module.
[0005] This invention combines control methods such as brain-computer interfaces, electromyography (EMG) interfaces, and force sensors with pelvic floor magnetic stimulation, robotics, and exoskeletons. By pre-collecting various physiological signals from the training subject, it actively identifies the subject's movement intentions and assists them in performing specific rehabilitation exercises. This transforms the existing passive training method into an active one, achieving convenient, rapid, and comprehensive active rehabilitation. Furthermore, because the data acquisition module includes various physiological acquisition modules such as EEG acquisition and decoding, EMG acquisition and decoding, and muscle strength acquisition and decoding, the weights of EEG, EMG, and muscle strength signals allow for targeted analysis of training subjects at different rehabilitation stages, generating corresponding control methods. This makes it applicable to different training subjects at various stages, accelerating the rehabilitation process.
[0006] Preferably, the system also includes: a data acquisition and decoding module, an EEG acquisition and decoding module for acquiring EEG signals during patient rest and motor imagery, an EMG acquisition and decoding module for acquiring EMG signals during patient rest and movement, a muscle strength acquisition and decoding module for acquiring muscle strength data during patient rest and intention to move, and other sensor acquisition and decoding modules for acquiring other types of rehabilitation data.
[0007] Preferably, the evaluation and analysis module has a corresponding evaluation model for the type of data acquired by the data acquisition and decoding module; the evaluation and analysis module includes: an EEG evaluation model for acquiring EEG frequency band characteristic data of different frequency bands based on EEG signals; an EMG evaluation model for acquiring EMG characteristic data of different muscle parts based on EMG signals; and a muscle strength evaluation model for acquiring muscle strength characteristic data based on muscle strength signals.
[0008] Preferably, the intelligent cognition and parameter mining module includes a model training unit that trains a deep reinforcement learning model using historical data and a parameter mining unit that mines the optimal rehabilitation training parameters based on the model output, wherein the historical data is stored in a rehabilitation database.
[0009] Preferably, the deep reinforcement learning model employs the DoubleDQN algorithm.
[0010] Preferably, the human-computer interaction module includes a user information management submodule, a rehabilitation training data management submodule, and an external detection data management submodule. The user data generated by the user information management submodule, the rehabilitation training data management submodule, and the external detection data management submodule is stored in the rehabilitation database.
[0011] A rehabilitation training data mining method based on deep reinforcement learning, employing the aforementioned rehabilitation training data mining system based on deep reinforcement learning, is characterized by including the following steps: Collect various physiological data from users, including electroencephalogram (EEG) data, electromyogram (EMG) data, and muscle strength data; Physiological data evaluation indicators are obtained by evaluating various types of physiological data, and the initial threshold parameters are calculated using the physiological data evaluation indicators and the initial motion threshold calculation model. User information data, various physiological data, and initial threshold parameters are uploaded to the rehabilitation database. The deep reinforcement learning rehabilitation training model is trained using the rehabilitation database to discover the optimal rehabilitation training parameters of the deep reinforcement learning rehabilitation training model. The optimal rehabilitation training parameters are then input into the rehabilitation training scoring model to obtain the corresponding control signals for the audiovisual feedback module and the rehabilitation training module. The training data generated during the training process is input into the rehabilitation database, and the rehabilitation database is used to update the optimal rehabilitation training parameters of the deep reinforcement learning rehabilitation training model. After training is completed, user feedback data, training data, and result data are stored in the rehabilitation database.
[0012] As a preferred approach, the initial motion threshold calculation model is: Where: k1, k2, k3, and k4 are characteristic coefficients of EEG, EMG, muscle strength, and other features; a1 and a2 are EEG characteristic coefficients; F eeg_image It is the EEG characteristic of imagining or attempting to imagine, F eeg_rest These are the EEG characteristics at rest, where N is the length of the EEG data and F is the length of the data. eeg_means b1 and b2 are the current mean values of EEG characteristics; b1 and b2 are the electromyographic characteristic coefficients; F emg_motor It is the electromyographic characteristic during movement, F emg_rest These are the electromyographic characteristics at rest, where M is the length of the electromyographic data and F is the length of the data. emg_means It is the current mean electromyographic characteristic; c1 and c2 are the muscle strength characteristic coefficients; F ms_motor It refers to the muscle strength characteristics during the intentional movement state, F ms_rest It refers to the muscle strength characteristics at rest, where T is the length of the muscle strength characteristic, and F is the muscle strength characteristic at rest. ms_means d1 and d2 are the mean values of current muscle strength characteristics; d1 and d2 are the coefficients of other characteristics; F others_imageOrmotor It is imagination, attempting to imagine or other characteristics in a state of motion, F others_rest Other features in the resting state, S is the data length, F others_means It is the current feature mean.
[0013] Preferably, the rehabilitation training scoring model receives control signals generated from the output of the intelligent cognition and parameter mining module and transmits them to the control module for corresponding rehabilitation training.
[0014] Preferably, the rehabilitation training scoring model is: Where: Score is the rehabilitation training score; C is the preset auxiliary score; α is the score conversion coefficient; T is the active movement threshold; k1, k2, k3, and k4 are the feature coefficients of EEG, EMG, muscle strength, and other characteristics; F eeg It represents the characteristics of the brainwave state; N is the length of the brainwave data, F eeg_means This is the current average EEG characteristic; F emg It represents the electromyographic state characteristics, M is the length of the electromyographic data, and F is the value of the electromyographic data. emg_means This is the current mean value of electromyographic characteristics; F ms It is a muscle strength status feature, where T is the length of the muscle strength data, and F is the muscle strength status feature. ms_means This is the current mean of muscle strength characteristics; F others Other features, S is the length of other feature data, F others_means It is the mean of other characteristics.
[0015] The beneficial effects of this invention are: by using this system and method, it is possible to adjust training parameters in real time and adaptively according to the patient's condition, motor function and training status during rehabilitation training. This not only reduces the workload of doctors in manually assessing and adjusting parameters, but also makes the adjustment more dynamic and precise compared to fixed-cycle parameter adjustment, which is conducive to improving the efficiency of rehabilitation training and increasing patient participation and training experience. Attached Figure Description
[0016] Figure 1 This is a simplified diagram of the system configuration of the present invention.
[0017] Figure 2 This is a simplified flowchart of the method of the present invention. Detailed Implementation
[0018] Example 1: This embodiment proposes a rehabilitation training data mining system based on deep reinforcement learning, referencing... Figure 1 ,include: Data acquisition and decoding module: Collects various physiological signals from patients during rehabilitation training from different sensors and devices, decodes and extracts their features, and inputs them into the assessment and analysis module to generate an initial assessment model; Evaluation and analysis module: performs preliminary processing and analysis on the various collected signals, generates initial training parameters, and inputs them into the intelligent cognition and parameter mining module to generate a deep reinforcement learning rehabilitation training model; Intelligent cognition and parameter mining module: After receiving signals from the data acquisition and decoding module, evaluating and analyzing the signals from the user information management sub-module, external detection data management sub-module and rehabilitation training data management sub-module in the human-computer interaction module, the module uses a deep reinforcement learning model to mine the optimal rehabilitation training parameters. Human-computer interaction module: After receiving the signal from the data acquisition and decoding module, the control submodule transmits the control signal generated by the training parameters mined by the deep reinforcement learning model to the audiovisual feedback module and the rehabilitation training module.
[0019] The data acquisition and decoding module of this invention comprises: an EEG acquisition and decoding module that acquires and decodes EEG signals from the patient during rest and motor imagery, extracting useful feature information; an EMG acquisition and decoding module that acquires and decodes EMG signals from the patient during rest and movement, extracting useful feature information; a muscle strength acquisition and decoding module that acquires and decodes muscle strength data from the patient during rest and when intending to move, extracting useful feature information; and other sensor acquisition and decoding modules that acquire and decode other types of rehabilitation data, such as heart rate and blood pressure, extracting useful feature information. The EEG acquisition and decoding module decodes the trainee's motor imagery intentions in real time by receiving the EEG information from the trainee. The electromyography (EMG) acquisition and decoding module receives the EMG information from the training subject and decodes the subject's movement intention in real time. The muscle force acquisition and decoding module receives dynamic muscle force information from the training subject as they intend to move their limbs in the positive and negative directions of each coordinate axis in a three-dimensional Cartesian coordinate system, and decodes the force data of the training subject in real time. Other sensor acquisition and decoding modules receive other physiological information (such as heart rate and blood pressure) from the training subjects and decode the physiological information features of the training subjects in real time.
[0020] The EEG acquisition and decoding module uses professional EEG equipment to non-invasively record the patient's brain activity during rest and imagined movement. The raw EEG signal is preprocessed (e.g., filtered, denoised), and then time-frequency analysis techniques are applied to extract feature vectors, such as power spectral density at different frequency bands (δ, θ, α, β, γ). The electromyography (EMG) acquisition and decoding module captures changes in the electrical signals of muscles during rest and actual movement using a surface EMG sensor. Similarly, preprocessing (such as bandpass filtering) is performed first, and then time-domain features (such as mean absolute value, slope change exponent) or frequency-domain features (such as center frequency, root mean square) are used to characterize the activity state of the muscles. The muscle strength acquisition and decoding module uses a high-precision force gauge to measure the force exerted by the patient when attempting to move their limb. Muscle strength characteristic values can be directly obtained through simple mathematical calculations. Other sensor acquisition and decoding modules integrate information from additional health monitoring devices such as heart rate monitors and blood pressure monitors to provide a comprehensive understanding of the patient's overall physiological condition. This data also requires appropriate preprocessing steps before it can be used for subsequent analysis.
[0021] The evaluation and analysis module of the present invention calculates an evaluation model based on the physiological signal characteristics of the training object when it is relaxed, when it is imagining movement, when it is moving, and when it intends to move, which are obtained from the data acquisition and decoding module.
[0022] The evaluation and analysis module of this invention includes at least an electroencephalogram (EEG) evaluation model, an electromyogram (EMG) evaluation model, and a muscle strength evaluation model.
[0023] The EEG assessment model acquires EEG frequency band feature data of different frequency bands based on EEG signals, extracts the motor imagination state features and rest state features of the training subjects, and substitutes them into the EEG threshold calculation model to obtain the optimal EEG threshold for each EEG frequency band.
[0024] The electromyography (EMG) assessment model obtains EMG feature data of different muscle parts based on EMG signals, extracts the movement state features and rest state features of the training object, and substitutes them into the EMG threshold calculation model to obtain the optimal EMG threshold for each muscle part.
[0025] The muscle strength assessment model obtains muscle strength feature data based on muscle strength signals, extracts the features of the training object's intended movement state and resting state, and substitutes them into the muscle strength threshold calculation model to obtain the optimal muscle strength threshold.
[0026] In the intelligent cognition and parameter mining module: the rehabilitation database receives and stores signals from the data acquisition and decoding module, assessment and analysis module, user information management module, external detection data management module, and rehabilitation training data management module. It uses historical data combined with real-time generated data to train a deep reinforcement learning model, enabling it to adaptively adjust rehabilitation parameters. Based on the model's output, it mines the optimal rehabilitation training parameters. Training data generated during training is processed by the rehabilitation training data management module and combined with previous data from the rehabilitation database, then input as indicators back into the deep reinforcement learning model to train the model and mine the optimal rehabilitation training parameters. This continuous optimization of training parameters leads to more efficient rehabilitation results. The rehabilitation training scoring model receives control signals from the output of the intelligent cognition and parameter mining module, which are then transmitted to the control module for corresponding rehabilitation training.
[0027] The intelligent cognition and parameter mining module is the core of the entire system, and its main functions include: Rehabilitation database: Stores patients' historical rehabilitation data, assessment data, and training data; Deep reinforcement learning model: Employs the DoubleDQN algorithm to mine suitable rehabilitation training parameters for patients from the rehabilitation database; Model training unit: Trains the deep reinforcement learning model using historical data, enabling it to adaptively adjust rehabilitation parameters; Parameter mining unit: Mines the optimal rehabilitation training parameters based on the model output. The DoubleDQN algorithm is employed as the core technology for intelligent cognition and parameter mining. During training, the generated training data is processed by the rehabilitation training data management module and combined with previous data from the rehabilitation database. This data is then used as indicators to input into the deep reinforcement learning model, training the model and identifying optimal rehabilitation training parameters. Continuous optimization of these parameters leads to more efficient rehabilitation outcomes. The rehabilitation training scoring model receives control signals from the intelligent cognition and parameter mining module, which are then transmitted to the control module for corresponding rehabilitation training.
[0028] This embodiment also proposes a rehabilitation training data mining method based on deep reinforcement learning, employing the aforementioned rehabilitation training data mining system based on deep reinforcement learning, with reference to... Figure 2 This includes the following steps: 1) Collect patient information and data from external testing equipment; 2) Collect relevant biological signals such as EEG, EMG, and muscle strength from the patient using equipment for acquiring related physiological signals such as EEG, EMG, and muscle strength; 3) The patient's EEG assessment index is obtained using the EEG assessment model, the patient's EMG assessment index is obtained using the EMG assessment model, the patient's muscle strength assessment index is obtained using the muscle strength assessment model, etc. Finally, the initial threshold parameters are obtained by combining various indicators and using the initial threshold calculation model. The initial motion threshold calculation model of the present invention is as follows: Where T is the characteristic coefficient of EEG, EMG, muscle strength, and other features, and the sum of the four is 1; a1 and a2 are the characteristic coefficients of EEG, and the sum of the two is 1; F is the characteristic coefficient of EEG. eeg_image For the brainwave characteristics of imagining or attempting to imagine a state, F eeg _ rest The EEG characteristics are shown at rest, where N is the length of the EEG data and F is the length of the data. eeg_means b1 and b2 are the mean values of the current EEG characteristics; b1 and b2 are the electromyographic characteristic coefficients, and their sum is 1; F emg_motor For electromyographic characteristics during movement, F emg_rest The electromyography (EMG) characteristics are shown at rest, where M is the length of the EMG data and F is the length of the data. emg_means The current mean electromyographic characteristics are given; c1 and c2 are the muscle strength characteristic coefficients, and their sum is 1; F ms_motorF represents the muscle strength characteristics during the intentional movement state. ms_rest The muscle strength characteristics at rest, where T is the length of the muscle strength characteristic, and F is the muscle strength characteristic at rest. ms_means d1 and d2 are the mean values of the current muscle strength characteristics; d1 and d2 are the coefficients of other characteristics, and their sum is 1; F others_imageOrmotor For imagining, attempting to imagine, or other characteristics in a state of motion, F others_rest Other features in the resting state, S is the data length, F others_means This represents the current feature mean.
[0029] During the evaluation process, the system calculates and updates the relevant feature results under various states in real time and provides feedback to the training subjects in the form of line graphs, energy maps, games, etc. through the display of the audiovisual feedback module 441.
[0030] 4) Upload the patient information data, related EEG, EMG, muscle strength and other biological signals, and initial threshold parameters obtained in steps 1), 2), and 3) to the rehabilitation database, and input them as indicators into our deep reinforcement learning model to train the deep reinforcement learning model and discover the optimal rehabilitation training parameters. 5) The output of step 4) is used as the input to the rehabilitation training scoring model, and then relevant control signals are generated and transmitted to the control module for corresponding rehabilitation training. The rehabilitation training scoring model receives control signals generated from the output of the intelligent cognition and parameter mining module and transmits them to the control module for corresponding rehabilitation training.
[0031] The rehabilitation training scoring model of the present invention is as follows: Where Score is the rehabilitation training score; C is the preset auxiliary score; α is the score conversion coefficient; T is the active movement threshold; k1, k2, k3, and k4 are the feature coefficients of EEG, EMG, muscle strength, and other features, and the sum of the four is 1; F eeg The EEG state characteristics; N is the length of the EEG data, F eeg_means The mean of the current EEG characteristics; F emg For electromyographic state characteristics, M is the length of the electromyographic data, and F is the value of the electromyographic data. emg_means The mean of the current electromyographic characteristics; F ms For muscle strength status features, T is the length of muscle strength data, and F is the value of the muscle strength data. ms_means F represents the mean of the current muscle strength characteristics. others For other features, S is the length of the other feature data, and F is the length of the other feature data. others_meansThe mean of other features. It should be noted that the preset auxiliary score can be set according to actual needs. This auxiliary score can be used as a preset score to judge various motor states; for example, it will be modified according to the different types of EEG frequency bands considered, and will be modified according to the location of different muscles, etc.
[0032] 6) Based on the training data generated during the training process in step 5), combine it with the previous data in the rehabilitation database as an indicator and input it into our deep reinforcement learning model again to train the deep reinforcement learning model and discover the optimal rehabilitation training parameters, so as to continuously optimize our training parameters and achieve a more efficient rehabilitation effect. 7) After training is completed, the feedback data is stored in the rehabilitation database based on user feedback, and finally the relevant training data and result data are uploaded to the rehabilitation database.
[0033] The rehabilitation training system is activated; online assessments are conducted, including EEG, EMG, muscle strength, and external testing instrument assessments; the assessment data is stored in the rehabilitation database; suitable rehabilitation training parameters are mined using an intelligent cognitive and parameter mining algorithm (using DoubleDQN); rehabilitation training begins based on the mined parameters; during training, real-time data is combined with past data in the rehabilitation database, and the relevant parameters are dynamically adjusted using the intelligent cognitive and parameter mining algorithm; the operator can manually adjust the relevant parameters based on real-time training data; after training, user feedback information is generated and stored in the rehabilitation database.
[0034] Example 2: This embodiment, based on Embodiment 1, defines the intelligent cognition and parameter mining algorithm and proposes a rehabilitation training data mining system based on deep reinforcement learning, including: Data acquisition and decoding module: Collects various physiological signals from patients during rehabilitation training from different sensors and devices, decodes and extracts their features, and inputs them into the assessment and analysis module to generate an initial assessment model; Evaluation and analysis module: performs preliminary processing and analysis on the various collected signals, generates initial training parameters, and inputs them into the intelligent cognition and parameter mining module to generate a deep reinforcement learning rehabilitation training model; Intelligent cognition and parameter mining module: After receiving signals from the data acquisition and decoding module, evaluating and analyzing the signals from the user information management sub-module, external detection data management sub-module and rehabilitation training data management sub-module in the human-computer interaction module, the module uses a deep reinforcement learning model to mine the optimal rehabilitation training parameters. Human-computer interaction module: After receiving the signal from the data acquisition and decoding module, the control submodule transmits the control signal generated by the training parameters mined by the deep reinforcement learning model to the audiovisual feedback module and the rehabilitation training module.
[0035] The data acquisition and decoding module of this invention comprises: an EEG acquisition and decoding module that acquires EEG signals from the patient during rest and motor imagery, and performs decoding processing to extract useful feature information; an EMG acquisition and decoding module that acquires EMG signals from the patient during rest and movement, and performs decoding processing to extract useful feature information; a muscle strength acquisition and decoding module that acquires muscle strength data from the patient during rest and when intending to move, and performs decoding processing to extract useful feature information; and other sensor acquisition and decoding modules that acquire other types of rehabilitation data such as heart rate and blood pressure, and perform decoding processing to extract useful feature information.
[0036] The evaluation and analysis module of the present invention calculates an evaluation model based on the physiological signal characteristics of the training object when it is relaxed, when it is imagining movement, when it is moving, and when it intends to move, which are obtained from the data acquisition and decoding module.
[0037] The evaluation and analysis module of this invention includes at least an electroencephalogram (EEG) evaluation model, an electromyogram (EMG) evaluation model, and a muscle strength evaluation model.
[0038] The EEG assessment model acquires EEG frequency band feature data of different frequency bands based on EEG signals, extracts the motor imagination state features and rest state features of the training subjects, and substitutes them into the EEG threshold calculation model to obtain the optimal EEG threshold for each EEG frequency band.
[0039] The electromyography (EMG) assessment model obtains EMG feature data of different muscle parts based on EMG signals, extracts the movement state features and rest state features of the training object, and substitutes them into the EMG threshold calculation model to obtain the optimal EMG threshold for each muscle part.
[0040] The muscle strength assessment model obtains muscle strength feature data based on muscle strength signals, extracts the features of the training object's intended movement state and resting state, and substitutes them into the muscle strength threshold calculation model to obtain the optimal muscle strength threshold.
[0041] In the intelligent cognition and parameter mining module: the rehabilitation database receives and stores signals from the data acquisition and decoding module, assessment and analysis module, user information management module, external detection data management module, and rehabilitation training data management module. It uses historical data combined with real-time generated data to train a deep reinforcement learning model, enabling it to adaptively adjust rehabilitation parameters. Based on the model's output, it mines the optimal rehabilitation training parameters. Training data generated during training is processed by the rehabilitation training data management module and combined with previous data from the rehabilitation database, then input as indicators back into the deep reinforcement learning model to train the model and mine the optimal rehabilitation training parameters. This continuous optimization of training parameters leads to more efficient rehabilitation results. The rehabilitation training scoring model receives control signals from the output of the intelligent cognition and parameter mining module, which are then transmitted to the control module for corresponding rehabilitation training.
[0042] The intelligent cognition and parameter mining module is the core of the entire system, and its main functions include: Rehabilitation database: Stores patients' historical rehabilitation data, assessment data, and training data; Deep reinforcement learning model: Employs the DoubleDQN algorithm to mine suitable rehabilitation training parameters for patients from the rehabilitation database; Model training unit: Trains the deep reinforcement learning model using historical data, enabling it to adaptively adjust rehabilitation parameters; Parameter mining unit: Mines the optimal rehabilitation training parameters based on the model output. The DoubleDQN algorithm is employed as the core technology for intelligent cognition and parameter mining. During training, the generated training data is processed by the rehabilitation training data management module and combined with previous data from the rehabilitation database. This data is then used as indicators to input into the deep reinforcement learning model, training the model and identifying optimal rehabilitation training parameters. Continuous optimization of these parameters leads to more efficient rehabilitation outcomes. The rehabilitation training scoring model receives control signals from the intelligent cognition and parameter mining module, which are then transmitted to the control module for corresponding rehabilitation training.
[0043] The deep reinforcement learning algorithm is as follows: State Definition: In a rehabilitation training scenario, "state" S refers to a multidimensional vector composed of various physiological signals such as electroencephalogram (EEG), electromyography (EMG), and muscle strength. These signals are acquired through a data acquisition and decoding module and converted into digital form for computer processing. Specifically, state S can be represented as: Where s1, s2, ..., s nCorrespond to different physiological signal characteristic values respectively. For example, s1 represents the electroencephalogram alpha wave characteristic signal, s2 represents the electromyogram characteristic signal of the biceps brachii, etc., a1 and a2 are electroencephalogram characteristic coefficients, and the sum of the two is 1; F eeg_α_image is the electroencephalogram alpha wave characteristic in the state of imagination or attempted imagination, F eeg_α_rest is the electroencephalogram alpha wave characteristic in the resting state, N is the length of the electroencephalogram data, F eeg_α_means is the current electroencephalogram characteristic mean value; b1 and b2 are electromyogram characteristic coefficients, and the sum of the two is 1; F emg_biceps_motor is the electromyogram characteristic in the movement state of the biceps brachii, F emg_biceps_rest is the electromyogram characteristic in the resting state of the biceps brachii, M is the length of the electromyogram data, F emg_biceps_means is the current electromyogram characteristic mean value.
[0044] Action definition: "Action" A represents various measures or parameter settings that may be taken during the rehabilitation training process, such as adjusting the working mode, intensity, frequency, stimulation time interval, movement trajectory, etc. of the rehabilitation equipment. Specifically, the action space can be defined as: where, a1, a2,..., a m represent different rehabilitation training actions respectively. For example, a1 is to change the stimulation time interval of the magnetic stimulator, and a2 is to adjust the movement trajectory of the rehabilitation robot, etc.
[0045] Reward definition: "Reward" R is used to quantify the impact of a certain action on the rehabilitation effect. It can be determined according to the patient's ability to complete specific tasks, the degree of pain reduction, or other clinical indicators. For example, if the patient can successfully complete a certain rehabilitation task, a positive reward can be obtained; otherwise, a negative reward is given. The reward function can be defined as: R = (s, a) where, s is the current state and a is the executed action. The design of the reward function should take into account the goals of the rehabilitation training, such as increasing muscle strength, improving joint range of motion, etc. For different rehabilitation goals, the specific form of the reward function will also be different.
[0046] Calculation and update of Q value: In the deep reinforcement learning algorithm, the Q value represents the expected return that can be obtained by executing a certain action in a given state. Its calculation formula is as follows: Q(s, a) = r + γmax a′ Q(s′, a′) where, s represents the current state, a is the executed action, r is the immediate reward, γ is the discount factor (0 < γ < 1), s′ is the next state, and a′ is the action that may be taken in the next state.
[0047] It is worth noting that this is a deep reinforcement learning model specifically designed for rehabilitation training. The states include, but are not limited to, various physiological signals such as brain waves, electromyography signals, and muscle strength; the actions include, but are not limited to, stimulus intensity, frequency, and movement trajectory.
[0048] This embodiment also proposes a rehabilitation training data mining method based on deep reinforcement learning, which employs the aforementioned rehabilitation training data mining system based on deep reinforcement learning and includes the following steps: 1) Collect patient information and data from external testing equipment; 2) Collect relevant biological signals such as EEG, EMG, and muscle strength from the patient using equipment for acquiring related physiological signals such as EEG, EMG, and muscle strength; 3) The patient's EEG assessment index is obtained using the EEG assessment model, the patient's EMG assessment index is obtained using the EMG assessment model, the patient's muscle strength assessment index is obtained using the muscle strength assessment model, etc. Finally, the initial threshold parameters are obtained by combining various indicators and using the initial threshold calculation model. The initial motion threshold calculation model of the present invention is as follows: Where T is the characteristic coefficient of EEG, EMG, muscle strength, and other features, and the sum of the four is 1; a1 and a2 are the characteristic coefficients of EEG, and the sum of the two is 1; F is the characteristic coefficient of EEG. eeg_image For the brainwave characteristics of imagining or attempting to imagine a state, F eeg_rest The EEG characteristics are shown at rest, where N is the length of the EEG data and F is the length of the data. eeg_means b1 and b2 are the mean values of the current EEG characteristics; b1 and b2 are the electromyographic characteristic coefficients, and their sum is 1; F emg_motor For electromyographic characteristics during movement, F emg_rest The electromyography (EMG) characteristics are shown at rest, where M is the length of the EMG data and F is the length of the data. emg_means The current mean electromyographic characteristics are given; c1 and c2 are the muscle strength characteristic coefficients, and their sum is 1; F ms_motor F represents the muscle strength characteristics during the intentional movement state. ms_rest The muscle strength characteristics at rest, where T is the length of the muscle strength characteristic, and F is the muscle strength characteristic at rest. ms_means d1 and d2 are the mean values of the current muscle strength characteristics; d1 and d2 are the coefficients of other characteristics, and their sum is 1; F others_imageOrmotor For imagining, attempting to imagine, or other characteristics in a state of motion, F others_rest Other features in the resting state, S is the data length, F others_means This represents the current feature mean.
[0049] 4) Upload the patient information data, related EEG, EMG, muscle strength and other biological signals, and initial threshold parameters obtained in steps 1), 2), and 3) to the rehabilitation database, and input them as indicators into our deep reinforcement learning model to train the deep reinforcement learning model and discover the optimal rehabilitation training parameters. 5) The output of step 4) is used as the input to the rehabilitation training scoring model, and then relevant control signals are generated and transmitted to the control module for corresponding rehabilitation training. The rehabilitation training scoring model receives control signals generated from the output of the intelligent cognition and parameter mining module and transmits them to the control module for corresponding rehabilitation training.
[0050] The rehabilitation training scoring model of the present invention is as follows: Where Score is the rehabilitation training score; C is the preset auxiliary score; α is the score conversion coefficient; T is the active movement threshold; k1, k2, k3, and k4 are the feature coefficients of EEG, EMG, muscle strength, and other features, and the sum of the four is 1; F eeg The EEG state characteristics; N is the length of the EEG data, F eeg_means The mean of the current EEG characteristics; F emg For electromyographic state characteristics, M is the length of the electromyographic data, and F is the value of the electromyographic data. emg_means The mean of the current electromyographic characteristics; F ms For muscle strength status features, T is the length of muscle strength data, and F is the value of the muscle strength data. ms_means F represents the mean of the current muscle strength characteristics. others For other features, S is the length of the other feature data, and F is the length of the other feature data. others_means The mean of other features. It should be noted that the preset auxiliary score can be set according to actual needs. This auxiliary score can be used as a preset score to judge various motor states; for example, it will be modified according to the different types of EEG frequency bands considered, and will be modified according to the location of different muscles, etc.
[0051] 6) Based on the training data generated during the training process in step 5), combine it with the previous data in the rehabilitation database as indicators and input it into our deep reinforcement learning model again to train the deep reinforcement learning model and discover the optimal rehabilitation training parameters. Continuously optimize our training parameters to achieve a more efficient rehabilitation effect; 7) After training is completed, store the feedback data in the rehabilitation database based on user feedback, and finally upload the relevant training data and result data to the rehabilitation database.
[0052] The rehabilitation training system is activated; online assessments are conducted, including EEG, EMG, muscle strength, and external testing instrument assessments; the assessment data is stored in the rehabilitation database; suitable rehabilitation training parameters are mined using an intelligent cognitive and parameter mining algorithm (using DoubleDQN); rehabilitation training begins based on the mined parameters; during training, real-time data is combined with past data in the rehabilitation database, and the relevant parameters are dynamically adjusted using the intelligent cognitive and parameter mining algorithm; the operator can manually adjust the relevant parameters based on real-time training data; after training, user feedback information is generated and stored in the rehabilitation database.
Claims
1. A rehabilitation training data mining system based on deep reinforcement learning, characterized in that, include: The intelligent cognition and parameter mining module is equipped with a deep reinforcement learning rehabilitation training model and a rehabilitation database. The deep reinforcement learning rehabilitation training model is trained using the rehabilitation database. The assessment and analysis module evaluates and analyzes various physiological data collected from patients during rehabilitation training to obtain initial training parameters. The initial training parameters are input into the trained deep reinforcement learning rehabilitation training model to obtain the optimal rehabilitation training parameters; The human-computer interaction module includes a control module, which receives optimal rehabilitation training parameters and generates control signals to the audiovisual feedback module and the rehabilitation training module.
2. The rehabilitation training data mining system based on deep reinforcement learning according to claim 1, characterized in that, include: The data acquisition and decoding module includes an EEG acquisition and decoding module for acquiring EEG signals during patient rest and motor imagery, an EMG acquisition and decoding module for acquiring EMG signals during patient rest and movement, a muscle strength acquisition and decoding module for acquiring muscle strength data during patient rest and intention to move, and other sensor acquisition and decoding modules for acquiring other types of rehabilitation data.
3. The rehabilitation training data mining system based on deep reinforcement learning according to claim 2, characterized in that, The evaluation and analysis module has corresponding evaluation models for the types of data collected by the data acquisition and decoding module; the evaluation and analysis module includes: an EEG evaluation model that acquires EEG frequency band characteristic data of different frequency bands based on EEG signals; an EMG evaluation model that acquires EMG characteristic data of different muscle parts based on EMG signals; and a muscle strength evaluation model that acquires muscle strength characteristic data based on muscle strength signals.
4. The rehabilitation training data mining system based on deep reinforcement learning according to claim 1, characterized in that, The intelligent cognition and parameter mining module includes a model training unit that trains a deep reinforcement learning model using historical data and a parameter mining unit that mines the optimal rehabilitation training parameters based on the model output. The historical data is stored in a rehabilitation database.
5. A rehabilitation training data mining system based on deep reinforcement learning according to claim 4, characterized in that, The deep reinforcement learning model employs the DoubleDQN algorithm.
6. A rehabilitation training data mining system based on deep reinforcement learning according to any one of claims 1-5, characterized in that, The human-computer interaction module includes a user information management submodule, a rehabilitation training data management submodule, and an external detection data management submodule. The user data generated by the user information management submodule, the rehabilitation training data management submodule, and the external detection data management submodule is stored in the rehabilitation database.
7. A method for mining rehabilitation training data based on deep reinforcement learning, employing the rehabilitation training data mining system based on deep reinforcement learning as described in claim 1, characterized in that, Includes the following steps: Collect various physiological data from users, including electroencephalogram (EEG) data, electromyogram (EMG) data, muscle strength data, and other characteristic data; Physiological data evaluation indicators are obtained by evaluating various types of physiological data, and the initial threshold parameters are calculated using the physiological data evaluation indicators and the initial motion threshold calculation model. User information data, various physiological data, and initial threshold parameters are uploaded to the rehabilitation database. The deep reinforcement learning rehabilitation training model is trained using the rehabilitation database to discover the optimal rehabilitation training parameters of the deep reinforcement learning rehabilitation training model. The optimal rehabilitation training parameters are then input into the rehabilitation training scoring model to obtain the corresponding control signals for the audiovisual feedback module and the rehabilitation training module. The training data generated during the training process is input into the rehabilitation database, and the rehabilitation database is used to update the optimal rehabilitation training parameters of the deep reinforcement learning rehabilitation training model. After training is completed, user feedback data, training data, and result data are stored in the rehabilitation database.
8. The rehabilitation training data mining method based on deep reinforcement learning according to claim 7, characterized in that, The initial motion threshold calculation process is as follows: Calculate the ratio of various physiological data to their data length in the state of imagination or attempted imagination and the ratio of their data length in the resting state. Obtain the characteristic mean of various physiological data. Set a first characteristic coefficient for the ratio of their data length in the state of imagination or attempted imagination and the ratio of their data length in the resting state, and a second characteristic coefficient for various physiological data. Calculate the product of the first characteristic coefficient and the set ratio of their data length in the state of imagination or attempted imagination, and sum the product of the first characteristic coefficient and the ratio of their data length in the resting state to obtain a first calculation result. Divide the first calculation result of various physiological data by the characteristic mean of various physiological data to obtain a second calculation result. The initial exercise threshold is equal to the sum of the products of the second calculation result of various physiological data and the second characteristic coefficient.
9. A method for mining rehabilitation training data based on deep reinforcement learning according to claim 7, characterized in that, The rehabilitation training scoring model receives control signals generated from the output of the intelligent cognition and parameter mining module and transmits them to the control module for corresponding rehabilitation training.
10. A method for mining rehabilitation training data based on deep reinforcement learning according to claim 8, characterized in that, The rehabilitation training score calculation process is as follows: Calculate the ratio of the state characteristics of various physiological data to its data length, multiply the ratio by the second feature coefficient, and then divide it by the mean of the state characteristics to obtain the third calculation result. Add the third calculation results of various physiological data and then subtract the active movement threshold to obtain the fourth calculation result. The rehabilitation training score is equal to the product of the score conversion coefficient and the fourth calculation result and the sum of the preset auxiliary score.
Citation Information
Cited By
Rehabilitation evaluation and training method for neurosurgical patient based on reinforcement learning
CN121565464A
Behavior pattern mining and guiding method and system based on user data
CN121722832A
User Data-Based Behavioral Pattern Mining and Guidance Methods and Systems
CN121722832B