An upper limb motion rehabilitation system based on emotional interaction
The upper limb motor rehabilitation system based on emotional interaction utilizes multi-sensor and deep learning to simulate interpersonal interaction, solving the problem of lack of emotional support in rehabilitation equipment, improving rehabilitation efficiency and patients' positive emotions, and shortening the rehabilitation cycle.
Patent Information
- Application Number
- CN202510103384.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Existing rehabilitation equipment cannot replace the emotional support of rehabilitation physicians, resulting in low rehabilitation efficiency for patients and a tedious and monotonous rehabilitation process.
Design an upper limb motor rehabilitation system based on emotional interaction. Utilize multi-sensor fusion and deep learning to simulate interpersonal interaction, and combine visual and auditory devices to provide patients with an immersive rehabilitation experience. Simulate the emotional support of rehabilitation physicians through virtual characters and voice interaction.
It improved patients' rehabilitation outcomes and training motivation, shortened the rehabilitation cycle, and enhanced their quality of life.
Smart Images

Figure CN119943273B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure belongs to the field of elderly care and rehabilitation equipment and human-computer interaction technology, and particularly relates to an upper limb movement rehabilitation system based on emotional interaction. BACKGROUND
[0002] With the development of global deep aging, the number of disabled, demented and cognitively impaired elderly people is constantly expanding. Most of them have motor disorders and are subjected to physical and psychological torture. Medical theory and clinical medicine have proved that correct and scientific rehabilitation training plays a very important role in the recovery and improvement of motor function. To solve the problems of lack of professional nursing personnel and high medical costs, safe, quantitative, effective and repeatable rehabilitation training devices have emerged. However, the rehabilitation process brought by pure mechanical equipment is often dull and monotonous, and the existing rehabilitation equipment still cannot replace the emotional support provided by doctors in the patient's rehabilitation process, which greatly affects the patient's emotional state and their rehabilitation efficiency.
[0003] An excellent rehabilitation doctor can not only flexibly adjust the training task according to the current condition of the patient to stimulate their rehabilitation confidence and motivation, but also can provide them with great emotional help through appropriate communication and exchange. Good doctor-patient relationship has been proved to play a crucial role in the improvement of medical results including motor learning performance, therefore, to solve the problem that the existing rehabilitation device still cannot replace the rehabilitation doctor, to endow the robot with emotional interaction ability comparable to the rehabilitation doctor, to introduce the positive effects brought by interpersonal emotional exchange, to create an immersive and positive rehabilitation environment to enhance the rehabilitation effect, is an important development direction of motor rehabilitation technology, but there is no corresponding technology formed at present. SUMMARY
[0004] The present disclosure aims to at least solve one of the technical problems existing in the prior art to some extent.
[0005] To this end, the present disclosure provides an upper limb movement rehabilitation system based on emotional interaction, which uses multi-sensor fusion, deep learning and generative large language model to simulate the positive emotional effects in interpersonal interaction process, adds interpersonal emotional interaction function in the rehabilitation system design, and combines visual, auditory interaction devices and rehabilitation training devices to provide patients with immersive rehabilitation experience, stimulate patients' positive emotions and training motivation, thereby improving the rehabilitation effect and promoting the recovery of upper limb movement function.
[0006] To achieve the above purpose, the upper limb movement rehabilitation system based on emotional interaction provided by the present disclosure comprises:
[0007] a sensing unit configured to acquire physiological signals and behavior signals of a patient;
[0008] The central controller comprises a data processing module and an interaction strategy generation module; the data processing module is used to determine the physical state and psychological state of the patient in real time according to the physiological signals and behavior signals obtained by the sensing unit, and outputs the determination result as the input of the interaction strategy generation module; the interaction strategy generation module simulates the professional knowledge and thinking mode of a rehabilitation physician to understand and fuse the current state of the patient and make a decision, and outputs a multi-dimensional interaction strategy of the patient, including real-time generated visual and voice interaction content and upper limb rehabilitation training content.
[0009] The execution unit comprises a loudspeaker, a visual interaction device and an upper limb rehabilitation robot; the loudspeaker is used to play the voice with the tonal characteristics familiar to the patient and the voice with the positive emotional style according to the voice interaction content generated by the interaction strategy generation module during the training; the visual interaction device is used to generate a virtual character according to the visual interaction content generated by the interaction strategy generation module, as one of the carriers of simulated interpersonal emotional interaction and support, and make the virtual character produce corresponding facial movements and expressions with the voice generated by the loudspeaker; the upper limb rehabilitation robot is used to guide the patient to complete the upper limb rehabilitation training action according to the upper limb rehabilitation training content generated by the interaction strategy generation module.
[0010] In some embodiments, the upper limb rehabilitation robot can guide the patient to complete a two-dimensional plane from simple to complex motion trajectory, with passive, assisted and active training modes and difficulty levels; the change of the difficulty level is realized by adjusting the complexity of the motion trajectory, the training speed in the passive and assisted mode training, and the resistance value in the active mode training; the upper limb rehabilitation robot feeds back the information including the motion position, speed and interaction force of the mechanism itself to the sensing unit through the internal joint force sensors and motor encoders, as the basic data for quantitative evaluation of the training task performance of the patient.
[0011] In some embodiments, the physiological signals obtained by the sensing unit include electroencephalogram, electromyogram, electrocardiogram and skin electricity signals, and the behavior signals obtained by the sensing unit include facial expression, voice, eye movement and motion task performance information; the process of determining the psychological state of the patient in real time by the data processing module according to the physiological and behavior signals obtained by the sensing unit comprises: respectively pre-processing and preliminarily extracting the obtained signals; then, deep feature extraction is performed on the extracted preliminary features by using a classification model based on deep learning to obtain the psychological state of the patient.
[0012] In some embodiments, the psychological state of the patient is divided into emotional state and mental load degree, the emotional state is divided into three levels of positive, neutral and negative, and the mental load degree is divided into three levels of high, medium and low.
[0013] The extracted preliminary features of each type are subjected to deep feature extraction by using a deep learning-based classification model to obtain the psychological state of the patient, specifically including:
[0014] Each type of extracted preliminary feature is taken as a modality, and a deep learning-based classification model is used for single-modality feature fusion and classification to obtain the preliminary psychological state classification result of the patient; wherein each modality uses a corresponding classification model, and in the training process of each classification model, a labeled data set is used for supervised learning, the label is evaluated based on the subjective emotional scale and the perceived stress scale, and through the scoring of the subjects, the score is divided into high, medium and low training labels according to the scale standard to classify emotions and mental load;
[0015] All single-modality preliminary psychological state classification results are subjected to multi-modality decision fusion based on rules, wherein the rules are to dynamically adjust the weight of each modality in decision fusion according to the individual classification performance of each modality, and the modality with better classification performance is given a higher weight. The classification labels of each modality are linearly weighted and summed to obtain a weighted result, and the weighted result is normalized to obtain the final psychological state classification result of the patient as the discrimination result of the patient's psychological state.
[0016] In some embodiments, the physical state of the patient is divided into peripheral fatigue degree and rehabilitation task performance;
[0017] The peripheral fatigue degree is used to intuitively reflect the muscle state of the patient, which is divided into high, medium and low degrees, and the data processing module obtains the peripheral fatigue information according to the electromyographic signal of the patient, specifically including: obtaining the peak electromyographic amplitude A of the current period after pre-processing the obtained electromyographic signal data, comparing the peak electromyographic amplitude A with the peak electromyographic amplitude A max within the first 10s of training, when A / A max is between 80%-100%, it is determined as a low fatigue state, when A / A max is between 60%-80%, it is determined as a moderate fatigue state, when A / A max is 60% or below, it is determined as a high fatigue state;
[0018] The rehabilitation task performance is related to the type and content of the upper limb rehabilitation training task, and the data processing module takes the motion accuracy and motion smoothness as evaluation indexes, the motion accuracy is represented by the consistency degree of the motion trajectory of the upper limb rehabilitation robot itself and the target trajectory path, and the motion smoothness is represented by the ratio of the average speed to the peak speed in the current training stage collected by the upper limb rehabilitation robot;
[0019] The data processing module integrates the physical state and the psychological state of the patient, and outputs in a unified format of natural language form:
[0020] "Psychological state--
[0021] Emotional state: positive / neutral / negative;
[0022] Mental load: high / medium / low;
[0023] Physical state--
[0024] Peripheral fatigue: high / medium / low;
[0025] Task performance: movement accuracy %a, movement smoothness %b"
[0026] Wherein, "%a" and "%b" are specific numerical values.
[0027] In some embodiments, the large language model used by the interaction strategy generation module is a first pre-trained large language model, and the pre-training process of the first large language model includes:
[0028] Step S100, thought chain training: first fine-tuning training is performed on the thought chain of the first large language model, and the first large language model is trained to gradually think and reason the decision-making ability of upper limb rehabilitation training content and the ability to communicate content and attitude through input of typical cases of exercise rehabilitation and exercise rehabilitation science and disease psychology theory, the upper limb rehabilitation training content includes training duration, training mode and task difficulty;
[0029] Step S200, interaction text training: collecting aging corpus and exercise training motivation corpus, the text content output by the first large language model after the first fine-tuning is further fine-tuned to make the output interaction text more suitable for the upper limb rehabilitation training scene and meet the patient's expectations.
[0030] In some embodiments, step S100 specifically includes:
[0031] Step S110, existing case collection:
[0032] Collecting existing decision-making cases and decision-making ideas of licensed professional rehabilitation physicians when training patients on upper limbs, and each existing case i should record at least: (1) patient condition s i , including injury time, lesion area and movement assessment; (2) initial artificial rehabilitation training scheme, which is equivalent to the initial rehabilitation training scheme p i adapted to the upper limb rehabilitation robot, and the initial rehabilitation training scheme p iThe standardized representation is a text sequence: "Training content: passive / assisted / active, resistance level, training speed, trajectory difficulty"; (3) the jth real-time adjustment strategy a of the training content and the communication content in the upper limb rehabilitation training process ij and the adjustment reason r ij , wherein the text standardized representation of the adjustment strategy a ij is: "Training content adjustment: passive / assisted / active, resistance level, training speed, trajectory difficulty; Communication content adjustment: communication attitude, communication text", the communication attitude has two labels: "comfort" or "encouragement"; the adjustment reason r ij includes changes in the patient's psychological state, changes in the patient's physical state, and / or professional theoretical knowledge referred to by the physician when making the adjustment;
[0033] Step S120, retrieve professional knowledge using the second large language model, and expand the real-time adjustment reason r of the training content and the communication content in the existing case through reasoning analysis ij :
[0034] By designing prompt words, prompting the second large language model to retrieve professional rehabilitation theory and disease psychology theory, and based on the zero-shot thinking chain training method, adding instructions in the prompt words that make the second large language model think step by step, so that the second large language model can combine the retrieved knowledge to analyze and reason step by step on the adjustment strategy a ij in the existing case i, to supplement and perfect the adjustment reason r ij corresponding to the single adjustment strategy a ij : ij The reasoning text character number is limited to not more than the first maximum length Lr1;
[0035] Step S130, data standardization:
[0036] All texts s i , p i , a ij , r ij of the case i are generated into text combinations w i according to the specified format, and the first text sequence W = {w1, w2, …, w i , …, w N} is generated using all case text combinations;
[0037] Step S140, first fine-tuning training:
[0038] The standardized first text sequence W is arranged into a first fine-tuning data set, and the first large language model is trained using an adapter fine-tuning method and based on the framework of Hugging Face Transformers; wherein a first cross-entropy loss function L LOSS1The adapter's parameter θ1 is optimized as the objective function:
[0039]
[0040] In the formula, |r i | represents the total number of training adjustment strategies - adjustment reason groups that case i possesses; P(p i |p i ,s i ;θ1) represents the first input of the patient's condition s into the language model after the initial training. i The initial rehabilitation training plan p is then obtained from the output. i The probability of P(a) is used to quantify the ability to derive an initial training plan from patient condition inferences; i,j |p i ,r i,≤j ,s i ,a i,<j ;θ1) is used to quantify the first large language model after the initial training. Based on the current patient condition, the initial rehabilitation training plan, the rehabilitation training adjustment strategy and reasons before the j-th adjustment, and the reasons for the j-th adjustment, the adjustment strategy a for the j-th adjustment is inferred. i,j The probability of r i,≤j Indicates the reason for the adjustment of rehabilitation training in the jth and previous sessions, a i,<j This represents the rehabilitation training adjustment strategy before the j-th adjustment.
[0041] In some embodiments, step S200 specifically includes:
[0042] Step S210, Corpus Collection:
[0043] We collected age-appropriate language data and motivational language data for rehabilitation exercise training from the internet using web crawling.
[0044] Step S220, Data Cleaning and Preprocessing:
[0045] The collected corpus was cleaned, including removing duplicate text, correcting punctuation errors and grammatical inconsistencies, and correcting typos and non-standard language. The cleaned corpus was then categorized according to semantic function, including motivational statements, comforting statements, task description statements, and everyday communication statements. The maximum number of characters in a single corpus fragment was limited to a second maximum length Lr2. All resulting corpus fragments were then formatted into a second text sequence V = {v1, v2, ..., v...}. l ,…,v M}, v l Represents a single segment of a corpus;
[0046] Step S230, Secondary Fine-tuning Training
[0047] The standardized second text sequence V is arranged into a second fine-tuning dataset, and the first large language model after the first fine-tuning is trained using an adapter fine-tuning method and based on the framework of Hugging Face Transformers; wherein the second cross-entropy loss function L LOSS2 The model parameters θ2 are fine-tuned as the objective function:
[0048]
[0049] wherein P(v l ; θ2) represents the probability of the first large language model after the second fine-tuning predicting the corpus segment v l .
[0050] In some embodiments, the pre-trained first large language model is caused to generate a standardized output by designing a prompt word, and the output content includes:
[0051] “-- training content--
[0052] Training duration: %c minutes;
[0053] Training mode: active / assistive / passive;
[0054] Training impedance: %dN;
[0055] Training trajectory: %e.
[0056] -- virtual human interaction content--
[0057] Voice interaction text: %f;
[0058] Interaction attitude: comfort / encouragement.
[0059] -- adjustment reason--
[0060] %g.”.
[0061] Wherein “%c” and “%d” are specific numerical values, “%e” is a specific mathematical formula, and “%f” and “%g” are specific text content.
[0062] In some embodiments, the voice interaction text and the interaction attitude generated by the interaction strategy generation module are converted into a voice with a tone feature and a positive emotional style familiar to the patient, and the patient is interacted with through the loudspeaker. The specific steps include:
[0063] The voice sample of the target speaker is collected, the text-to-speech model is used to extract the timbre features of the collected voice sample, including the fundamental frequency, the formant and the voice duration, and the feature vector of the target timbre is established; the text-to-speech model is used to realize the conversion from text to audio, and the extracted target timbre feature vector is embedded into the voice generated by the text-to-speech model, so that the audio output by the text-to-speech model has the timbre features of the target speaker through feature fusion; the generative adversarial network is used to extract the emotional features of the sample voice library with positive encouragement and comfort emotions; based on the extracted emotional features, the fundamental frequency value, the mel spectrum and the energy distribution of the audio output by the text-to-speech model are adjusted, so that the audio output by the loudspeaker has the timbre features familiar to the patient and the positive emotional style.
[0064] In some embodiments, the visual interaction device is selected from a display screen, a virtual reality device or an augmented reality device; the virtual character image generated by the visual interaction device includes three aspects of character appearance, facial movements corresponding to the voice of the character, and facial expressions of the character; wherein,
[0065] The character appearance is from a preset database or customized based on the preferences of the patient;
[0066] The facial movements corresponding to the voice of the character are generated by: extracting features from the voice signal output by the loudspeaker, capturing key prosody, syllable intensity and speech rate information in the voice; then dividing the extracted voice features into time steps to ensure the synchronization of audio and subsequent facial animation sequences on the time axis; encoding the audio signal into a low-dimensional voice embedding vector to ensure the correspondence with dynamic mouth shape and expression movements, and mapping the extracted voice features to the action parameter space of the FLAME model to control the key point movement of the mouth, the FLAME model cooperates with the reference expression parameters and voice features to generate a frame-by-frame dynamically changing facial movement grid; finally, the rendering engine is used to render the generated facial movement grid into facial movements adapted to the appearance of the virtual character;
[0067] The facial expressions of the character are determined by the interaction attitude output by the interaction strategy generation module, and based on the adjusted mouth shape movement of the extracted voice features, the expression layer tool of the FLAME model is used to superimpose the 50-dimensional expression control parameters of the virtual character on the basis of the mouth shape adjustment parameters, and the expression control parameters are obtained by fine-tuning the expression parameter library of the FLAME model.
[0068] In some embodiments, the upper limb rehabilitation robot adjusts its state in real time according to the training duration, training mode, training impedance and training trajectory in the training content generated by the interaction strategy generation module.
[0069] The present disclosure has the following characteristics and beneficial effects:
[0070] The upper limb movement rehabilitation system based on emotional interaction provided by the embodiment of the present disclosure introduces the positive effect of interpersonal emotional interaction in the process of human-computer interaction in movement rehabilitation training, activates the rehabilitation confidence and motivation of the user from the flexible adjustment of the task, establishes an image of a virtual doctor to provide visual feedback and social emotional interaction and support of voice for the patient, optimizes the function and effect of the traditional rehabilitation system, greatly improves the effectiveness of the rehabilitation training, helps the patient to improve the positive emotion and the compliance and participation of the training, and performs high-quality upper limb movement rehabilitation under the guidance of the positive emotion, shortens the rehabilitation cycle, and improves the life quality of the patient. BRIEF DESCRIPTION OF DRAWINGS
[0071] Figure 1 is a structural schematic diagram of the upper limb movement rehabilitation system based on emotional interaction provided by the embodiment of the present disclosure;
[0072] Figure 2 is Figure 1 is a schematic diagram of the psychological and physical state classification and rating of the upper limb movement rehabilitation system to the patient;
[0073] Figure 3 is Figure 1 is a specific process schematic diagram of the psychological and physical state analysis of the upper limb movement rehabilitation system to the patient;
[0074] Figure 4 is Figure 1 is a specific process schematic diagram of the generation of the multi-dimensional interaction strategy of the upper limb movement rehabilitation system;
[0075] Figure 5 is Figure 1 is a schematic diagram of the execution of the interaction process of the upper limb movement rehabilitation system. DETAILED DESCRIPTION
[0076] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.
[0077] On the contrary, the present application covers any substitution, modification, equivalent method and scheme defined by the claims within the essence and scope of the present application. Further, in order to make the public have a better understanding of the present application, some specific details are described in detail in the following detailed description of the present application. The present application can also be completely understood without the description of these details by those skilled in the art.
[0078] As Figure 1 shown, the upper limb movement rehabilitation system based on emotional interaction provided by the embodiment of the present disclosure comprises:
[0079] a sensing unit 1 for acquiring physiological signals and behavioral signals of the patient comprehensively to provide data support for rehabilitation training, the acquired physiological signals including but not limited to electroencephalogram, electromyogram, electrocardiogram and skin electricity signal, and the acquired behavioral signals including but not limited to facial expression, voice, eye movement and motor task performance information;
[0080] a central controller 2 including a data processing module 21 and an interactive strategy generation module 22; the data processing module 21 is configured to determine the physical state and psychological state of the patient in real time according to the physiological signals and behavioral signals acquired by the sensing unit 1, and output the determination result as the input of the interactive strategy generation module 22, and the interactive strategy generation module 22 simulates the professional knowledge and thinking mode of the rehabilitation physician to understand and fuse the current state of the patient and make reasoning decisions, and outputs a multi-dimensional interactive strategy for the patient, including real-time generated visual and voice interactive content and upper limb rehabilitation training content;
[0081] an execution unit 3 including a loudspeaker 31, a visual interactive device 32 and an upper limb rehabilitation robot 33; the loudspeaker 31 is configured to play a voice with a timbre feature familiar to the patient and with an encouraging or comforting emotional style according to the voice interactive content during the training; the visual interactive device 32 is configured to generate a virtual character according to the visual interactive content as one of the carriers simulating interpersonal emotional interaction and support, and generate corresponding facial movements and expressions of the virtual character in cooperation with the voice generated by the loudspeaker 31; and the upper limb rehabilitation robot 33 is configured to guide the patient to complete the upper limb rehabilitation training action according to the upper limb rehabilitation training content.
[0082] In some embodiments, the voice emitted by the loudspeaker 31 has the feature that the timbre approaches the familiar person or close person of the user, and has positive effects such as encouragement, comfort and care in emotion.
[0083] In some embodiments, the upper limb rehabilitation robot 33 can help the user to gradually reconstruct the fine motor function of the shoulder, elbow and wrist joints of the upper limb, and the basic structure thereof can be a two-degree-of-freedom mechanical arm, the end of which is connected to the palm of the patient and has the function of upper limb motor rehabilitation training, and can directly guide the patient to complete a two-dimensional plane from simple to complex movement trajectory, and has multiple training modes such as passive, assisted and active, and difficulty levels, and the change of the difficulty level can be realized by adjusting the complexity of the movement trajectory, the training speed in the passive and assisted mode training, and the resistance value in the active mode training.
[0084] Further, the movement trajectory guided by the upper limb rehabilitation robot 33 for the patient to complete can be straight line, circle and "8" shape trajectory from simple to difficult, which can be generated by giving key point coordinates and combining mathematical functions.
[0085] Further, the passive, assisted, and active training modes of the upper limb rehabilitation robot 33 can be implemented by any one of PID, fuzzy control, admittance control, and impedance control algorithms, and the training speed and training resistance value can be transformed by adjusting the parameters of the control algorithm. The range of the training speed can be 0-20 mm / s, and the range of the resistance value can be 0-10 N. In addition, the upper limb rehabilitation robot 33 also obtains the information of the joint force sensors and motor encoders in the internal mechanism, and feeds back the information such as the motion position, speed, and interaction force of the mechanism to the sensing unit 1, as the basic data for quantitatively evaluating the performance of the training task of the patient.
[0086] In some embodiments, the sensing unit 1 includes a physiological signal acquisition device and a behavior signal acquisition device. The physiological signal acquisition device includes an electroencephalogram acquisition instrument, an electrocardiogram acquisition instrument, an electromyogram acquisition instrument, and a galvanic skin response acquisition instrument; the behavior signal acquisition device includes an eye tracking instrument, a voice acquisition instrument, a facial expression capture instrument, and can also include force sensors and motor encoders integrated in the upper limb rehabilitation robot 33. Various devices in the sensing unit 1 are commercially available products.
[0087] In some embodiments, in order to provide the patient with an immersive rehabilitation experience, stimulate the patient's positive emotions and training motivation, and thus improve the rehabilitation effect, the current physiological condition and psychological condition of the patient need to be accurately distinguished before generating the multi-dimensional interaction strategy. Referring to Figure 2 , the psychological state of the patient can be further divided into emotional state and degree of mental load, and divided into three levels, namely positive, neutral, and negative emotions, and high, medium, and low mental load. In order to identify the emotional state and degree of mental load label of the patient, the physiological and behavior signals of the patient obtained by the sensing unit 1 are fully utilized, including but not limited to electroencephalogram, electrocardiogram, galvanic skin response, facial expression, voice, eye tracking, motion task performance, and the like.
[0088] In some embodiments, referring to Figure 3 , the process of the data processing module 21 for real-time distinguishing the psychological state of the patient according to the physiological and behavior signals obtained by the sensing unit 1 includes: respectively pre-processing and preliminarily extracting features of the obtained signals; then using a deep learning-based classification model to extract deep features from the preliminarily extracted features of various types, to obtain the psychological state of the patient.
[0089] Further, the specific process of the data processing module 21 for respectively pre-processing and preliminarily extracting features of the obtained signals includes:
[0090] For the electroencephalogram signal data, the data is denoised, and the eye movement and heartbeat artifacts are removed by independent principal component analysis (ICA). The power spectral density of alpha, beta, theta and delta bands in the electroencephalogram signal is extracted by Fourier transform or wavelet transform.
[0091] For the electrocardiogram signal data, baseline drift removal, denoising and R-wave detection preprocessing operations are performed, followed by extraction of heart rate variability, R-R interval and other time domain features, and extraction of low and high frequency components of the signal using Fourier transform.
[0092] For the skin conductance signal data, denoising and sliding average artifact removal preprocessing operations are first performed, and then the amplitude characteristics of the skin conductance response in the target time period are extracted.
[0093] For the facial expression signal data, denoising and normalization preprocessing operations are performed on each collected image, and the key point positions of the patient's facial features are obtained by combining the opencv facial feature detection method. Subsequently, the facial action unit features are extracted based on the Facial Action Coding System (FACS).
[0094] For the speech signal data, denoising preprocessing is first performed, and then the pitch frequency, tone features and mel-frequency cepstral coefficients are extracted based on Fourier transform. The speech signal is segmented into short frames, the number of speech frames per unit time is calculated to obtain the speech rate feature, and the energy of each speech signal frame is calculated to estimate the volume feature.
[0095] For the eye tracking signal data, background denoising preprocessing is first performed, and then the eye movement information is obtained. The position coordinates of the eye at each time are output, the trajectory sequence formed by the position coordinates in the time step is marked, and the Kalman filter method can be used for smoothing interpolation processing of the trajectory sequence to obtain the final eye movement trajectory feature.
[0096] Further, when the data processing module 21 extracts deep features from each type of preliminary feature, each type of preliminary feature is taken as a modality. First, a classification model based on deep learning is used for single-modality feature fusion and classification to obtain the preliminary psychological state label of the patient, i.e., the emotional state level (including three levels of positive, neutral and negative) and the mental load level (including three levels of high, medium and low). Subsequently, all single-modality classification results are decision-fused based on rules to obtain the final psychological state label of the patient. This process can be implemented by selecting a classification model based on convolutional neural network or other deep learning to fuse and classify each type of preliminary feature. To avoid the phenomenon of redundancy and model overfitting caused by too high dimension of multi-modal features, a hybrid multi-modal fusion strategy can be adopted to combine feature fusion and decision fusion for psychological state classification. The specific steps are as follows:
[0097] Single-modal feature fusion stage: n preliminary features f k1 ,f k2 ,…,f kn are fused and spliced to form a feature vector f k for the signal of each modality, and the classification models corresponding to different modalities can be homogeneous (i.e., the same model structure) or heterogeneous (different model structures). In an embodiment of the present application, a convolutional neural network (CNN), a recurrent neural network (RNN), and a long short-term memory (LSTM) are trained according to the characteristics of the modal features, respectively, to implement classification, and the classification labels y k of the single modality k are obtained. The classification labels correspond to the preliminary psychological state labels of the patient, i.e., the emotional state level (divided into positive, neutral, and negative) and the mental load level (divided into high, medium, and low). In the training process of each classification model, a labeled data set is used for supervised learning, and the labels are evaluated based on the subjective mood scale (such as PANAS) and the perceived stress scale. Through the subjective scoring of the subjects, the scores are divided into high, medium, and low training labels according to the scale standard to classify emotions and mental load; in the training process, a cross-entropy loss function can be used to optimize the classification accuracy of the classification model, and a K-fold cross-validation method is used to verify the classification effect of the pre-trained classification model.
[0098] Multi-modal decision fusion stage: based on the classification results of each modality obtained in the single-modal feature fusion stage, rules are formulated for decision fusion. Specifically, the classification effect of each modality can be evaluated based on the aforementioned K-fold cross-validation method, and indicators such as accuracy, precision, recall, and F1 score are used to dynamically adjust the weight r k of each modality in decision fusion according to the individual classification performance of each modality. The modality with better classification performance is given a higher weight. Then, the classification labels y k of each modality are linearly weighted and summed to obtain a weighted result Σr k y k . The weighted result is input into a fully connected layer with a softmax operation function to output the final psychological state classification label, which is used as the judgment result of the patient's psychological state.
[0099] In some embodiments, referring to Figure 2The data processing module 21 categorizes the patient's physical state into peripheral fatigue level and rehabilitation task performance based on the acquired physiological signals (mainly electromyographic signals) and behavioral signals. Peripheral fatigue level directly reflects the patient's muscle state and is also categorized into high, medium, and low levels. Peripheral fatigue information can be obtained using electromyographic signals. The specific steps are as follows: After preprocessing the acquired electromyographic signal data by removing baseline drift, denoising, and applying a moving average, the peak electromyographic amplitude A for the current time period is obtained. This peak electromyographic amplitude A is compared with the peak electromyographic amplitude A within the first 10 seconds of training. max When comparing, when A / A max When the fatigue level is between 80% and 100%, it is considered a low fatigue state. When A / A... max Between 60% and 80%, it is judged as a moderate fatigue state, when A / A max A state of high fatigue is defined as a fatigue level of 60% or below. Performance on rehabilitation tasks is directly related to the type and content of the task. For the upper limb rehabilitation robot of this embodiment, task performance may include motion accuracy (the degree of consistency between the motion trajectory and the target trajectory path) and motion smoothness (the patient's ability to control muscle exertion evenly and stably during rehabilitation exercises). The raw data for task performance, such as motion accuracy and motion smoothness, are provided by the sensor module built into the upper limb rehabilitation robot, and specific values can be calculated in each stage of training. For example, motion smoothness can be calculated using the following formula:
[0100]
[0101]
[0102] Among them, v average v represents the average speed during the current training phase. max v represents the peak speed during the current training phase. t Let t be the velocity at the t-th time step in the current training phase, and T be the total time step in the current training phase.
[0103] In some embodiments, the data processing module 21 ultimately integrates the determined physical and psychological states of the patient, and outputs the determination results in natural language form by writing Python code. The output natural language is written in the following unified format:
[0104] --Psychological state--
[0105] Emotional state: positive / neutral / negative;
[0106] Mental workload: High / Medium / Low;
[0107] --Physical condition--
[0108] Peripheral fatigue: high / medium / low
[0109] Task performance: movement accuracy %a, movement smoothness %b.
[0110] wherein both “%a” and “%b” are specific numerical values.
[0111] In some embodiments, referring to Figure 4 , the interactive strategy generation module 22 is centered on a generative large language model, and the output multi-dimensional interactive strategy includes voice and visual interactive content rich in positive emotions such as comfort and encouragement, as well as training content for the upper limb rehabilitation robot 33. The flexible interactive control mode of the multi-dimensional interactive strategy is significantly superior to the preset control program used in the past human-computer interaction, and has great potential in simulating human thinking decision-making process and real interpersonal emotional interaction. Among them:
[0112] The voice interactive content includes language text and voice audio generated based on the language text, and has the following features: the language text content is fully combined with the actual training performance of the patient, and can strengthen the patient's correct or good behavior through positive language, correct the patient's wrong or poor behavior, and the voice audio has the features of tone similar to the tone of the person close to the patient, and the emotion has the style of encouragement, encouragement, comfort. Further, in the voice interactive content, the tone feature of the person close to the patient is extracted before rehabilitation training, the encouragement and comfort emotion features are preset by the system, and the language text content and the final interactive audio with the tone and emotion features are generated by the generative large language model in real time during the upper limb rehabilitation training task.
[0113] The visual interactive content is a virtual rehabilitation doctor or other virtual character image, and has the following features: the virtual character image can communicate with the patient through facial expressions, gestures, and body postures, and establish a positive “interpersonal” relationship with the patient through voice audio, thereby stimulating the patient's positive emotions during upper limb rehabilitation training. In addition, the features of the virtual character image can be selected as preset images according to the patient's wishes, or can be customized according to the patient's needs.
[0114] The training content for the upper limb rehabilitation robot 33 can include training duration, training mode, task difficulty, etc., and has the following features: the training content should be reasonable, and should be maintained at a certain challenge level, but can give the patient a sense of success in achieving the goal, thereby stimulating the patient's training confidence and motivation.
[0115] Further, in order to improve the performance and adaptability of the generative large language model in the upper limb rehabilitation training scene, and make the output multi-dimensional interactive strategy more meet the actual needs of the patient group, the large language model needs to be pre-trained.
[0116] In some embodiments, the generative large language model adopts a first large language model, which is a lightweight large language model, such as a Llama3-8B large language model. Since the parameter order of the lightweight large language model is small, the computing power requirement of the pre-training process thereof is relatively easy to meet, and the real-time performance of the lightweight large language model in generating data is relatively high, which meets the real-time requirement of interaction with the upper limb rehabilitation training patient. The steps of pre-training the first large language model include:
[0117] Step S100, thought chain training, i.e., first fine-tuning training of the thought chain (CoT) of the first large language model. By inputting typical cases of motor rehabilitation and theories of motor rehabilitation and disease psychology, the first large language model is trained to gradually think and reason the content of upper limb rehabilitation training and the ability to communicate content and attitude;
[0118] Step S200, interactive text training. The interactive text content output by the first large language model is collected and secondarily fine-tuned by using suitable aging corpus and motor training motivation corpus, so that the output interactive text is more suitable for the upper limb rehabilitation training scene and meets the expectations of the patient.
[0119] Further, the specific steps of step S100, thought chain training, include:
[0120] Step S110, existing case collection.
[0121] The existing decision-making cases and decision-making ideas of the licensed professional rehabilitation physician when training the patient are collected. Optionally, one existing case i should at least record: (1) patient condition s i , such as injury time, lesion area, and motor assessment, (2) initial artificial rehabilitation training scheme, such as motor training method, training speed, training time, etc. The initial artificial motor rehabilitation scheme is approximately equivalent to the initial rehabilitation training scheme p i adapted to the upper limb rehabilitation robot 33 in the embodiment of the present disclosure, and p i is standardized as a text sequence: “training content: passive / assistance / active, resistance level, training speed, trajectory difficulty”, and the equivalent process can refer to existing experience in the rehabilitation field and suggestions of professional rehabilitation physicians; (3) the jth real-time adjustment strategy a ij and adjustment reason r ij of the training content and communication content in the upper limb rehabilitation training process, wherein the text of the adjustment strategy a ij is standardized as: “training content adjustment: passive / assistance / active, resistance level, training speed, trajectory difficulty; communication content adjustment: communication attitude, communication text”, and further, the communication attitude can be two labels of “comfort” or “encouragement”.ij The psychological state change and the physical state change of the patient can be included, and further, professional theoretical knowledge referenced by the physician when making the adjustment.
[0122] Step S120, retrieving professional knowledge by using the second large language model, expanding the real-time adjustment reason r of the training content and the communication content in the existing case through reasoning analysis ij .
[0123] By designing a prompt, the second large language model is prompted to search a wide range of professional rehabilitation theory and disease psychology theory, and based on the zero-shot thinking chain training method, instructions are added to the prompt to make the second large language model think step by step, so that the second large language model can combine the retrieved knowledge to analyze and reason the adjustment strategy a ij in the existing case i step by step. Based on the retrieved knowledge, the second large language model combines the original professional rehabilitation physician experience knowledge to analyze and reason the adjustment strategy a ij step by step to supplement and perfect the adjustment reason r ij . To improve the training efficiency, the adjustment reason r ij corresponding to the single adjustment strategy a ij is limited to no more than the first maximum length Lr1=256. Optionally, the second large language model used in the present embodiment is ChatGPT-4o, which can combine external knowledge and existing case information to efficiently reason and supplement, so as to expand the sample size required for offline pre-training of the first large language model.
[0124] Step S130, data standardization.
[0125] All texts s i , p i , a ij , r ij of the case i are generated into text combinations w i according to the specified format, and the first text sequence W={w1,w2,…, w i ,…, w N} is generated by using all case texts. Since the first text sequence can form a causal chain relationship, w i in the first text sequence W can be represented in the format (s i -->p i , j:r ij -->a ij ).
[0126] Step S140, first model fine-tuning training.
[0127] The standardized first text sequence W is arranged into a first fine-tuning data set, and a first fine-tuning training is performed on a fully open-source Llama3-8B large language model. Specifically, the first fine-tuning process adopts an adapter tuning method, that is, a lightweight adapter module is introduced in each layer of the pre-trained first large language model, the weight parameters of the adapter module are adjusted through training, and the original parameters of the first large language model are kept unchanged, so as to realize efficient fine-tuning of the first large language model. Adapter Tuning not only greatly reduces the amount of parameters and resource consumption required in the training process, but also has good task adaptability and scalability, and is suitable for multi-task or specific scene fine-tuning requirements. The fine-tuning process of the Llama3-8B large language model is implemented based on the framework of Hugging Face Transformers. HuggingFace provides convenient Trainer API and efficient fine-tuning tools, which can quickly integrate Adapter Tuning and optimize the training process. Further, the Accelerate library of Hugging Face can be combined to realize distributed training and automatic mixed precision training, improving the efficiency of fine-tuning and the performance of large language models. In a specific embodiment, the learning rate of the fine-tuning process is set to 3e-4, the number of training epochs is 20, and the first cross-entropy loss function L LOSS1 The parameters θ1 of the adapter module are optimized as the objective function, as shown below:
[0128]
[0129] In the formula, |r i | represents the total number of training adjustment strategy-adjustment reason groups that case i has, since different cases may have different numbers of training adjustment strategy-adjustment reasons, in order to ensure that the influence weight of the loss function on the thinking chain training of different cases is constant and not affected by the number of adjustment strategy-adjustment reasons, the second term of the formula is normalized by |r i i i i ; θ1) represents the probability of the initial rehabilitation training scheme p i output by the first large language model after training inputting the patient's condition s i , which quantifies the ability to infer the initial training scheme from the patient's condition; P(a i,j |p i , r i,≤j , s i , a i,<j ; θ1) a probability of the first large language model after training inferring the adjustment strategy a i,j of the jth adjustment based on the current patient condition, the initial rehabilitation training scheme, the rehabilitation training adjustment strategy before the jth adjustment, the adjustment reason before the jth adjustment, and the adjustment reason of the jth adjustment i,≤j , where a i,<j represents the rehabilitation training adjustment strategy before the jth adjustment.
[0130] Further, the specific steps of step S200, interactive text training, include:
[0131] Step S210, corpus collection.
[0132] The old-age adaptation corpus and the rehabilitation exercise training motivation corpus are collected from the network through a network crawler. Optionally, the old-age adaptation corpus involves language expressions commonly used by the elderly in daily life, health management related to the elderly, psychological intervention, social support, etc.; its sources include: online articles, forums and comments related to the health and psychological support of the elderly, case corpus of communication with the elderly in social support and accompanying services, corpus of the elderly group described in professional literature in rehabilitation medical institutions. The rehabilitation exercise training motivation corpus is mainly the motivation corpus for improving the confidence and motivation of patients in rehabilitation training; its sources include: the sorting of motivational sentences in professional rehabilitation training documents; the real dialogue records between doctors and patients in excellent rehabilitation cases; literature and corpus related to motivation in sports psychology research.
[0133] Step S220, data cleaning and preprocessing.
[0134] The collected corpus is cleaned, including removing duplicate texts, correcting punctuation errors and syntax confusion, and correcting errors and language irregularities in the corpus through automated tools. In some embodiments, the Chinese text correction can be implemented using the python library named pycorrector. After data cleaning is completed, the corpus is divided into different categories according to semantic functions, such as motivational sentences, comforting sentences, task description sentences, and daily communication sentences, so that the first large language model can output appropriate text to patients in the rehabilitation training scene. To ensure the uniformity and efficient processing of the corpus, the maximum number of characters of a single corpus segment is limited to Lr2=128 to avoid interference and influence of long corpus on the rehabilitation training process. Finally, the obtained corpus is arranged into a second text sequence V={v1, v2, …, v l ,…,v M}, where v l represents a single corpus segment, and the corpus segments do not have causal relationships.
[0135] Step S230, second model fine-tuning training.
[0136] The normalized second text sequence V is arranged into a second fine-tuning data set, and the fine-tuned Llama3-8B large language model obtained in step S140 is subjected to second fine-tuning training. For details of the fine-tuning training process, refer to step S140. The learning rate used in the second fine-tuning process can be 2e-5, epoch is set to 20, and the second cross-entropy loss function L LOSS2 The model parameter θ2 is fine-tuned as the target function of training, and the specific formula is as follows:
[0137]
[0138] In the formula, P(v l ; θ2) represents the probability of the first large language model after the second fine-tuning predicting the corpus segment v l .
[0139] The first large language model after the second fine-tuning has great performance improvement in the human-computer interaction strategy making in the rehabilitation training scene, can fully analyze the multi-dimensional psychological and physical conditions of the patient, and make interaction decisions in the thinking way of an excellent rehabilitation physician, including appropriately adjusting the content and difficulty level of the next stage of rehabilitation training tasks to stimulate the patient's confidence and motivation to actively participate in rehabilitation training, timely generating some comforting or encouraging texts to directly provide emotional support to the user, etc., by simulating the real doctor-patient and family-patient interpersonal interaction process in the process of human-computer interaction, fully integrating and utilizing the positive effect of interpersonal emotional interaction in motor learning, improving the psychological health level of the patient, and amplifying the benefits of rehabilitation training tasks on the patient's physical condition.
[0140] In some embodiments, the prompt is designed to enable the fine-tuned first large language model to generate standardized output after accepting standardized input, and the specific output content includes:
[0141] "Training content--
[0142] Training duration: %c minutes;
[0143] Training mode: active / assistance / passive;
[0144] Training impedance: %dN;
[0145] Training trajectory: %e.
[0146] "Virtual interpersonal interaction content--
[0147] Interaction text: %f;
[0148] Interaction attitude: comfort / encouragement.
[0149] --adjustment reason--
[0150] %g."
[0151] In the above text information, "%c" and "%d" are specific numerical values, "%e" is a specific mathematical formula (a function of the motion trajectory of the end of the upper limb rehabilitation robot), and "%f" and "%g" are specific text content.
[0152] In some embodiments, referring to Figure 5 , the implementation of the multi-dimensional interaction strategy is completed by the execution unit 3. Specifically, the multi-dimensional interaction strategy generated by the interaction strategy generation module 22 is embodied in two aspects: exercise training content and interpersonal interaction simulation content. Among them, the interpersonal interaction simulation is realized by the loudspeaker 31 and the visual interaction device 32, which are respectively responsible for training motivation voice generation and virtual character image generation, constituting auditory and visual interpersonal interaction simulation, which enhances the positive emotions of patients by providing them with the positive effects of interpersonal interaction, and further enhances their motivation and engagement.
[0153] In some embodiments, for the interaction text generated by the interaction strategy generation module 22, it can be further converted into audio and interacted with the patient through the loudspeaker 31, and the audio can be endowed with the timbre characteristics of the patient's close relatives and the emotional characteristics of comfort, encouragement and other positive effects. The specific steps can include: collecting the voice samples of the target speaker, extracting the timbre characteristics of the collected voice samples using a text-to-speech (TTS) model (such as a deep learning model VALL-E X), including fundamental frequency, formant and speech duration, etc. Key parameters to establish a feature vector of the target timbre. Then use the TTS model to realize the conversion of text to audio, and embed the target timbre characteristics extracted in the first step into the speech generated by the TTS model, and through feature fusion, the audio output by the TTS model has the timbre characteristics of the target speaker; then use the generative adversarial network to extract the emotional features of the existing sample voice library with positive encouragement and comfort emotions. The sample voice library is composed of a total of 400 audio data groups by recording 20 adults of different genders and timbres reading comfort and encouragement corpus.
[0154] In some embodiments, the visual interaction device 32 can be a traditional display screen, or an augmented reality or virtual reality device to provide a more immersive visual experience. The virtual character image generation based on the visual interaction device 32 includes three aspects: character appearance, character facial movements corresponding to the voice, and character facial expressions.
[0155] In some embodiments, the virtual character appearance can be from a preset database or customized based on the patient's preferences. The specific implementation is that the visual interaction device 32 creates a preset 3D role database through the Unity rendering engine, and the character image materials come from various open source or paid platforms MakeHuman, Adobe Mixamo, and Renderpeople. Further, the preset virtual character appearance parameters can be adjusted or remodeled according to the patient's needs to customize the desired image. Alternatively, the customization process can use a 3D scanning device (such as Artec Eva or Structure Sensor) to collect the real facial and body features of the target person, generate a high-precision 3D mesh model, and then use the Unity engine to realize virtual character model customization combined with the scanning data.
[0156] In some embodiments, the virtual character and voice-adapted facial movements mainly simulate the mouth shape when the character speaks. The specific implementation includes: using a pre-processing algorithm (such as MFCCs or Wav2Vec 2.0) to extract features from the speech signal of the training incentive voice generated by the loudspeaker 31, capturing key prosody, syllable intensity, and speech rate information in the voice; then dividing the extracted speech features into time steps to ensure the synchronization of audio and subsequent facial animation sequences on the time axis. Based on the VOCA (Voice Operated Character Animation) model, the audio signal is encoded into a low-dimensional speech embedding vector to ensure the corresponding relationship with dynamic mouth shape and expression movements, and then the extracted speech features are mapped to the expression parameter space of the FLAME (Faces Learned with an Articulated Model and Expressions) model to control the key point movements of the mouth (such as opening, closing, round lips, etc.). The FLAME model cooperates with the reference expression parameters and speech features to generate a frame-by-frame dynamically changing facial action mesh, and finally uses the Unity rendering engine to render the generated dynamic facial action mesh into the facial action of the virtual character appearance.
[0157] In some embodiments, the virtual character facial expression is determined according to the interaction attitude output by the interaction strategy generation module 22, mainly for comfort and encouragement. Based on the voice-driven mouth shape generation, the Expression Layer tool of the FLAME model is used to superimpose 50-dimensional expression control parameters of the virtual character on the basis of the mouth shape adjustment parameters, which can be obtained by fine-tuning the expression parameter library of the FLAME model.
[0158] In some embodiments, the rehabilitation training content generated by the interaction strategy generation module 22 is adjusted by the upper limb rehabilitation robot 33. The training content includes training duration, training mode, training impedance, and training trajectory. For training impedance, the upper limb rehabilitation robot 33 dynamically adjusts the output torque to change the muscle force required by the patient when performing the task to achieve the impedance control goal of different upper limb rehabilitation stages. For training trajectory, the upper limb rehabilitation robot 33 adjusts the motion range and trajectory shape of the mechanical arm according to the preset path function through the trajectory planning algorithm, gradually transitioning from simple straight line motion to complex trajectories such as curves. As previously described, the interaction strategy generation module 22 sets reasonable training content by considering the physical and psychological conditions of the patient, for example, when the user exhibits any of the conditions of “low emotional state”, “high mental load”, “high peripheral fatigue level”, and “significant decline in task performance value”, the training intensity and difficulty can be reduced, thereby improving the patient's sense of achievement and self-efficacy during upper limb rehabilitation training, and inspiring their rehabilitation confidence and motivation.
[0159] In summary, the upper limb movement rehabilitation system based on emotional interaction provided by the embodiments of the present disclosure discriminates the psychological and physical states of the patient by collecting multi-source physiological and behavioral signals of the patient, and adjusts the interaction strategy of the rehabilitation system in real time based on the large language model. The system introduces the positive effect of interpersonal emotional interaction into the traditional human-computer interaction process of movement rehabilitation training dominated by rehabilitation robots, from activating the rehabilitation confidence and motivation of the patient by personalizing the task content, to establishing a virtual character image to provide visual feedback and voice encouragement for the patient. Therefore, the embodiments of the present disclosure can optimize the function and effect of the traditional rehabilitation system, improve the effectiveness of rehabilitation training, help the patient improve positive emotions and their adherence and participation in training, and perform high-quality upper limb movement rehabilitation under the guidance of positive emotions, shorten the rehabilitation period, and improve the patient's quality of life.
[0160] In the description of the present specification, the description of the terms “one embodiment”, “some embodiments”, “illustrative embodiment”, “example”, “specific example”, or “some examples” means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In the present specification, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0161] Although embodiments of the disclosure have been shown and described, it will be apparent to those having ordinary skill in the art that a number of changes, modifications, alternatives, and variations can be made to the embodiments without departing from the principles and spirit of the disclosure, the scope of which is defined by the claims and their equivalents.
Claims
1. An upper limb motor rehabilitation system based on affective interaction, characterized in that, The application relates to a patient upper limb rehabilitation interactive system, which comprises the following: a sensing unit for acquiring physiological signals and behavioral signals of a patient; a central controller comprising a data processing module and an interactive strategy generation module; the data processing module is used for judging the physical state and psychological state of the patient in real time according to the physiological signals and behavioral signals acquired by the sensing unit and outputting the judgment result as the input of the interactive strategy generation module, the interactive strategy generation module simulates the professional knowledge and thinking mode of a rehabilitation physician to understand and fuse the current state of the patient and make a reasoning decision, and outputs a multi-dimensional interactive strategy of the patient, including real-time generated visual and voice interactive content and upper limb rehabilitation training content; an execution unit comprising a loudspeaker, a visual interactive device and an upper limb rehabilitation robot; the loudspeaker is used for playing voice with a tone feature familiar to the patient and a positive emotional style according to the voice interactive content generated by the interactive strategy generation module during training; the visual interactive device is used for generating a virtual character according to the visual interactive content generated by the interactive strategy generation module, so as to serve as one of carriers for simulating interpersonal emotional interaction and support, and the virtual character is caused to produce corresponding facial movements and expressions in cooperation with the voice generated by the loudspeaker; and the upper limb rehabilitation robot is used for guiding the patient to complete upper limb rehabilitation training actions according to the upper limb rehabilitation training content generated by the interactive strategy generation module; the large language model used by the interactive strategy generation module is a pre-trained first large language model, and the pre-training process of the first large language model comprises the following steps: step S100, thinking chain training: the first large language model is subjected to first fine-tuning training of a thinking chain, and the first large language model is trained to gradually think and reason the ability of upper limb rehabilitation training content and communication content and attitude through input of typical cases of movement rehabilitation and theories of movement rehabilitation science and disease psychology; the upper limb rehabilitation training content comprises training duration, training mode and task difficulty; specifically comprising: step S110, existing case collection: Collecting existing decision cases and decision ideas of licensed professional rehabilitation physicians when they conduct upper limb training for patients, one existing case i should record at least: (1) patient condition s i , including injury time, lesion area and motion assessment; (2) initial artificial rehabilitation training scheme, which is equivalent to the initial rehabilitation training scheme p i adapted to the upper limb rehabilitation robot, and the initial rehabilitation training scheme p i is normalized as a text sequence: "training content: passive / assistance / initiative, resistance level, training speed, trajectory difficulty"; (3) the jth real-time adjustment strategy a ij and the adjustment reason r ij of the training content and the communication content in the upper limb rehabilitation training process, wherein the text normalized representation of the adjustment strategy a ij is: "training content adjustment: passive / assistance / initiative, resistance level, training speed, trajectory difficulty; communication content adjustment: communication attitude, communication text", the communication attitude has two labels of "comfort" or "encouragement"; the adjustment reason r ij includes the change of the patient's psychological state, the change of the patient's physical state and / or the professional theoretical knowledge referred to by the physician when making the adjustment; Step S120, retrieving professional knowledge by using the second large language model, and expanding the real-time adjustment reason r of the training content and the communication content in the existing case through reasoning analysis ij : By designing a prompt word, prompting the second large language model to retrieve professional rehabilitation theory and disease psychology theory, and based on the zero-shot thinking chain training method, an instruction is added to the prompt word to make the second large language model think step by step, so that the second large language model can combine the retrieved knowledge to adjust the strategy a in the existing case i ij Carrying out step-by-step analysis reasoning to supplement and perfect the adjustment reason r ij , the single-time adjustment strategy a ij Corresponding adjustment reason r ij The number of characters in the reasoning text is limited to not more than the first maximum length Lr1; step S130, data standardization: all the texts s of the cases i i , p i , a ij , r ij generate text combinations w in a prescribed format i and generate a first text sequence W = {w1, w2,..., w i ,..., w N} using all the case text combinations step S140, first fine-tuning training: The standardized first text sequence W is arranged into a first fine-tuning data set, and a first large language model is trained using an adapter fine-tuning method and based on a framework of Hugging Face Transformers; wherein a first cross-entropy loss function L LOSS1 The parameters θ1 of the adapter are optimized as an objective function: wherein |r i | represents the total number of training adjustment strategies-adjustment reasons group that case i has; P(p i |p i ,s i ; θ1) represents the probability of the first large language model outputting the initial rehabilitation training scheme p i after the first training based on the patient's condition s i , which is used to quantify the ability to infer the initial training scheme from the patient's condition; P(a i,j |p i ,r i,≤j ,s i ,a i,<j ; θ1) is used to quantify the probability of the first large language model inferring the jth adjustment strategy a i,j based on the current patient's condition, the initial rehabilitation training scheme, the rehabilitation training adjustment strategy before the jth adjustment, and the jth adjustment reason after the first training; r i,≤j represents the rehabilitation training adjustment reason before the jth adjustment; and a i,<j represents the rehabilitation training adjustment strategy before the jth adjustment. step S200, interactive text training: the text content output by the first large language model after the first fine-tuning is subjected to second fine-tuning training of the text content by collecting aging adaptation corpus and movement training motivation corpus, so that the output interactive text is more suitable for the upper limb rehabilitation training scene and meets the expectations of the patient; specifically comprising: step S210, corpus collection: aging adaptation corpus and rehabilitation movement training motivation corpus are collected from the network through a network crawler; step S220, data cleaning and preprocessing: The collected corpus is cleaned, including removing duplicate texts, correcting punctuation errors and grammatical chaos, and correcting misspelled words and language irregularities in the corpus; the cleaned corpus is divided into different categories according to semantic functions, including encouraging statements, comforting statements, task description statements and daily communication statements; the maximum number of characters of a single corpus segment is limited to not more than the second maximum length Lr2; all corpus segments obtained are arranged into a second text sequence V={v1,v2,…,v l ,…,v M} according to the format, v l represents a single corpus segment; step S230, second fine-tuning training The standardized second text sequence V is arranged into a second fine-tuning data set, and the first large language model after the first fine-tuning is trained by using an adapter fine-tuning method and based on a framework of Hugging Face Transformers; wherein a second cross-entropy loss function L LOSS2 The model parameters θ2 are fine-tuned as a target function: where P(v l ; θ2) denotes the probability of the first large language model predicting the corpus segment v l after the second fine tuning.
2. The upper limb motion rehabilitation system of claim 1, wherein, The upper limb rehabilitation robot can guide the patient to complete a two-dimensional plane from simple to complex motion trajectory, and has passive, assisted and active training modes and difficulty levels. The change of the difficulty level is realized by adjusting the complexity of the motion trajectory, the training speed in the passive and assisted mode training, and the resistance value in the active mode training. The upper limb rehabilitation robot obtains the information of the internal joint force sensor and motor encoder, and feeds back the information including the motion position, speed and interaction force of the mechanism to the sensing unit as the basic data for quantitative evaluation of the performance of the patient's training task.
3. The upper limb motor rehabilitation system of claim 1, wherein, The physiological signals obtained by the sensing unit include electroencephalogram, electromyogram, electrocardiogram and skin electricity signals, and the behavior signals include facial expression, voice, eye movement and motion task performance information. The process of real-time judgment of the psychological state of the patient by the data processing module according to the physiological and behavior signals obtained by the sensing unit includes: pre-processing and preliminary feature extraction of the obtained signals; then deep feature extraction of the extracted preliminary features by using a classification model based on deep learning to obtain the psychological state of the patient.
4. The upper limb motor rehabilitation system of claim 3, wherein, The psychological state of the patient is divided into emotional state and mental load degree, the emotional state is divided into three levels of positive, neutral and negative, and the mental load degree is divided into three levels of high, medium and low. The deep feature extraction of the extracted preliminary features by using a classification model based on deep learning to obtain the psychological state of the patient specifically includes: Each type of preliminary feature is used as a modality, and a classification model based on deep learning is used for single-modality feature fusion and classification to obtain the preliminary psychological state classification result of the patient. Each modality uses a corresponding classification model, and in the training process of each classification model, a labeled data set is used for supervised learning. The label is evaluated based on the subjective emotion scale and the perceived stress scale. Through the scoring of the subjects, the score is divided into high, medium and low training labels according to the scale standard to classify emotions and mental load. All single-modality preliminary psychological state classification results are multi-modality decision fusion based on rules. The rule is to dynamically adjust the weight of each modality in decision fusion according to the classification performance of each modality. The better the classification performance of a modality, the higher the weight. The classification labels of each modality are linearly weighted and summed to obtain a weighted result. The weighted result is normalized to obtain the final psychological state classification result of the patient as the judgment result of the psychological state of the patient.
5. The upper limb motor rehabilitation system of claim 1, wherein, The physical state of the patient is divided into peripheral fatigue degree and rehabilitation task performance; The peripheral fatigue degree is used to intuitively reflect the muscle state of the patient, and is divided into three degrees of high, medium and low. The data processing module obtains peripheral fatigue information according to the electromyographic signal of the patient, and specifically includes: obtaining the peak electromyographic amplitude A of the current period after pre-processing the obtained electromyographic signal data, comparing the peak electromyographic amplitude A with the peak electromyographic amplitude A within the first 10s of training max , when A / A max is between 80%-100%, it is determined as a low fatigue state, when A / A max is between 60%-80%, it is determined as a medium fatigue state, and when A / A max is 60% or below, it is determined as a high fatigue state. The rehabilitation task performance is related to the type and content of the upper limb rehabilitation training task. The data processing module uses motion accuracy and motion smoothness as evaluation indexes. The motion accuracy is represented by the consistency of the motion trajectory of the mechanism of the upper limb rehabilitation robot and the target trajectory path. The motion smoothness is represented by the ratio of the average speed to the peak speed in the current training stage. The data processing module integrates the physical state and the psychological state of the patient, and outputs in a unified format of natural language form: "--Psychological state-- Emotional state: positive / neutral / negative; Mental load: high / medium / low; --Physical state-- Peripheral fatigue: high / medium / low; Task performance: motion accuracy %a, motion smoothness %b" Wherein, "%a" and "%b" are specific numerical values.
6. The upper limb motor rehabilitation system of claim 1, wherein, By designing prompt words, the first large language model pre-trained generates standardized output, and the output content includes: "--Training content-- Training duration: %c minutes; Training mode: active / assistance / passive; Training impedance: %dN; Training trajectory: %e; --Virtual interpersonal interaction content-- Voice interaction text: %f; Interaction attitude: comfort / encouragement; --Adjustment reason-- %g;” Wherein, "%c" and "%d" are specific numerical values, "%e" is a specific mathematical formula, and "%f" and "%g" are specific text content.
7. The upper limb motor rehabilitation system of claim 6, wherein, For the voice interaction text and interaction attitude generated by the interaction strategy generation module, it is converted into a voice with the patient's familiar timbre characteristics and positive emotional style, and interacts with the patient through the loudspeaker, the specific steps including: Collecting the voice samples of the target speaker, extracting the timbre characteristics of the collected voice samples using a text-to-speech model, including fundamental frequency, formant and voice duration, establishing a feature vector of the target timbre; using the text-to-speech model to realize the conversion of text to audio, and embedding the extracted target timbre feature vector into the voice generated by the text-to-speech model, through feature fusion, the audio output by the text-to-speech model has the timbre characteristics of the target speaker; using a generative adversarial network, extracting the emotional features of a sample voice library with positive encouragement and comfort emotions; based on the extracted emotional features, adjusting the fundamental frequency value, mel spectrum and energy distribution of the audio output by the text-to-speech model, so that the audio output by the loudspeaker has the timbre characteristics and positive emotional style familiar to the patient.
8. The upper limb motor rehabilitation system of claim 6, wherein, The visual interaction device selects a display screen, a virtual reality device or an augmented reality device; the virtual character image generated by the visual interaction device includes three aspects of character appearance, character facial movements corresponding to voice, and character facial expressions, wherein, The character appearance comes from a pre-set database or is customized based on the patient's preferences; The face movement of the character corresponding to the voice is generated by: feature extraction on the voice signal output by the speaker, capturing key prosody, syllable intensity and speech rate information in the voice; then dividing the extracted voice features into time steps, ensuring the synchronization of audio and subsequent face animation sequence on the time axis; encoding the audio signal into a low-dimensional speech embedding vector, ensuring the correspondence with dynamic mouth shape and expression movement, and mapping the extracted voice features to the action parameter space of the FLAME model, controlling the key point movement of the mouth, the FLAME model cooperates with the reference expression parameter and the voice feature to generate a frame-by-frame dynamically changing face movement grid; finally, the generated face movement grid is rendered into a face movement suitable for the appearance of the virtual character by using the rendering engine; The expression of the character is determined by the interaction attitude output by the interaction strategy generation module, and on the basis of the adjusted mouth shape movement obtained by extracting the voice feature, the expression layer tool of the FLAME model is used to superimpose the 50-dimensional expression control parameters of the virtual character on the basis of the mouth shape adjustment parameters, and the expression control parameters are obtained by fine-tuning the expression parameter library of the FLAME model.
9. The upper limb motor rehabilitation system of claim 6, wherein, The upper limb rehabilitation robot adjusts its state in real time according to the training duration, training mode, training impedance and training trajectory in the training content generated by the interaction strategy generation module.
Citation Information
Patent Citations
Virtual scene interaction-based rehabilitation training robot system and use method thereof
CN106779045A