Lower limb motion rehabilitation system based on emotional interaction

The lower limb motor rehabilitation system based on emotional interaction utilizes multi-sensor and deep learning to simulate interpersonal interaction, solving the problem of lack of emotional support in rehabilitation equipment, realizing an immersive rehabilitation experience, and improving rehabilitation efficiency and patient motivation.

CN119943274BActive Publication Date: 2026-04-14TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TSINGHUA UNIVERSITY
Filing Date
2025-01-22
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing rehabilitation equipment cannot replace the emotional support of rehabilitation physicians, resulting in low rehabilitation efficiency for patients and a tedious and monotonous rehabilitation process.

Method used

The lower limb motor rehabilitation system based on emotional interaction utilizes multi-sensor fusion and deep learning to simulate interpersonal interaction, combined with visual and auditory devices to provide an immersive rehabilitation experience, and simulates the emotional support of rehabilitation physicians through virtual characters and voice interaction.

Benefits of technology

It improved patients' rehabilitation outcomes, stimulated positive emotions and training motivation, shortened the rehabilitation cycle, and improved patients' quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943274B_ABST
    Figure CN119943274B_ABST
Patent Text Reader

Abstract

The disclosure provides a lower limb exercise rehabilitation system based on emotional interaction, comprising: a sensing unit for acquiring physiological signals and behavioral signals of a patient; in the central controller, the data processing module determines the physical state and psychological state of the patient in real time, the interaction strategy generation module simulates the professional knowledge and thinking mode of the rehabilitation physician to understand and integrate the current state of the patient, and outputs a multi-dimensional interaction strategy with the patient; in the execution unit, the loudspeaker is used to play the voice with the timbre characteristics and positive emotional style familiar to the patient according to the voice interaction content during training; the visual interaction device generates a virtual character according to the visual interaction content, and makes it produce corresponding facial movements and expressions with the voice generated by the loudspeaker; the lower limb rehabilitation robot guides the patient to complete the lower limb rehabilitation training action according to the lower limb rehabilitation training content. The disclosure provides emotional support for patient training, enhancing the rehabilitation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of elderly care and rehabilitation equipment and human-computer interaction technology, specifically relating to a lower limb motor rehabilitation system based on emotional interaction. Background Technology

[0002] With the deepening of global aging, the number of elderly people suffering from disability, dementia, and cognitive impairment is constantly increasing. Most of them also experience motor impairments, enduring both physical and psychological suffering. Medical theory and clinical medicine have proven that correct and scientific rehabilitation training plays a crucial role in the recovery and improvement of motor function. To address the shortage of professional caregivers and the high costs of medical care, safe, quantitative, effective, and repeatable rehabilitation training devices have emerged. However, the rehabilitation process using purely mechanical equipment is often tedious and monotonous, and existing rehabilitation equipment is far from replacing the emotional support provided by physicians during the patient's rehabilitation process. This negatively impacts the patient's emotional state and significantly affects their rehabilitation efficiency.

[0003] An excellent rehabilitation therapist can not only flexibly adjust training tasks according to the patient's current condition to stimulate their confidence and motivation for rehabilitation, but also provide significant emotional support through appropriate communication. A good doctor-patient relationship has been proven to play a crucial role in improving medical outcomes, including motor learning performance. Therefore, addressing the issue that existing rehabilitation devices cannot replace rehabilitation therapists, endowing robots with emotional interaction capabilities comparable to therapists, introducing the positive effects of interpersonal emotional communication, and creating an immersive and positive rehabilitation environment to enhance rehabilitation effects are important directions for the development of sports rehabilitation technology. However, no corresponding technology has yet been developed. Summary of the Invention

[0004] This disclosure aims to at least partially address one of the technical problems existing in the prior art.

[0005] To this end, this disclosure proposes a lower limb motor rehabilitation system based on emotional interaction. It utilizes multi-sensor fusion, deep learning, and generative large language models to simulate the positive emotional effects in interpersonal interaction. The system incorporates interpersonal emotional interaction functions and combines visual and auditory interaction devices with rehabilitation training equipment to provide patients with an immersive rehabilitation experience, stimulate their positive emotions and training motivation, thereby improving rehabilitation outcomes and promoting the recovery of lower limb motor function.

[0006] To achieve the above objectives, the present disclosure provides a lower limb motor rehabilitation system based on emotional interaction, comprising:

[0007] The sensing unit is used to acquire the patient's physiological and behavioral signals;

[0008] The central controller includes a data processing module and an interaction strategy generation module. The data processing module is used to determine the patient's physical and psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit, and outputs the determination results as input to the interaction strategy generation module. The interaction strategy generation module uses a large language model to simulate the professional knowledge and thinking of rehabilitation physicians to understand and integrate reasoning decisions about the patient's current state, and outputs multi-dimensional interaction strategies with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content.

[0009] The execution unit includes a speaker, a visual interaction device, and a lower limb rehabilitation robot. The speaker plays voices with familiar timbre and a positive emotional style based on the voice interaction content generated by the interaction strategy generation module during training. The visual interaction device generates a virtual character based on the visual interaction content generated by the interaction strategy generation module, serving as one of the carriers for simulating interpersonal emotional interaction and support. The virtual character then performs corresponding facial movements and expressions in conjunction with the voice generated by the speaker. The lower limb rehabilitation robot guides the patient to complete lower limb rehabilitation training movements based on the lower limb rehabilitation training content generated by the interaction strategy generation module.

[0010] In some embodiments, the lower limb rehabilitation robot can guide patients to complete a full cycle of cycling motion, and has multiple training modes and difficulty levels, including passive, assisted, and active modes. The difficulty level is changed by adjusting the training speed in passive and assisted modes and the resistance value in active mode. The lower limb rehabilitation robot obtains information from its internal pedal pressure sensor and motor encoder, and feeds back information, including its own movement speed and interaction force, to the sensing unit as basic data for quantitatively evaluating the patient's training task performance.

[0011] In some embodiments, the physiological signals acquired by the sensing unit include electroencephalogram (EEG), electromyogram (EMG), electrocardiogram (ECG), and electrodermal conductance (EDC) signals, and the behavioral signals acquired include facial expressions, speech, eye movements, and motor task performance information. The process by which the data processing module determines the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit includes: preprocessing and extracting preliminary features from the acquired signals; and then using a deep learning-based classification model to extract deep features from the extracted preliminary features to obtain the patient's psychological state.

[0012] In some embodiments, the patient's psychological state is divided into emotional state and mental load level, the emotional state is divided into three levels: positive, neutral and negative, and the mental load level is divided into three levels: high, medium and low.

[0013] The process involves using a deep learning-based classification model to extract deeper features from the various preliminary features to obtain the patient's psychological state, specifically including:

[0014] Each extracted preliminary feature is treated as a modality. A deep learning-based classification model is used to perform feature fusion and classification of the single modality to obtain the preliminary psychological state classification results of the patients. Each modality uses a corresponding classification model. During the training of each classification model, a labeled dataset is used for supervised learning. The labels are based on the subjective emotion scale and the perceived stress scale. Based on the scores given by the subjects, the scores are divided into high, medium and low training labels according to the scale standards to classify emotions and mental load.

[0015] Based on rules, the preliminary psychological state classification results of all single modalities are fused into a multimodal decision. The rules are as follows: the weight of each modality in the decision fusion is dynamically adjusted according to the individual classification performance of each modality. The modality with better classification performance is given a higher weight. The classification labels of each modality are linearly weighted and summed to obtain a weighted result. The weighted result is then normalized to obtain the final psychological state classification result of the patient, which serves as the judgment result of the patient's psychological state.

[0016] In some embodiments, the patient's physical condition is divided into peripheral fatigue level and rehabilitation task performance;

[0017] The peripheral fatigue level is used to visually reflect the patient's muscle state and is divided into three levels: high, medium, and low. The data processing module obtains peripheral fatigue information based on the patient's electromyography (EMG) signals, specifically including: preprocessing the acquired EMG signal data to obtain the peak EMG amplitude A for the current time period, and comparing this peak EMG amplitude A with the peak EMG amplitude A within the first 10 seconds of training. max When comparing, when A / A max When the fatigue level is between 80% and 100%, it is considered a low fatigue state. When A / A... max Between 60% and 80%, it is judged as a moderate fatigue state, when A / A max When the fatigue level is 60% or below, it is considered a state of high fatigue.

[0018] The performance of the rehabilitation task is related to the type and content of the lower limb rehabilitation training task. The data processing module uses movement speed and movement symmetry as evaluation indicators. The movement speed is characterized by the circumferential rotation speed of the lower limb rehabilitation robot's own mechanism, and the movement symmetry is characterized by the ratio of the average positive pressure applied by the left and right legs to the pedals during the current full cycle of pedaling, collected by the lower limb rehabilitation robot.

[0019] The data processing module integrates the assessed physical and psychological states of the patient and outputs them in a unified natural language format:

[0020] --Psychological state--

[0021] Emotional state: positive / neutral / negative;

[0022] Mental workload: High / Medium / Low;

[0023] --Physical condition--

[0024] Peripheral fatigue: High / Medium / Low;

[0025] Task performance: Motion speed %a, Motion symmetry %b

[0026] In this context, "%a" and "%b" are both specific numerical values.

[0027] In some embodiments, the large language model used by the interaction strategy generation module is a pre-trained first large language model, and the pre-training process of the first large language model includes:

[0028] Step S100, Mind Chain Training: The mind chain of the first major language model is fine-tuned for the first time. By inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, the first major language model is trained to gradually think and reason about the content of lower limb rehabilitation training, as well as the content and attitude of communication. The content of lower limb rehabilitation training includes training duration, training mode and task difficulty.

[0029] Step S200, Interactive Text Training: Collect age-appropriate language data and exercise training motivational language data to perform secondary fine-tuning training on the text content output by the first language model after the initial fine-tuning, so that the output interactive text is more in line with the lower limb rehabilitation training scenario and meets the patient's expectations.

[0030] In some embodiments, step S100 specifically includes:

[0031] Step S110: Collection of existing cases:

[0032] Collect existing decision-making cases and decision-making processes of licensed professional rehabilitation physicians when performing lower limb training on patients. Each existing case i should record at least: (1) the patient's condition s i (2) An initial artificial rehabilitation training program, which is equivalent to an initial rehabilitation training program p adapted to the lower limb rehabilitation robot. i and the initial rehabilitation training plan p iThe standardized representation is a text sequence: "Training content: passive / assisted / active, resistance level, training speed"; (3) The j-th real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and the reasons for the adjustment ij Among them, strategy a will be adjusted ij The text normalization representation is as follows: "Training content adjustment: passive / assisted / active, resistance level, training speed; Communication content adjustment: communication attitude, communication text", where the communication attitude has two labels: "comforting" or "encouraging"; the reason for adjustment is r. ij This includes changes in the patient's psychological state, changes in their physical state, and / or the professional theoretical knowledge that physicians refer to when making adjustments;

[0033] Step S120: Utilize the second language model to retrieve professional knowledge, and expand the existing cases through reasoning analysis to explain the real-time adjustments to training content and communication content. ij :

[0034] By designing prompts, the second language model is guided to retrieve professional rehabilitation theories and theories of disease psychology. Based on the zero-shot thought chain training method, instructions for the second language model to think step by step are added to the prompts, enabling the second language model to combine the retrieved knowledge with the adjustment strategy a in the existing case i. ij A step-by-step analysis and reasoning process was conducted to supplement and refine the reasons for the adjustment. ij , and single adjustment strategy a ij The corresponding reason for the adjustment r ij The number of characters in the reasoning text is limited to no more than the first maximum length Lr1;

[0035] Step S130, Data Standardization:

[0036] All texts in case i i p i a ij r ij Generate a text combination according to the specified format. i And use all the case texts to generate the first text sequence W = {w1, w2, ..., w i ,…,w N};

[0037] Step S140, Initial Fine-tuning Training:

[0038] The standardized first text sequence W was organized into the first fine-tuning dataset. An adapter fine-tuning method was used, and the first large language model was trained based on the Hugging Face Transformers framework; specifically, the first cross-entropy loss function L was employed. LOSS1The adapter's parameter θ1 is optimized as the objective function:

[0039]

[0040] In the formula, |r i | represents the total number of training adjustment strategies - adjustment reason groups that case i possesses; P(p i |p i ,s i ;θ1) represents the first input of the patient's condition s into the language model after the initial training. i The initial rehabilitation training plan p is then obtained from the output. i The probability of P(a) is used to quantify the ability to derive an initial training plan from patient condition inferences; i,j |p i ,r i,≤j ,s i ,a i,<j ;θ1) is used to quantify the first large language model after the initial training. Based on the current patient condition, the initial rehabilitation training plan, the rehabilitation training adjustment strategy and reasons before the j-th adjustment, and the reasons for the j-th adjustment, the adjustment strategy a for the j-th adjustment is inferred. i,j The probability of r i,≤j Indicates the reason for the adjustment of rehabilitation training in the jth and previous sessions, a i,<j This represents the rehabilitation training adjustment strategy before the j-th adjustment.

[0041] In some embodiments, step S200 specifically includes:

[0042] Step S210, Corpus Collection:

[0043] We collected age-appropriate language data and motivational language data for rehabilitation exercise training from the internet using web crawling.

[0044] Step S220, Data Cleaning and Preprocessing:

[0045] The collected corpus was cleaned, including removing duplicate text, correcting punctuation errors and grammatical inconsistencies, and correcting typos and non-standard language. The cleaned corpus was then categorized according to semantic function, including motivational statements, comforting statements, task description statements, and everyday communication statements. The maximum number of characters in a single corpus fragment was limited to a second maximum length Lr2. All resulting corpus fragments were then formatted into a second text sequence V = {v1, v2, ..., v...}. l ,…,v M}, v l Represents a single segment of a corpus;

[0046] Step S230, Secondary Fine-tuning Training

[0047] The standardized second text sequence V was organized into a second fine-tuning dataset. An adapter fine-tuning method was used, and the first large language model, after its initial fine-tuning, was trained based on the Hugging Face Transformers framework. The second cross-entropy loss function L was employed. LOSS2 Used as the objective function to fine-tune the model parameters θ2:

[0048]

[0049] In the formula, P(v l ;θ2) represents the first large language model prediction corpus segment v after secondary fine-tuning. l The probability of.

[0050] In some embodiments, prompt words are designed to enable a pre-trained first language model to generate standardized output, the output of which includes:

[0051] --Training Content--

[0052] Training duration: %c minutes;

[0053] Training modes: Active / Assisted / Passive;

[0054] Training impedance: %dN.

[0055] --Virtual Interpersonal Interaction Content--

[0056] Voice-interactive text: %e;

[0057] Interaction attitude: comforting / encouraging.

[0058] --Reason for Adjustment--

[0059] %f.

[0060] In this context, "%c" and "%d" represent specific numerical values, while "%e" and "%f" represent specific textual content.

[0061] In some embodiments, the voice interaction text and interaction attitude generated by the interaction strategy generation module are converted into voice with timbre characteristics familiar to the patient and a positive emotional style, and then interacted with the patient through the speaker. Specific steps include:

[0062] Speech samples from the target speaker are collected, and a text-to-speech model is used to extract timbre features from the collected speech samples, including fundamental frequency, formants, and speech duration, to establish a feature vector of the target timbre. The text-to-speech model is then used to convert text to audio, and the extracted target timbre feature vector is embedded into the speech generated by the text-to-speech model. Through feature fusion, the audio output by the text-to-speech model possesses the timbre features of the target speaker. Generative adversarial networks are used to extract emotional features from a sample speech library with positive motivational and comforting emotions. Based on the extracted emotional features, the fundamental frequency, Mel spectrum, and energy distribution of the audio output by the text-to-speech model are adjusted so that the audio output by the speaker has timbre features familiar to the patient and a positive emotional style.

[0063] In some embodiments, the visual interaction device is selected from a display screen, a virtual reality device, or an augmented reality device; the virtual character generated by the visual interaction device includes three aspects: the character's appearance, facial movements corresponding to voice, and facial expressions; wherein...

[0064] The appearance of the character is derived from a pre-set database or customized based on the patient's preferences;

[0065] The facial movements corresponding to the voice of the character are generated in the following way: Features are extracted from the voice signal output by the speaker to capture key prosody, syllable intensity, and speech rate information; then, the extracted voice features are divided into time steps to ensure synchronization between the audio and subsequent facial animation sequences on the time axis; the audio signal is encoded into a low-dimensional voice embedding vector to ensure correspondence with dynamic lip movements and facial expressions, and the extracted voice features are mapped to the motion parameter space of the FLAME model to control the movement of key points of the mouth; the FLAME model, in conjunction with reference to facial expression parameters and voice features, generates a dynamically changing facial motion mesh frame by frame; finally, the generated facial motion mesh is rendered into facial movements adapted to the appearance of the virtual character using a rendering engine.

[0066] The facial expressions of the character are determined by the interactive attitude output by the interaction strategy generation module. Based on the lip movement adjustment obtained from the extracted speech features, the expression layer tool of the FLAME model is used to overlay 50-dimensional expression control parameters of the virtual character on the lip adjustment parameters. The expression control parameters are obtained by fine-tuning the expression parameter library built into the FLAME model.

[0067] In some embodiments, the lower limb rehabilitation robot adjusts its own state in real time based on the training duration, training mode, and training impedance in the training content generated by the interaction strategy generation module.

[0068] This disclosure has the following characteristics and beneficial effects:

[0069] This disclosure provides a lower limb motor rehabilitation system based on emotional interaction. By introducing the positive effects of interpersonal emotional interaction into the human-computer interaction process of motor rehabilitation training, it optimizes the functions and effects of traditional rehabilitation systems. This is achieved by flexibly adjusting tasks to activate users' confidence and motivation in rehabilitation, and by establishing a virtual doctor image to provide patients with visual feedback and voice-based socio-emotional interaction and support. This greatly improves the effectiveness of rehabilitation training, helps patients enhance their positive emotions and their compliance and participation in training, and enables high-quality lower limb motor rehabilitation under the guidance of positive emotions, shortening the rehabilitation cycle and improving patients' quality of life. Attached Figure Description

[0070] Figure 1 This is a schematic diagram of the structure of a lower limb motor rehabilitation system based on emotional interaction provided in an embodiment of this disclosure;

[0071] Figure 2 yes Figure 1 The diagram shows a lower limb motor rehabilitation system that classifies and rates the patient's psychological and physical condition.

[0072] Figure 3 yes Figure 1 The diagram shows the specific process of how the lower limb motor rehabilitation system analyzes the patient's psychological and physical state.

[0073] Figure 4 yes Figure 1 The diagram illustrates the specific process by which the lower limb motor rehabilitation system generates multidimensional interactive strategies.

[0074] Figure 5 yes Figure 1 The diagram shows the interactive process of the lower limb motor rehabilitation system. Detailed Implementation

[0075] To make the objectives, technical solutions, and advantages of this application clearer, the application will be described in further detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely for explaining this application and are not intended to limit this application.

[0076] Conversely, this application covers any alternatives, modifications, equivalent methods, and solutions made within the spirit and scope of this application as defined by the claims. Furthermore, to provide the public with a better understanding of this application, certain specific details are described in detail below. However, those skilled in the art will fully understand this application even without these detailed descriptions.

[0077] like Figure 1 As shown, an embodiment of this disclosure provides a lower limb motor rehabilitation system based on emotional interaction, comprising:

[0078] Sensing unit 1 is used to acquire the patient's physiological and behavioral signals in all directions to provide data support for rehabilitation training. The acquired physiological signals include, but are not limited to, electroencephalogram (EEG), electromyogram (EMG), electrocardiogram (ECG), and electrodermal conductance signals. The acquired behavioral signals include, but are not limited to, facial expressions, speech, eye movements, and motor task performance information.

[0079] The central controller 2 includes a data processing module 21 and an interaction strategy generation module 22. The data processing module 21 is used to judge the patient's physical and psychological state in real time based on the physiological and behavioral signals obtained by the sensing unit 1, and output the judgment results as the input of the interaction strategy generation module 22. The interaction strategy generation module 22 uses a large language model to simulate the professional knowledge and thinking of rehabilitation physicians to understand and integrate reasoning decisions about the patient's current state, and outputs multi-dimensional interaction strategies with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content.

[0080] The execution unit 3 includes a speaker 31, a visual interaction device 32, and a lower limb rehabilitation robot 33. The speaker 31 is used to play voices with familiar timbre characteristics and encouraging or comforting emotional styles based on the voice interaction content during training. The visual interaction device 32 is used to generate virtual characters based on the visual interaction content, serving as one of the carriers for simulating interpersonal emotional interaction and support. The generated virtual characters, in conjunction with the voices generated by the speaker 31, produce corresponding facial movements and expressions. The lower limb rehabilitation robot 33 is used to guide patients to complete lower limb rehabilitation training movements based on the lower limb rehabilitation training content.

[0081] In some embodiments, the voice emitted by the speaker 31 is characterized by a timbre that closely resembles that of someone familiar or close to the user, and has positive effects such as encouragement, comfort, and care.

[0082] In some embodiments, the lower limb rehabilitation robot 33 can help patients gradually rebuild the motor function of the hip, knee and ankle joints and gait function of the lower limbs. Its basic structure can be a seated or recumbent treadmill, with its end connected to the patient's foot. It has the function of lower limb motor rehabilitation training and can directly guide the patient to complete a full cycle of treadmill exercise. It also has multiple training modes and difficulty levels, including passive, assisted and active modes. The difficulty level can be changed by adjusting the training speed during passive and assisted mode training and the resistance value during active mode training.

[0083] Furthermore, the passive, assisted, and active training modes of the lower limb rehabilitation robot 33 can be implemented using any of the algorithms such as PID, fuzzy control, admittance control, and impedance control. The training speed and resistance value can be varied by adjusting the parameters of the control algorithm. The training speed range can be 0-20 mm / s, and the resistance value range can be 0-10 N. In addition, the lower limb rehabilitation robot 33 also obtains information from its pedal pressure sensor and motor encoder, feeding back information such as its own mechanism movement speed and interaction force to the sensing unit 1, which serves as the basis for quantitatively evaluating the patient's training task performance.

[0084] In some embodiments, the sensing unit 1 includes physiological signal acquisition devices and behavioral signal acquisition devices. The physiological signal acquisition devices include an electroencephalogram (EEG) signal acquisition device, an electrocardiogram (ECG) signal acquisition device, an electromyogram (EMG) signal acquisition device, and a skin conductance signal acquisition device; the behavioral signal acquisition devices include an eye tracker, a voice acquisition device, and a facial expression capture device, and may also include force sensors and motor encoders integrated into the lower limb rehabilitation robot 33. All devices in the sensing unit 1 are commercially available products.

[0085] In some embodiments, to provide patients with an immersive rehabilitation experience, stimulate their positive emotions and training motivation, and thereby improve rehabilitation outcomes, it is necessary to accurately assess the patient's current physiological and psychological state before generating multidimensional interaction strategies. See also... Figure 2 The patient's psychological state can be further subdivided into emotional state and mental load level, and divided into three levels: positive, neutral, and negative emotions, and high, medium, and low mental load. To identify the patient's emotional state and mental load level, this embodiment of the disclosure fully utilizes the patient's physiological and behavioral signals acquired by the sensing unit 1, including but not limited to data such as electroencephalogram (EEG), electrocardiogram (ECG), electrodermal conductance, facial expressions, speech, eye tracking, and motor task performance.

[0086] In some embodiments, see Figure 3 The process by which the data processing module 21 determines the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit 1 includes: preprocessing and extracting preliminary features from the acquired signals; and then using a deep learning-based classification model to extract deep features from the extracted preliminary features to obtain the patient's psychological state.

[0087] Furthermore, the specific process of preprocessing and preliminary feature extraction of the acquired signals by the data processing module 21 includes:

[0088] For EEG signal data, noise is removed, and independent principal component analysis (ICA) is used to remove eye movement and heartbeat artifacts. Fourier transform or wavelet transform is used to extract frequency domain features such as power spectral density of the alpha, beta, theta and delta bands in the EEG signal.

[0089] For ECG signal data, preprocessing operations such as baseline drift removal, noise reduction, and R wave detection are performed. Then, time-domain features such as heart rate variability and RR interval are extracted, and low-frequency and high-frequency components of the signal are extracted using Fourier transform.

[0090] For the skin conductance signal data, we first perform denoising and moving average artifact removal preprocessing operations, and then extract the amplitude features of the skin conductance response within the target time period.

[0091] For facial expression signal data, each frame of the acquired image is preprocessed by denoising and normalization, and the key points of the patient's facial features are obtained by combining the OpenCV facial feature detection method. Then, facial action unit features are extracted based on the Facial Action Coding System (FACS).

[0092] For speech signal data, noise reduction preprocessing is first performed. Then, based on Fourier transform, the fundamental frequency, pitch features and Mel frequency cepstral coefficients are extracted to perform short-time frame segmentation of the speech signal, calculate the number of speech frames per unit time, obtain speech rate features, calculate the energy of each frame of speech signal to estimate its volume features, etc.

[0093] For eye-tracking signal data, background denoising preprocessing is first performed, then eye movement information is acquired, the position coordinates of the eye at each moment are output, the trajectory sequence formed by the position coordinates in the time step is marked, and the Kalman filter method can be used to smooth the trajectory sequence to obtain the final eye-tracking trajectory features.

[0094] Furthermore, when the data processing module 21 performs deep feature extraction on various preliminary features, it treats each type of preliminary feature as a modality. First, it uses a deep learning-based classification model to perform feature fusion and classification of the single modality to obtain the patient's preliminary psychological state label, namely, emotional state level (including positive, neutral, and negative levels) and mental burden level (including high, medium, and low levels). Then, based on rules, it performs decision fusion on all single-modality classification results to obtain the patient's final psychological state label. This process can be achieved by using a classification model based on deep learning such as convolutional neural networks to fuse and classify various preliminary features. To avoid redundancy and model overfitting due to excessively high multimodal feature dimensionality, a hybrid multimodal fusion strategy can be adopted, combining feature fusion and decision fusion for psychological state classification. The specific steps are as follows:

[0095] Single-mode feature fusion stage: n preliminary features f of the k-mode signal k1 ,f k2 ,…,f kn The features are fused and concatenated to form the feature vector f corresponding to mode k. k Independent classification models are trained for each modality of signal. The classification models corresponding to different modalities can be homogeneous (i.e., have the same model structure) or heterogeneous (have different model structures). In one embodiment of this application, convolutional neural networks (CNN), recurrent neural networks (RNN), and long short-term memory (LSTM) are trained according to the characteristics of the modality features to achieve classification and obtain the classification label y of a single modality k. k The classification labels correspond to the patient's initial psychological state labels, namely, the emotional state level (divided into positive, neutral, and negative) and the mental burden level (divided into high, medium, and low). During the training of each classification model, supervised learning is performed using labeled datasets. The labels are based on subjective emotion scales (such as PANAS) and perceived stress scales. Based on the subjects' subjective scores, the scores are divided into high, medium, and low training labels according to the scale standards to classify emotions and mental burden. During training, minimizing the cross-entropy loss function can be used to optimize the classification accuracy of the classification model, and the K-fold cross-classification method is used to verify the classification effect of the pre-trained classification model.

[0096] Multimodal decision fusion stage: Based on the classification results of each modality obtained in the single-modal feature fusion stage, rules are formulated for decision fusion. Specifically, the classification effect on each modality can be evaluated based on the aforementioned K-fold cross-validation method, using metrics such as accuracy, precision, recall, and F1 score. The weight r of each modality in the decision fusion is dynamically adjusted according to its individual classification performance. k Modalities with better classification performance are assigned higher weights. Then, the classification labels y for each modality are assigned... k Perform a linear weighted summation to obtain the weighted result Σr k y k The weighted result is passed through a fully connected layer with softmax as the operation function to obtain the final psychological state classification label, which is used as the discrimination result of the patient's psychological state.

[0097] In some embodiments, see Figure 2The data processing module 21 categorizes the patient's physical state into peripheral fatigue level and rehabilitation task performance based on the acquired physiological signals (mainly electromyographic signals) and behavioral signals. Peripheral fatigue level directly reflects the patient's muscle state and is also categorized into high, medium, and low levels. Peripheral fatigue information can be obtained using electromyographic signals. The specific steps are as follows: After preprocessing the acquired electromyographic signal data by removing baseline drift, denoising, and applying a moving average, the peak electromyographic amplitude A for the current time period is obtained. This peak electromyographic amplitude A is compared with the peak electromyographic amplitude A within the first 10 seconds of training. max When comparing, when A / A max When the fatigue level is between 80% and 100%, it is considered a low fatigue state. When A / A... max Between 60% and 80%, it is judged as a moderate fatigue state, when A / A max A state of high fatigue is defined as a fatigue level of 60% or below. Performance on rehabilitation tasks is directly related to the type and content of the task. For the lower limb rehabilitation robot of this embodiment, task performance may include movement speed, which reflects the continuity and proficiency of the patient's lower limb movements, and movement symmetry, i.e., the patient's ability to control the muscles of both legs to exert force evenly and stably during rehabilitation exercises. The raw data for task performance, such as movement speed and movement symmetry, are provided by the sensor module built into the lower limb rehabilitation robot, and definite values ​​can be calculated in each stage of training. For example, movement speed can be characterized by the circumferential rotational speed of the mechanism, while movement symmetry can be calculated using the following formula:

[0098]

[0099] Among them, F left F represents the average pressure exerted on the pedal by the patient's left leg after a full cycle of cycling. right This represents the average pressure applied to the pedal by the patient's right leg after a full cycle of cycling.

[0100] In some embodiments, the data processing module 21 ultimately integrates the determined physical and psychological states of the patient, and outputs the determination results in natural language form by writing Python code. The output natural language is written in the following unified format:

[0101] --Psychological state--

[0102] Emotional state: positive / neutral / negative;

[0103] Mental workload: High / Medium / Low;

[0104] --Physical condition--

[0105] Peripheral fatigue: High / Medium / Low;

[0106] Task performance: Motion speed %a, Motion symmetry %b.

[0107] In this context, "%a" and "%b" are both specific numerical values.

[0108] In some embodiments, see Figure 4 The interaction strategy generation module 22, based on a generative large language model, outputs multi-dimensional interaction strategies, including voice and visual interaction content rich in positive emotions such as comfort and encouragement, as well as training content for the lower limb rehabilitation robot 33. The flexible interaction control method of the multi-dimensional interaction strategy is significantly superior to the preset control programs used in previous human-computer interactions, demonstrating great potential in simulating human thought and decision-making processes and real interpersonal emotional interactions. Among them:

[0109] The voice interaction content includes spoken text and audio generated from the spoken text. Its characteristics include: the spoken text content fully integrates with the patient's actual training performance, reinforcing correct or well-performing behaviors and correcting incorrect or poorly performed behaviors through positive language; and the audio features a timbre similar to that of someone close to the patient, with an encouraging, motivating, and comforting tone. Furthermore, the timbre characteristics of someone close to the patient are extracted before rehabilitation training, and the encouraging and comforting emotional characteristics are preset by the system. Both the spoken text content and the final interactive audio, incorporating timbre and emotional features, are generated in real-time by a generative large language model during the lower limb rehabilitation training task.

[0110] The visual interaction content consists of virtual rehabilitation physicians or other virtual avatars. These avatars are characterized by their ability to communicate with patients through facial expressions, gestures, and body postures, and to establish a positive interpersonal relationship with patients through voice and audio, thereby stimulating positive emotions in patients during lower limb rehabilitation training. Furthermore, the virtual avatars can be preset according to the patient's wishes or customized to meet their specific needs.

[0111] The training content for the lower limb rehabilitation robot 33 may include information such as training duration, training mode, and task difficulty. Its characteristics include: the training content should be set reasonably so that it is challenging enough to give patients a sense of success in achieving their goals, thereby stimulating their training confidence and motivation.

[0112] Furthermore, in order to improve the performance and adaptability of generative large language models in lower limb rehabilitation training scenarios, and to make their output multidimensional interaction strategies better meet the actual needs of the patient population, it is necessary to pre-train the large language model.

[0113] In some embodiments, the generative large language model employs a first large language model, which is a lightweight large language model, such as the Llama3-8B large language model. Because the lightweight large language model has a smaller number of parameters, the computational requirements for its pre-training process are more easily met. Furthermore, the lightweight large language model generates data with high real-time performance, meeting the real-time requirements for interaction with patients undergoing lower limb rehabilitation training. The steps for pre-training the first large language model include:

[0114] Step S100, Chain of Thought Training, is the first fine-tuning training of the chain of thought (CoT) of the first major language model. By inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, the first major language model is trained to gradually think and reason to make decisions on the content of lower limb rehabilitation training, as well as the content and attitude of communication.

[0115] Step S200: Interactive text training. Collect age-appropriate language data and exercise training motivational language data to perform secondary fine-tuning training on the interactive text content output by the first language model, so that the output interactive text is more in line with the lower limb rehabilitation training scenario and meets the patient's expectations.

[0116] Furthermore, the specific steps of step S100, the mind chain training, include:

[0117] Step S110: Collection of existing cases.

[0118] Collect existing decision-making cases and decision-making processes of licensed professional rehabilitation physicians when performing lower limb training on patients. Optionally, each existing case i should record at least: (1) the patient's condition s i (1) Information such as injury time, lesion area, and motor assessment status; (2) Initial artificial rehabilitation training program, such as information on exercise training methods, training speed, and training duration. The initial artificial exercise rehabilitation program should be approximately equivalent to the initial rehabilitation training program p adapted to the lower limb rehabilitation robot 33 in this embodiment. i And combining the characteristics of the lower limb rehabilitation robot 33, p i The standardized representation is a text sequence: “Training content: passive / assisted / active, resistance level, training speed”. The equivalent process can refer to existing experience in the rehabilitation field and the advice of professional rehabilitation physicians; (3) The j-th real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and the reasons for the adjustment ij Among them, strategy a will be adjusted ij The text normalization representation is as follows: "Training content adjustment: passive / assisted / active, resistance level, training speed; Communication content adjustment: communication attitude, communication text," where the communication attitude can be labeled as either "comforting" or "encouraging"; the reason for adjustment is r. ijThis may include changes in the patient's psychological state and physical state, and further include the professional theoretical knowledge that the physician refers to when making adjustments.

[0119] Step S120: Utilize the second language model to retrieve professional knowledge, and expand the existing cases through reasoning analysis to explain the real-time adjustments to training content and communication content. ij .

[0120] By designing prompts, the second language model is guided to conduct a broad search of professional rehabilitation theories and disease psychology theories. Based on the zero-shot thought chain training method, instructions are added to the prompts to guide the second language model to think step by step, enabling it to combine the retrieved knowledge with the adjustment strategy a in the existing case i. ij A step-by-step analysis and reasoning process is then conducted. The second language model, based on retrieved knowledge and combined with the existing professional rehabilitation physician experience, analyzes adjustment strategy a. ij A step-by-step analysis and reasoning process was conducted to supplement and refine the reasons for the adjustment. ij To improve training efficiency, compared with the single-adjustment strategy a... ij The corresponding reason for the adjustment r ij The number of characters in the reasoning text is limited to no more than the first maximum length Lr1 = 256. Optionally, the second large language model used in this embodiment is ChatGPT-4o, which fully combines external knowledge and existing case information for efficient reasoning and supplementation, thereby expanding the sample size required for offline pre-training of the first large language model.

[0121] Step S130: Data standardization.

[0122] All text in case i i p i a ij r ij Generate a text combination according to the specified format. i And use all the case texts to generate the first text sequence W = {w1, w2, ..., w i ,…,w N Since the first text sequence can form a causal chain relationship, w in the first text sequence W... i It can be represented as format (s) i -->p i ,j:r ij -->a ij ).

[0123] Step S140: Initial model fine-tuning training.

[0124] The standardized first text sequence W was organized into the first fine-tuning dataset for the initial fine-tuning training of the fully open-source Llama3-8B large language model. Specifically, the initial fine-tuning process employed the adapter tuning method, which introduces lightweight adapter modules into each layer of the pre-trained large language model. The weight parameters of these adapter modules are adjusted during training while maintaining the original parameters of the large language model, thus achieving efficient fine-tuning. Adapter tuning significantly reduces the number of parameters and resource consumption required during training and also possesses good task adaptability and scalability, making it suitable for fine-tuning needs in multi-task or specific scenarios. The fine-tuning process of the Llama3-8B large language model was implemented using the Hugging Face Transformers framework. Hugging Face provides a convenient Trainer API and efficient fine-tuning tools, enabling rapid integration of adapter tuning and optimization of the training process. Furthermore, it can be combined with Hugging Face's Accelerate library to achieve distributed training and automatic mixed-precision training, improving fine-tuning efficiency and the performance of the large language model. In one specific embodiment, the fine-tuning process sets the learning rate to 3e-4, the number of training epochs to 20, and uses the first cross-entropy loss function L. LOSS1 The parameter θ1 of the adapter module is optimized as the objective function, as shown below:

[0125]

[0126] In the formula, |r i | represents the total number of training adjustment strategies-adjustment reasons groups in case i. Since different cases may have different numbers of training adjustment strategies-adjustment reasons, to ensure that the weight of the influence on the loss function during mind chain training remains constant and is not affected by the number of adjustment strategies-adjustment reasons, |r is used before the second term of the formula. i Perform a normalization; P(p i |p i ,s i ;θ1) represents the first large language model input after training, which is the patient's condition s i The initial rehabilitation training plan p is then obtained from the output. i The probability of P(a) is used to quantify the ability to derive an initial training plan from patient condition inferences; i,j |p i ,r i,≤j ,s i ,a i,<j;θ1) The first large language model after quantitative training is used to infer the adjustment strategy a for the jth time based on the current patient condition, the initial rehabilitation training plan, the rehabilitation training adjustment strategy and reasons before the jth adjustment, and the reasons for the jth adjustment. i,j The probability of r i,≤j Indicates the reason for the adjustment of rehabilitation training in the jth and previous sessions, a i,<j This represents the rehabilitation training adjustment strategy before the j-th adjustment.

[0127] Furthermore, step S200, the specific steps of interactive text training, include:

[0128] Step S210: Corpus collection.

[0129] The aging-friendly language corpus and rehabilitation exercise training motivational language corpus were collected from the internet using web crawling. Optionally, the aging-friendly language corpus includes commonly used language expressions of the elderly in daily life, health management related to the elderly, psychological intervention, and social support; its sources include: online articles, forums, and comments related to elderly health and psychological support; case studies of communication with the elderly in social support and companionship services; and language describing the elderly from rehabilitation medical institutions and professional literature. The rehabilitation exercise training motivational language corpus mainly consists of motivational language to enhance patients' confidence and motivation in rehabilitation training; its sources include: compilations of motivational statements from professional rehabilitation training documents; real dialogue records between physicians and patients in excellent rehabilitation cases; and literature and corpus related to motivation in sports psychology research.

[0130] Step S220: Data cleaning and preprocessing.

[0131] The collected corpus undergoes cleaning, including removing duplicate text, correcting punctuation errors and grammatical inconsistencies, and using automated tools to correct typos and grammatical irregularities. In some embodiments, a Python library called pycorrector can be used for Chinese text correction. After data cleaning, the corpus is categorized according to semantic function, such as motivational statements, comforting statements, task description statements, and everyday communication statements, so that the primary language model can output appropriate text to patients in rehabilitation training scenarios. To ensure the uniformity and efficient processing of the corpus, the maximum number of characters in a single corpus segment is limited to Lr2 = 128 to avoid interference and impact on the rehabilitation training process from excessively long corpora. Finally, the obtained corpus is formatted into a second text sequence V = {v1, v2, ..., v...}. l ,…,v M}, where v l This represents a single corpus segment, where there is no causal relationship between the corpus segments.

[0132] Step S230: Secondary model fine-tuning training.

[0133] The standardized second text sequence V is organized into a second fine-tuning dataset. The fine-tuned Llama3-8B large language model obtained in step S140 is then subjected to secondary fine-tuning training. The specific fine-tuning training process is detailed in step S140. The learning rate used in the secondary fine-tuning process can be 2e-5, the epoch is set to 20, and a second cross-entropy loss function L is introduced. LOSS2 The model parameters θ2 are fine-tuned using the objective function for training, as shown below:

[0134]

[0135] In the formula, P(v l ;θ2) represents the first large language model prediction corpus segment v after secondary fine-tuning. l The probability of.

[0136] The first major language model, after secondary fine-tuning, shows significant performance improvement in human-computer interaction strategy formulation in rehabilitation training scenarios. It can fully analyze the multidimensional psychological and physical conditions of patients and make interactive decisions in the mindset of an excellent rehabilitation physician. This includes appropriately adjusting the content and difficulty level of the next stage of rehabilitation training tasks to stimulate patients' confidence and motivation to actively participate in rehabilitation training, and generating comforting or encouraging texts in a timely manner to directly provide emotional support to users. By simulating the real doctor-patient and family-patient interpersonal interaction process during human-computer interaction, it fully integrates and utilizes the positive effects of interpersonal emotional interaction in motor learning, improves patients' mental health, and amplifies the benefits of rehabilitation training tasks for patients' physical condition.

[0137] In some embodiments, the prompt is designed so that the fine-tuned first language model, after receiving standardized input, can generate standardized output, specifically including:

[0138] --Training Content--

[0139] Training duration: %c minutes;

[0140] Training modes: Active / Assisted / Passive;

[0141] Training impedance: %dN.

[0142] --Virtual Interpersonal Interaction Content--

[0143] Interactive text: %e;

[0144] Interaction attitude: comforting / encouraging.

[0145] --Reason for Adjustment--

[0146] %f.

[0147] In the above text information, "%c" and "%d" are specific numerical values, and "%e" and "%f" are specific text content.

[0148] In some embodiments, see Figure 5 The multi-dimensional interaction strategy is implemented by the execution unit 3. Specifically, the multi-dimensional interaction strategy generated by the interaction strategy generation module 22 is reflected in two aspects: exercise training content and interpersonal interaction simulation content. Among them, the interpersonal interaction simulation is realized by the speaker 31 and the visual interaction device 32, which are responsible for generating training incentive speech and virtual character image respectively, constituting auditory and visual interpersonal interaction simulation. By providing patients with the positive effects of interpersonal interaction, it enhances their positive emotions for rehabilitation, thereby enhancing their motivation and level of engagement.

[0149] In some embodiments, the interactive text generated by the interaction strategy generation module 22 can be further converted into audio and used to interact with the patient through the speaker 31. The audio can be endowed with the timbre characteristics of a person close to the patient, as well as emotional features with positive effects such as comfort and encouragement. Specific steps may include: collecting speech samples from the target speaker; using a text-to-speech (TTS) model (such as the deep learning model VALL-EX) to extract timbre features from the collected speech samples, including key parameters such as fundamental frequency, formants, and speech duration, to establish a feature vector of the target timbre; subsequently, using the TTS model to convert text to audio, and embedding the target timbre features extracted in the first step into the speech generated by the TTS model; through feature fusion, the audio output by the TTS model possesses the timbre characteristics of the target speaker; and then using a generative adversarial network to extract emotional features from an existing sample speech library with positive encouragement and comfort. The sample speech library consists of 400 audio data points generated by recording audio of 20 adults of different genders and timbres reading comforting and encouraging text. Based on the extracted emotional features, the fundamental frequency, Mel spectrum, and energy distribution of the audio output by the TTS model are further adjusted to make the audio output by speaker 31 present a gentle and positive tone, conveying feelings of comfort and encouragement.

[0150] In some embodiments, the visual interaction device 32 can be a traditional display screen, or an augmented reality or virtual reality device to provide a more immersive visual experience. The generation of virtual characters based on the visual interaction device 32 includes three aspects: character appearance, facial movements corresponding to voice, and facial expressions.

[0151] In some embodiments, the virtual character's appearance can be derived from a preset database or customized based on the patient's preferences. Specifically, the visual interaction device 32 uses the Unity rendering engine to create a preset 3D character database, with character images sourced from various open-source or paid platforms such as MakeHuman, Adobe Mixamo, and Renderpeople. Furthermore, the preset virtual character parameters can be adjusted or a new model can be created to customize the desired appearance based on the patient's needs. Optionally, the customization process can utilize 3D scanning equipment (such as Artec Eva or Structure Sensor) to collect the target person's real facial and body features, generate a high-precision 3D mesh model, and then use the Unity engine in conjunction with the scan data to customize the virtual character model.

[0152] In some embodiments, the facial movements of the virtual character adapted to the voice mainly involve simulating the mouth shapes when a person speaks. Specifically, this is achieved by: using a preprocessing algorithm (such as MFCCs or Wav2Vec 2.0) to extract features from the speech signal of the training stimulus speech generated by the speaker 31, capturing key prosody, syllable intensity, and speech rate information; then dividing the extracted speech features into time steps to ensure synchronization between the audio and subsequent facial animation sequences on the time axis. Based on the VOCA (Voice Operated Character Animation) model, the audio signal is encoded into a low-dimensional speech embedding vector, ensuring a correspondence with dynamic lip shapes and facial expressions. The extracted speech features are then mapped to the expression parameter space of the FLAME (Faces Learned with an Articulated Model and Expressions) model to control the key point movements of the mouth (such as opening, closing, and lip rounding). The FLAME model then coordinates with the expression parameters and speech features to generate a dynamically changing facial motion mesh frame by frame. Finally, the Unity rendering engine is used to render the generated dynamic facial motion mesh as facial movements adapted to the virtual character's appearance.

[0153] In some embodiments, the facial expressions of the virtual character are determined by the interaction attitude output by the interaction strategy generation module 22, mainly consisting of two emotions: comfort and encouragement. Based on voice-driven lip-sync generation, the Expression Layer tool of the FLAME model is used to overlay 50-dimensional expression control parameters of the virtual character on top of the lip-sync adjustment parameters. These expression control parameters can be obtained by fine-tuning the expression parameter library built into the FLAME model.

[0154] In some embodiments, the rehabilitation training content generated by the interaction strategy generation module 22 is adjusted by the lower limb rehabilitation robot 33. The training content includes training duration, training mode, and training impedance. For training impedance, the lower limb rehabilitation robot 33 dynamically adjusts its output torque to change the degree of muscle exertion required by the patient when performing tasks, thereby achieving impedance control targets at different stages of lower limb rehabilitation. As mentioned earlier, the interaction strategy generation module 22 establishes reasonable training content by considering the patient's physical and psychological condition. For example, when the user exhibits any of the following conditions: "low emotional state," "high mental load," "high peripheral fatigue," or "significant decline in task performance," the training intensity and difficulty can be reduced, thereby enhancing the patient's sense of accomplishment and self-efficacy during lower limb rehabilitation training and stimulating their confidence and motivation for rehabilitation.

[0155] In summary, the lower limb motor rehabilitation system with emotional interaction function provided in this disclosure collects multi-source physiological and behavioral signals from patients to determine their psychological and physical state, and then uses a large language model to make real-time decisions to adjust the interaction strategy of the rehabilitation system. This system introduces the positive effects of interpersonal emotional interaction into the traditional human-computer interaction process of motor rehabilitation training, which is primarily led by rehabilitation robots. From personalized adjustment of task content to activate patients' confidence and motivation in rehabilitation, to establishing a virtual avatar to provide patients with visual feedback and verbal encouragement, this system provides socio-emotional interaction and support. Therefore, this disclosure can optimize the function and effect of traditional rehabilitation systems, improve the effectiveness of rehabilitation training, help patients enhance positive emotions and their compliance and participation in training, and conduct high-quality lower limb motor rehabilitation under the guidance of positive emotions, shortening the rehabilitation cycle and improving patients' quality of life.

[0156] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0157] Although embodiments of this disclosure have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A lower limb motor rehabilitation system based on emotional interaction, characterized in that, include: The sensing unit is used to acquire the patient's physiological and behavioral signals; The central controller includes a data processing module and an interaction strategy generation module; The data processing module is used to determine the patient's physical and psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit, and output the determination result as the input of the interaction strategy generation module. The interaction strategy generation module uses a large language model to simulate the professional knowledge and thinking of rehabilitation physicians to understand and integrate reasoning decisions about the patient's current state, and outputs multi-dimensional interaction strategies with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content. The execution unit includes a speaker, a visual interaction device, and a lower limb rehabilitation robot; The speaker is used to play voices with familiar timbre and positive emotional style based on the voice interaction content generated by the interaction strategy generation module during training; the visual interaction device is used to generate virtual characters based on the visual interaction content generated by the interaction strategy generation module, serving as one of the carriers for simulating interpersonal emotional interaction and support, and enabling the virtual characters to produce corresponding facial movements and expressions in conjunction with the voice generated by the speaker; the lower limb rehabilitation robot is used to guide the patient to complete lower limb rehabilitation training movements based on the lower limb rehabilitation training content generated by the interaction strategy generation module. The interaction strategy generation module uses a pre-trained first large language model as its large language model. The pre-training process of the first large language model includes: Step S100, Mind Chain Training: The mind chain of the first major language model is fine-tuned for the first time. By inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, the first major language model is trained to gradually think and reason about the content of lower limb rehabilitation training, as well as the content and attitude of communication. The content of lower limb rehabilitation training includes training duration, training mode and task difficulty. Step S200: Interactive text training: Collect age-appropriate language data and exercise training motivational language data to perform secondary fine-tuning training on the text content output by the first language model after the first fine-tuning, so that the output interactive text is more in line with the lower limb rehabilitation training scenario and meets the patient's expectations. Step S100 specifically includes: Step S110: Collection of existing cases: Collect existing decision-making cases and decision-making processes of licensed professional rehabilitation physicians when performing lower limb training on patients. Each existing case i should record at least: (1) the patient's condition s i (2) An initial artificial rehabilitation training program, which is equivalent to an initial rehabilitation training program p adapted to the lower limb rehabilitation robot. i and the initial rehabilitation training plan p i The normalized representation is a text sequence: "Training content: passive / assisted / active, resistance level, training speed"; (3) The j-th real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and the reasons for the adjustment ij Among them, strategy a will be adjusted ij The text normalization representation is as follows: "Training content adjustment: passive / assisted / active, resistance level, training speed; Communication content adjustment: communication attitude, communication text", where the communication attitude has two labels: "comforting" or "encouraging"; the reason for adjustment is r. ij This includes changes in the patient's psychological state, changes in their physical state, and / or the professional theoretical knowledge that physicians refer to when making adjustments; Step S120: Utilize the second language model to retrieve professional knowledge, and expand the existing cases through reasoning analysis to explain the real-time adjustments to training content and communication content. ij : By designing prompts, the second language model is guided to retrieve professional rehabilitation theories and theories of disease psychology. Based on the zero-shot thought chain training method, instructions for the second language model to think step by step are added to the prompts, enabling the second language model to combine the retrieved knowledge with the adjustment strategy a in the existing case i. ij A step-by-step analysis and reasoning process was conducted to supplement and refine the reasons for the adjustment. ij , and single adjustment strategy a ij The corresponding reason for the adjustment r ij The number of characters in the reasoning text is limited to no more than the first maximum length Lr1; Step S130, Data Standardization: All texts in case i i p i a ij r ij Generate a text combination according to the specified format. i And by combining all the case texts, a first text sequence W={w1, w2, …, w} is generated. i , …, w N }; Step S140, Initial Fine-tuning Training: The standardized first text sequence W was organized into the first fine-tuning dataset. An adapter fine-tuning method was used, and the first large language model was trained based on the Hugging Face Transformers framework; the first cross-entropy loss function was employed. The adapter's parameter θ1 is optimized as the objective function: In the formula, This represents the total number of training adjustment strategies - adjustment reason groups that case i possesses; This indicates the first large language model input after initial training, showing the patient's condition. The initial rehabilitation training plan is then obtained from the output. The probability is used to quantify the ability to derive an initial training plan from the patient's condition; The first large language model used for quantification after the initial training infers the adjustment strategy for the jth time based on the current patient condition, the initial rehabilitation training plan, the rehabilitation training adjustment strategy and reasons before the jth adjustment, and the reasons for the jth adjustment. The probability of; This indicates the reason for the adjustment of rehabilitation training in the jth and previous sessions. This represents the rehabilitation training adjustment strategy before the j-th adjustment; Step S200 specifically includes: Step S210, Corpus Collection: We collected age-appropriate language data and motivational language data for rehabilitation exercise training from the internet using web crawling. Step S220, Data Cleaning and Preprocessing: The collected corpus was cleaned, including removing duplicate text, correcting punctuation errors and grammatical inconsistencies, and correcting typos and non-standard language. The cleaned corpus was then categorized according to semantic function, including motivational statements, comforting statements, task description statements, and everyday communication statements. The maximum number of characters in a single corpus fragment was limited to a second maximum length Lr2. All resulting corpus fragments were then formatted into a second text sequence V={v1, v2, …, v l , …, v M }, v l Represents a single segment of a corpus; Step S230, Secondary Fine-tuning Training The standardized second text sequence V was organized into a second fine-tuning dataset. An adapter fine-tuning method was used, and the first large language model, after its initial fine-tuning, was trained using the Hugging Face Transformers framework. The second cross-entropy loss function was employed. Used as the objective function to fine-tune the model parameters θ2: In the formula, This represents the first language model prediction corpus segment after secondary fine-tuning. The probability of.

2. The lower limb motor rehabilitation system according to claim 1, characterized in that, The lower limb rehabilitation robot can guide patients to complete a full cycle of cycling motion. It also has multiple training modes and difficulty levels, including passive, assisted, and active modes. The difficulty level is adjusted by changing the training speed in passive and assisted modes and the resistance value in active mode. The lower limb rehabilitation robot obtains information from its internal pedal pressure sensor and motor encoder, and feeds back information, including its own movement speed and interaction force, to the sensing unit as the basis for quantitatively evaluating the patient's training performance.

3. The lower limb motor rehabilitation system according to claim 1, characterized in that, The physiological signals acquired by the sensing unit include electroencephalogram (EEG), electromyogram (EMG), electrocardiogram (ECG), and electrodermal signals, and the behavioral signals acquired include facial expressions, speech, eye movements, and motor task performance information. The process by which the data processing module determines the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit includes: preprocessing and preliminary feature extraction of the acquired signals respectively; Subsequently, a deep learning-based classification model was used to extract deeper features from the various preliminary features to obtain the patient's psychological state.

4. The lower limb motor rehabilitation system according to claim 3, characterized in that, The patient's psychological state is divided into emotional state and mental load level. The emotional state is divided into three levels: positive, neutral and negative. The mental load level is divided into three levels: high, medium and low. The process involves using a deep learning-based classification model to extract deeper features from the various preliminary features to obtain the patient's psychological state, specifically including: Each extracted preliminary feature is treated as a modality. A deep learning-based classification model is used to perform feature fusion and classification of the single modality to obtain the preliminary psychological state classification results of the patients. Each modality uses a corresponding classification model. During the training of each classification model, a labeled dataset is used for supervised learning. The labels are based on the subjective emotion scale and the perceived stress scale. Based on the scores given by the subjects, the scores are divided into high, medium and low training labels according to the scale standards to classify emotions and mental load. Based on rules, the preliminary psychological state classification results of all single modalities are fused into a multimodal decision. The rules are as follows: the weight of each modality in the decision fusion is dynamically adjusted according to the individual classification performance of each modality. The modality with better classification performance is given a higher weight. The classification labels of each modality are linearly weighted and summed to obtain a weighted result. The weighted result is then normalized to obtain the final psychological state classification result of the patient, which serves as the judgment result of the patient's psychological state.

5. The lower limb motor rehabilitation system according to claim 1, characterized in that, The patient's physical condition was divided into peripheral fatigue level and rehabilitation task performance; The peripheral fatigue level is used to visually reflect the patient's muscle state and is divided into three levels: high, medium, and low. The data processing module obtains peripheral fatigue information based on the patient's electromyography (EMG) signals, specifically including: preprocessing the acquired EMG signal data to obtain the peak EMG amplitude A for the current time period, and comparing this peak EMG amplitude A with the peak EMG amplitude A within the first 10 seconds of training. max When comparing, when A / A max When the fatigue level is between 80% and 100%, it is considered a low fatigue state. When A / A... max Between 60% and 80%, it is judged as a moderate fatigue state, when A / A max When the level is 60% or below, it is considered a state of high fatigue. The performance of the rehabilitation task is related to the type and content of the lower limb rehabilitation training task. The data processing module uses movement speed and movement symmetry as evaluation indicators. The movement speed is characterized by the circumferential rotation speed of the lower limb rehabilitation robot's own mechanism, and the movement symmetry is characterized by the ratio of the average positive pressure applied by the left and right legs to the pedals during the current full cycle of pedaling, collected by the lower limb rehabilitation robot. The data processing module integrates and identifies the patient's physical and psychological state, and outputs the results in a unified natural language format. --Psychological state-- Emotional state: positive / neutral / negative; Mental workload: High / Medium / Low; --Physical condition-- Peripheral fatigue: High / Medium / Low; Task performance: Motion speed %a, Motion symmetry %b In this context, "%a" and "%b" are both specific numerical values.

6. The lower limb motor rehabilitation system according to claim 1, characterized in that, By designing prompt words, the pre-trained first language model generates standardized output, which includes: --Training Content-- Training duration: %c minutes; Training modes: Active / Assisted / Passive; Training impedance: %d N; --Virtual Interpersonal Interaction Content-- Voice-interactive text: %e; Interaction attitude: comforting / encouraging; --Reason for Adjustment-- %f” In this context, "%c" and "%d" represent specific numerical values, while "%e" and "%f" represent specific text content.

7. The lower limb motor rehabilitation system according to claim 6, characterized in that, The voice interaction text and interaction attitude generated by the interaction strategy generation module are converted into voice with timbre characteristics familiar to the patient and a positive emotional style, and then interacted with the patient through the speaker. Specific steps include: Speech samples from the target speaker are collected, and a text-to-speech model is used to extract timbre features from the collected speech samples, including fundamental frequency, formants, and speech duration, to establish a feature vector of the target timbre. The text-to-speech model is then used to convert text to audio, and the extracted target timbre feature vector is embedded into the speech generated by the text-to-speech model. Through feature fusion, the audio output by the text-to-speech model possesses the timbre features of the target speaker. Generative adversarial networks are used to extract emotional features from a sample speech library with positive motivational and comforting emotions. Based on the extracted emotional features, the fundamental frequency, Mel spectrum, and energy distribution of the audio output by the text-to-speech model are adjusted so that the audio output by the speaker has timbre features familiar to the patient and a positive emotional style.

8. The lower limb motor rehabilitation system according to claim 6, characterized in that, The visual interaction device is selected from a display screen, a virtual reality device, or an augmented reality device; the virtual character generated by the visual interaction device includes three aspects: the character's appearance, facial movements corresponding to voice, and facial expressions. The appearance of the character is derived from a pre-set database or customized based on the patient's preferences; The facial movements corresponding to the voice of the character are generated in the following way: Features are extracted from the voice signal output by the speaker to capture key prosody, syllable intensity, and speech rate information; then, the extracted voice features are divided into time steps to ensure synchronization between the audio and subsequent facial animation sequences on the time axis; the audio signal is encoded into a low-dimensional voice embedding vector to ensure correspondence with dynamic lip movements and facial expressions, and the extracted voice features are mapped to the motion parameter space of the FLAME model to control the movement of key points of the mouth; the FLAME model, in conjunction with reference to facial expression parameters and voice features, generates a dynamically changing facial motion mesh frame by frame; finally, the generated facial motion mesh is rendered into facial movements adapted to the appearance of the virtual character using a rendering engine. The facial expressions of the character are determined by the interactive attitude output by the interaction strategy generation module. Based on the lip movement adjustment obtained from the extracted speech features, the expression layer tool of the FLAME model is used to overlay 50-dimensional expression control parameters of the virtual character on the lip adjustment parameters. The expression control parameters are obtained by fine-tuning the expression parameter library built into the FLAME model.

9. The lower limb motor rehabilitation system according to claim 6, characterized in that, The lower limb rehabilitation robot adjusts its own state in real time based on the training duration, training mode, and training impedance in the training content generated by the interaction strategy generation module.

Citation Information

Patent Citations

  • Exercise rehabilitation robot interactive control method based on emotion perception

    CN104287747A

  • Lower limb rehabilitation training robot system based on virtual reality

    CN107049702A