Lower limb exercise rehabilitation system based on emotional interaction

By introducing emotional interaction technology into the rehabilitation system, using deep learning and generative large language models to simulate interpersonal interactions, the problem that existing rehabilitation equipment cannot provide emotional support is solved, and the effect of improving rehabilitation efficiency and quality of life is achieved.

CN119943274AActive Publication Date: 2025-05-06TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202510103414.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

Existing rehabilitation equipment cannot replace the emotional support provided by the physician during the patient's rehabilitation process, resulting in a negative impact on the patient's emotional state and affecting the rehabilitation efficiency.

Method used

Design a lower limb motor rehabilitation system based on emotional interaction, using multi-sensing fusion, deep learning and generative large language model to simulate the positive emotional effects in the interpersonal interaction process, provide an immersive rehabilitation experience, and stimulate patients' positive emotions and training motivation.

Benefits of technology

By introducing interpersonal emotional interactions, the functions and effects of traditional rehabilitation systems are optimized, the effectiveness of rehabilitation training is improved, and patients can improve their positive emotions, adherence and participation in training, shorten the rehabilitation cycle, and improve their quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943274A_ABST
    Figure CN119943274A_ABST
Patent Text Reader

Abstract

The invention provides a lower limb exercise rehabilitation system based on emotion interaction, and the system comprises a sensing unit which is used for obtaining a physiological signal and a behavior signal of a patient; in the central controller, a data processing module discriminates the physical state and psychological state of a patient in real time, and an interaction strategy generation module simulates professional knowledge and thinking modes of a rehabilitation physician based on a large language model to understand the current state of the patient and fuse reasoning decisions, and outputs a multi-dimensional interaction strategy with the patient. In the execution unit, a loudspeaker is used for playing voice with timbre characteristics and forward emotion styles familiar to a patient according to voice interaction content during training; the visual interaction device generates a virtual character according to the visual interaction content, and the virtual character is matched with the voice generated by the loudspeaker to generate corresponding facial actions and expressions; the lower limb rehabilitation robot guides the patient to complete lower limb rehabilitation training actions according to the lower limb rehabilitation training content. Extra emotion support is provided for training of the patient, and the rehabilitation effect is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of elderly care and rehabilitation equipment and human-computer interaction, and specifically relates to a lower limb movement rehabilitation system based on emotional interaction. Background Art

[0002] With the development of global deep aging, the number of elderly people with disability, dementia, and cognitive impairment continues to increase. Most of them have limb movement disorders and suffer from both physical and psychological torture. Medical theory and clinical medicine have proved that correct and scientific rehabilitation training plays a very important role in the recovery and improvement of motor function. In order to solve the problems of lack of professional caregivers and expensive medical costs, safe, quantitative, effective and repetitive rehabilitation training devices have emerged. However, the rehabilitation process brought by purely mechanical equipment is often boring and monotonous, and the existing rehabilitation equipment is far from replacing the emotional support provided by doctors during the patient's rehabilitation process, which makes the patient's emotional state negatively affected, greatly affecting their rehabilitation efficiency.

[0003] An excellent rehabilitation physician can not only flexibly adjust the training tasks according to the patient's current condition to inspire their confidence and motivation for rehabilitation, but also provide them with great emotional help through appropriate communication. A good doctor-patient relationship has been proven to play a vital role in improving medical outcomes, including motor learning performance. Therefore, solving the problem that existing rehabilitation devices still cannot replace rehabilitation physicians, giving robots the emotional interaction capabilities comparable to rehabilitation physicians, introducing the positive effects brought about by interpersonal emotional communication, and creating an immersive and positive rehabilitation environment to enhance the rehabilitation effect are important development directions for sports rehabilitation technology, but the corresponding technology has not yet been formed. Summary of the invention

[0004] The present disclosure aims to solve one of the technical problems existing in the existing related technologies at least to a certain extent.

[0005] To this end, the disclosed embodiment proposes a lower limb motor rehabilitation system based on emotional interaction, which utilizes multi-sensor fusion, deep learning and a generative large language model to simulate the positive emotional effects in the process of interpersonal interaction, adds interpersonal emotional interaction functions to the design of the rehabilitation system, and combines visual, auditory interaction devices and rehabilitation training equipment to provide patients with an immersive rehabilitation experience, stimulate patients' positive emotions and training motivation, thereby improving the rehabilitation effect and promoting the recovery of lower limb motor function.

[0006] In order to achieve the above objectives, a lower limb movement rehabilitation system based on emotional interaction disclosed in the present invention includes:

[0007] A sensing unit, used to obtain physiological and behavioral signals of the patient;

[0008] The central controller includes a data processing module and an interaction strategy generation module; the data processing module is used to judge the patient's physical state and psychological state in real time according to the physiological signals and behavioral signals acquired by the sensor unit, and output the judgment result as the input of the interaction strategy generation module; the interaction strategy generation module simulates the professional knowledge and thinking mode of rehabilitation physicians based on the large language model to understand the patient's current state and make integrated reasoning decisions, and output a multi-dimensional interaction strategy with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content;

[0009] The execution unit includes a speaker, a visual interaction device and a lower limb rehabilitation robot; the speaker is used to play voice with timbre characteristics familiar to the patient and a positive emotional style according to the voice interaction content generated by the interaction strategy generation module during training; the visual interaction device is used to generate a virtual character according to the visual interaction content generated by the interaction strategy generation module, as one of the carriers for simulating interpersonal emotional interaction and support, and to make the virtual character produce corresponding facial movements and expressions in conjunction with the voice generated by the speaker; the lower limb rehabilitation robot is used to guide the patient to complete the lower limb rehabilitation training movements according to the lower limb rehabilitation training content generated by the interaction strategy generation module.

[0010] In some embodiments, the lower limb rehabilitation robot can guide the patient to complete a full week of cycling exercise, and has multiple training modes and difficulty levels, including passive, assisted, and active. The change in difficulty level is achieved by adjusting the training speed during passive and assisted mode training, and the resistance value during active mode training; the lower limb rehabilitation robot obtains information from its internal pedal pressure sensor and motor encoder, and feeds back information including its own mechanism movement speed and interaction force to the sensing unit as basic data for quantitative evaluation of the patient's training task performance.

[0011] In some embodiments, the physiological signals acquired by the sensing unit include electroencephalogram, electromyography, electrocardiogram and electrodermal signals, and the behavioral signals acquired include facial expressions, speech, eye movements and motor task performance information; the process of the data processing module determining the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit includes: preprocessing and preliminary feature extraction of each type of acquired signal; then using a deep learning-based classification model to perform deep feature extraction on the extracted preliminary features to obtain the patient's psychological state.

[0012] In some embodiments, the patient's psychological state is divided into an emotional state and a mental workload, the emotional state is divided into three levels: positive, neutral, and negative, and the mental workload is divided into three levels: high, medium, and low;

[0013] The deep learning-based classification model is used to perform deep feature extraction on various extracted preliminary features to obtain the patient's psychological state, specifically including:

[0014] Each type of extracted preliminary features is taken as a modality, and a classification model based on deep learning is used to perform single-modal feature fusion and classification to obtain the preliminary classification results of the patient's psychological state; wherein, each modality uses a corresponding classification model, and during the training process of each classification model, a labeled data set is used for supervised learning. The labels are assessed based on the subjective emotion scale and the perceived stress scale. Through the subjects' scores, the scores are divided into high, medium, and low training labels according to the scale standards to classify emotions and mental loads;

[0015] Based on the rules, all the preliminary psychological state classification results of single modalities are subjected to multimodal decision fusion, wherein the rule is to dynamically adjust the weight of each modality in the decision fusion according to its individual classification performance, and the modality with better classification performance is assigned a higher weight. The weighted result is obtained by linearly weighted summation of the classification labels of each modality, and the weighted result is normalized to obtain the patient's final psychological state classification result as the discrimination result of the patient's psychological state.

[0016] In some embodiments, the patient's physical condition is divided into peripheral fatigue level and rehabilitation task performance;

[0017] The peripheral fatigue degree is used to directly reflect the patient's muscle state and is divided into three levels: high, medium and low. The data processing module obtains peripheral fatigue information according to the patient's electromyographic signal, specifically including: performing preprocessing operations on the acquired electromyographic signal data to obtain the peak electromyographic amplitude A of the current period, and comparing the peak electromyographic amplitude A with the peak electromyographic amplitude A within the first 10 seconds of training. max When A / A max When it is between 80% and 100%, it is considered as low fatigue state. max Between 60% and 80%, it is considered a moderate fatigue state. max When it is 60% or below, it is considered as a high fatigue state;

[0018] The performance of the rehabilitation task is related to the type and content of the lower limb rehabilitation training task. The data processing module uses movement speed and movement symmetry as evaluation indicators. The movement speed is characterized by the circumferential speed of the lower limb rehabilitation robot itself collected by the lower limb rehabilitation robot, and the movement symmetry is characterized by the average positive pressure ratio of the left and right legs applied to the pedals during the current training cycle phase collected by the lower limb rehabilitation robot.

[0019] The data processing module integrates the identified physical and psychological states of the patient and outputs them in a unified natural language format:

[0020] "--Mental State--

[0021] Emotional state: positive / neutral / negative;

[0022] Mental workload: high / medium / low;

[0023] --Physical Condition--

[0024] Peripheral fatigue: high / medium / low;

[0025] Task performance: movement speed %a, movement symmetry %b"

[0026] Wherein, “%a” and “%b” are specific values.

[0027] In some embodiments, the large language model used by the interaction strategy generation module is a pre-trained first large language model, and the pre-training process of the first large language model includes:

[0028] Step S100, thinking chain training: performing the first fine-tuning training on the thinking chain of the first language model, by inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, training the first language model to gradually think and reason about the content of lower limb rehabilitation training and the ability to communicate content and attitude, wherein the lower limb rehabilitation training content includes training duration, training mode and task difficulty;

[0029] Step S200, interactive text training: collect aging-friendly corpus and sports training motivational corpus to perform secondary fine-tuning training on the text content output by the first language model after the first fine-tuning, so that the interactive text output by it is more suitable for the lower limb rehabilitation training scenario and meets the patient's expectations.

[0030] In some embodiments, step S100 specifically includes:

[0031] Step S110: Collection of existing cases:

[0032] Collect existing decision-making cases and decision-making ideas of licensed professional rehabilitation physicians when performing lower limb training on patients. An existing case should at least record: (1) the patient's condition i , including injury time, lesion area and movement assessment; (2) an initial artificial rehabilitation training program, which is equivalent to an initial rehabilitation training program p adapted for the lower limb rehabilitation robot i and the initial rehabilitation training program p iNormalized representation as a text sequence: "Training content: passive / assisted / active, resistance level, training speed"; (3) The jth real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and reasons for adjustment ij , where strategy a will be adjusted ij The text normalization of the training content is: "Training content adjustment: passive / assisted / active, resistance level, training speed; communication content adjustment: communication attitude, communication text", where the communication attitude has two labels: "comfort" or "encouragement"; the reason for adjustment is: ij Includes changes in the patient's mental state, physical state, and / or professional theoretical knowledge that the physician refers to when making adjustments;

[0033] Step S120: Retrieve professional knowledge using the second language model, and expand the real-time adjustment reasons for the training content and communication content in the existing cases through reasoning analysis. ij :

[0034] By designing prompt words, the second largest language model is prompted to retrieve professional rehabilitation theory and disease psychology theory. Based on the zero-sample thinking chain training method, instructions for the second largest language model to think step by step are added to the prompt words, so that the second largest language model can combine the retrieved knowledge to adjust the strategy a in the existing case i. ij Conduct step-by-step analytical reasoning to supplement and improve the reasons for adjustment ij , and the single adjustment strategy a ij Corresponding adjustment reason r ij The number of characters in the inference text is limited to not exceed the first maximum length Lr1;

[0035] Step S130: Data standardization:

[0036] All text s of case i i 、p i 、a ij 、r ij Generate text combination w according to the specified format i , and use all case text combinations to generate the first text sequence W = {w1,w2,…,w i ,…,w N};

[0037] Step S140: First fine-tuning training:

[0038] The standardized first text sequence W is organized into a first fine-tuning dataset, and the first language model is trained using the adapter fine-tuning method and the Hugging Face Transformers framework; wherein the first cross entropy loss function L is used. LOSS1The adapter parameter θ1 is optimized as the objective function:

[0039]

[0040] In the formula, |r i | represents the total number of training adjustment strategy-adjustment reason groups for case i; P(p i |p i ,s i ; θ1) represents the first language model input patient condition s after the first training i The output is then used to obtain the initial rehabilitation training program p i The probability of quantifying the ability to infer the initial training plan from the patient's condition; P(a i,j |p i ,r i,≤j ,s i ,a i,<j θ1) is used to quantify the first language model after the first training. Based on the current patient condition, the initial rehabilitation training program, the rehabilitation training adjustment strategy and adjustment reason before the jth adjustment, and the adjustment reason for the jth adjustment, the jth adjustment strategy a is inferred. i,j The probability of i,≤j represents the adjustment reason for the jth and previous rehabilitation training, a i,<j Represents the rehabilitation training adjustment strategy before the jth adjustment.

[0041] In some embodiments, step S200 specifically includes:

[0042] Step S210: Corpus collection:

[0043] Collect aging-friendly corpus and rehabilitation exercise training motivational corpus from the Internet through web crawlers;

[0044] Step S220: Data cleaning and preprocessing:

[0045] The collected corpus is cleaned, including removing duplicate texts, correcting punctuation errors and grammatical confusion, and correcting typos and language irregularities in the corpus; the cleaned corpus is divided into different categories according to semantic functions, including motivational sentences, comforting sentences, task description sentences and daily communication sentences; the maximum number of characters in a single corpus segment is limited to not more than the second maximum length Lr2; all the obtained corpus segments are formatted into a second text sequence V = {v1, v2, ..., v l ,…,v M},v l Represents a single corpus fragment;

[0046] Step S230: Second fine-tuning training

[0047] The standardized second text sequence V is organized into the second fine-tuning dataset, and the first language model after the first fine-tuning is trained using the adapter fine-tuning method and the framework of Hugging Face Transformers; wherein the second cross entropy loss function L is used LOSS2 As the objective function to fine-tune the model parameters θ2:

[0048]

[0049] In the formula, P(v l ; θ2) represents the first language model prediction corpus segment v after secondary fine-tuning l probability.

[0050] In some embodiments, prompt words are designed to enable the pre-trained first language model to generate standardized output, and the output content includes:

[0051] "--Training content--

[0052] Training duration: %c minutes;

[0053] Training mode: active / assisted / passive;

[0054] Training impedance: %dN.

[0055] --Virtual human interaction content--

[0056] Voice interaction text: %e;

[0057] Interaction attitude: comfort / motivation.

[0058] --Reason for adjustment--

[0059] %f."

[0060] Among them, “%c” and “%d” are specific values, and “%e” and “%f” are specific text contents.

[0061] In some embodiments, the speech interaction text and interaction attitude generated by the interaction strategy generation module are converted into speech with timbre characteristics and positive emotional style familiar to the patient, and the speaker is used to interact with the patient. The specific steps include:

[0062] The speech samples of the target speaker are collected, and the timbre features of the collected speech samples are extracted using a text-to-speech model, including fundamental frequency, resonance peaks and speech duration, to establish a feature vector of the target timbre; the text-to-speech model is used to realize text-to-audio conversion, and the extracted target timbre feature vector is embedded into the speech generated by the text-to-speech model, and through feature fusion, the audio output by the text-to-speech model has the timbre features of the target speaker; the generative adversarial network is used to extract the emotional features of the sample speech library with positive motivation and comfort emotions; based on the extracted emotional features, the fundamental frequency value, Mel spectrum and energy distribution of the audio output by the text-to-speech model are adjusted, so that the audio output by the speaker has the timbre features and positive emotional style familiar to the patient.

[0063] In some embodiments, the visual interaction device is a display screen, a virtual reality device or an augmented reality device; the virtual character image generated by the visual interaction device includes three aspects: the character's appearance, the character's facial movements corresponding to the voice, and the character's facial expressions; wherein,

[0064] The character appearance is from a preset database or customized based on the patient's preferences;

[0065] The facial movements corresponding to the speech of the character are generated in the following manner: extracting features from the speech signal output by the speaker to capture key rhythm, syllable strength and speech rate information in the speech; then dividing the extracted speech features into time steps to ensure synchronization of the audio and subsequent facial animation sequences on the time axis; encoding the audio signal into a low-dimensional speech embedding vector to ensure the correspondence with the dynamic mouth shape and facial expression, and mapping the extracted speech features to the action parameter space of the FLAME model to control the key point movement of the mouth, and the FLAME model collaboratively references the facial expression parameters and speech features to generate a facial movement mesh that changes dynamically frame by frame; finally, using a rendering engine to render the generated facial movement mesh into facial movements that match the appearance of the virtual character;

[0066] The facial expression of the character is determined by the interaction attitude output by the interaction strategy generation module. On the basis of adjusting the lip movement of the extracted speech features, the expression layer tool of the FLAME model is used to superimpose 50-dimensional expression control parameters of the virtual character on the basis of the lip adjustment parameters. The expression control parameters are obtained by fine-tuning the expression parameter library of the FLAME model.

[0067] In some embodiments, the lower limb rehabilitation robot adjusts its own state in real time according to the training duration, training mode and training impedance in the training content generated by the interaction strategy generation module.

[0068] The present disclosure has the following characteristics and beneficial effects:

[0069] The embodiment of the present disclosure provides a lower limb motor rehabilitation system based on emotional interaction. By introducing the positive effects of interpersonal emotional interaction in the human-computer interaction process of motor rehabilitation training, the system activates the user's confidence and motivation for rehabilitation through flexible task adjustment, establishes a virtual physician image to provide patients with visual feedback and voice social emotional interaction and support, optimizes the functions and effects of traditional rehabilitation systems, greatly improves the effectiveness of rehabilitation training, helps patients improve their positive emotions as well as their compliance and participation in training, and conducts high-quality lower limb motor rehabilitation under the guidance of positive emotions, shortens the rehabilitation cycle, and improves the quality of life of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] Figure 1 is a structural schematic diagram of a lower limb movement rehabilitation system based on emotional interaction provided by an embodiment of the present disclosure;

[0071] Figure 2 yes Figure 1 The schematic diagram of the lower limb motor rehabilitation system for classifying and rating the patient's psychological and physical status;

[0072] Figure 3 yes Figure 1 The schematic diagram of the specific process of the lower limb motor rehabilitation system analyzing the patient's psychological and physical state;

[0073] Figure 4 yes Figure 1 The schematic diagram of the specific process of the lower limb motor rehabilitation system generating a multi-dimensional interaction strategy;

[0074] Figure 5 yes Figure 1 Schematic diagram of the interaction process performed by the lower limb motor rehabilitation system shown. DETAILED DESCRIPTION

[0075] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0076] On the contrary, the present application covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present application as defined by the claims. Further, in order to make the public have a better understanding of the present application, some specific details are described in detail in the detailed description of the present invention below. For those skilled in the art, the present application can be fully understood without the description of these details.

[0077] like Figure 1 As shown, a lower limb movement rehabilitation system based on emotional interaction according to an embodiment of the present disclosure includes:

[0078] Sensor unit 1, which is used to obtain the patient's physiological signals and behavioral signals in an all-round way to provide data support for rehabilitation training. The acquired physiological signals include but are not limited to EEG, EMG, ECG and EGG signals, and the acquired behavioral signals include but are not limited to facial expressions, speech, eye movements and motor task performance information;

[0079] The central controller 2 includes a data processing module 21 and an interaction strategy generation module 22; the data processing module 21 is used to judge the patient's physical state and psychological state in real time according to the physiological signals and behavioral signals obtained by the sensor unit 1, and output the judgment result as the input of the interaction strategy generation module 22. The interaction strategy generation module 22 simulates the professional knowledge and thinking mode of the rehabilitation physician based on the large language model to understand the patient's current state and make integrated reasoning decisions, and output a multi-dimensional interaction strategy with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content;

[0080] The execution unit 3 includes a speaker 31, a visual interaction device 32 and a lower limb rehabilitation robot 33; the speaker 31 is used to play voices with timbre characteristics familiar to the patient and with a motivational or comforting emotional style according to the voice interaction content during training; the visual interaction device 32 is used to generate a virtual character according to the visual interaction content as one of the carriers for simulating interpersonal emotional interaction and support, and to produce corresponding facial movements and expressions for the generated virtual character in conjunction with the voice generated by the speaker 31; the lower limb rehabilitation robot 33 is used to guide the patient to complete the lower limb rehabilitation training movements according to the lower limb rehabilitation training content.

[0081] In some embodiments, the voice emitted by the speaker 31 has the characteristics of being close to the timbre of someone familiar or close to the user, and having positive effects such as encouragement, comfort, and care emotionally.

[0082] In some embodiments, the lower limb rehabilitation robot 33 can help patients gradually rebuild the motor function and gait function of the lower limb hip, knee and ankle joints. Its basic structure can be a seated or recumbent bicycle, the end of which is connected to the soles of the patient's feet. It has the function of lower limb motor rehabilitation training and can directly guide patients to complete a full week of bicycle exercise. It also has passive, assisted, and active training modes and difficulty levels. The change in difficulty level can be achieved by adjusting the training speed during passive and assisted mode training, and the resistance value during active mode training.

[0083] Furthermore, the passive, assisted, and active training modes of the lower limb rehabilitation robot 33 can be implemented by any one of the algorithms such as PID, fuzzy control, admittance control, and impedance control, and the training speed and training resistance value can be changed by adjusting the parameters of the control algorithm. The training speed can range from 0 to 20 mm / s, and the resistance value can range from 0 to 10 N. In addition, the lower limb rehabilitation robot 33 also obtains the information of its pedal pressure sensor and motor encoder, and feeds back the information such as its own mechanism movement speed and interaction force to the sensor unit 1 as the basic data for quantitative evaluation of the patient's training task performance.

[0084] In some embodiments, the sensing unit 1 includes a physiological signal acquisition device and a behavioral signal acquisition device. The physiological signal acquisition device includes an electroencephalogram signal acquisition device, an electrocardiogram signal acquisition device, an electromyography signal acquisition device, and an electrodermal signal acquisition device; the behavioral signal acquisition device includes an eye tracker, a voice acquisition device, a facial expression capture device, and may also include a force sensor and a motor encoder integrated in the lower limb rehabilitation robot 33. All types of equipment in the sensing unit 1 are commercially available products.

[0085] In some embodiments, in order to provide patients with an immersive rehabilitation experience, stimulate patients' positive emotions and training motivation, and thus improve rehabilitation effects, it is necessary to accurately identify the patient's current physiological and psychological conditions before generating a multi-dimensional interaction strategy. Figure 2 , the patient's psychological state can be further subdivided into emotional state and mental load degree, and divided into three levels, namely positive, neutral, and negative emotions, and high, medium, and low mental load. In order to identify the patient's emotional state and mental load degree label, the disclosed embodiment makes full use of the patient's physiological and behavioral signals obtained by the sensor unit 1, including but not limited to EEG, ECG, skin electricity, facial expression, voice, eye tracking, motor task performance and other data.

[0086] In some embodiments, see Figure 3 The process of the data processing module 21 determining the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensor unit 1 includes: preprocessing and preliminary feature extraction of various acquired signals respectively; then using a deep learning-based classification model to perform deep feature extraction on the extracted various preliminary features to obtain the patient's psychological state.

[0087] Furthermore, the specific process of the data processing module 21 preprocessing and extracting preliminary features of various acquired signals includes:

[0088] De-noising the EEG signal data, and using independent principal component analysis (ICA) to remove eye movement and heartbeat artifacts, and using Fourier transform or wavelet transform to extract frequency domain features such as the power spectral density of the alpha, beta, theta and delta frequency bands in the EEG signal;

[0089] For ECG signal data, it is preprocessed by removing baseline drift, denoising and detecting R waves, and then extracting time domain features such as heart rate variability and RR interval, and using Fourier transform to extract low-frequency and high-frequency components of the signal;

[0090] For the skin electrical signal data, firstly, the preprocessing operation of denoising and sliding average artifact removal is performed, and then the amplitude characteristics of the skin electrical response in the target time period are extracted;

[0091] For the facial expression signal data, each frame of the collected image is pre-processed by denoising and normalization, and the key point positions of the patient's facial features are obtained by combining the opencv facial feature detection method, and then the facial action unit features are extracted based on the Facial Action Coding System (FACS);

[0092] For speech signal data, denoising preprocessing is first performed, and then the fundamental frequency, pitch characteristics and Mel frequency cepstrum coefficients are extracted based on Fourier transform to segment the speech signal into short-time frames, calculate the number of speech frames per unit time, obtain the speech speed characteristics, calculate the energy of each frame of speech signal, and estimate its volume characteristics;

[0093] For the eye tracking signal data, background denoising preprocessing is first performed, and then the movement information of the eyeball is obtained, the position coordinates of the eye at each moment are output, and the trajectory sequence formed by the position coordinates in the time step is marked. The Kalman filter method can be used to perform smooth interpolation processing on the trajectory sequence to obtain the final eye movement trajectory characteristics.

[0094] Furthermore, when the data processing module 21 performs deep feature extraction on various preliminary features, each type of preliminary feature is taken as a modality, and a classification model based on deep learning is first used to perform single-modal feature fusion and classification to obtain the patient's preliminary psychological state label, that is, the emotional state level (including three levels: positive, neutral, and negative) and the mental load level (including three levels: high, medium, and low); then, all single-modal classification results are decision-fused based on rules to obtain the patient's final psychological state label. This process can be achieved by selecting a classification model based on deep learning such as convolutional neural networks to fuse and classify various preliminary features. In order to avoid redundancy and model overfitting caused by the high dimensionality of multimodal features, a hybrid multimodal fusion strategy can be adopted to combine feature fusion and decision fusion to classify psychological states. The specific steps are as follows:

[0095] Single-modal feature fusion stage: n preliminary features f of k-modal signals k1 ,f k2 ,…,f kn Perform fusion and splicing to form the feature vector f corresponding to mode k k , train independent classification models for signals of each modality, and the classification models corresponding to different modalities can be homogeneous (i.e., the same model structure) or heterogeneous (different model structures). In one embodiment of the present application, the convolutional neural network CNN, the recurrent neural network RNN, and the long short-term memory model LSTM are trained according to the characteristics of the modal features to achieve classification, and the classification label y of the single modality k is obtained. k , the classification label corresponds to the patient's preliminary psychological state label, that is, the emotional state level (divided into positive, neutral, and negative), and the mental workload level (divided into high, medium, and low). During the training of each classification model, a labeled data set is used for supervised learning, and the label is assessed based on the subjective emotion scale (such as PANAS) and the perceived stress scale. Through the subjective scoring of the subjects, the scores are divided into high, medium, and low training labels according to the scale standards for emotion and mental workload classification; during the training process, the cross entropy loss function can be minimized to optimize the classification accuracy of the classification model, and the K-fold cross classification method can be used to verify the classification effect of the pre-trained classification model.

[0096] Multimodal decision fusion stage: Based on the classification results of each modality obtained in the single-modal feature fusion stage, rules are formulated for decision fusion. Specifically, the classification effect of each modality can be evaluated based on the aforementioned K-fold cross-validation method. The indicators used are such as accuracy, precision, recall, and F1 score. The weight of each modality in decision fusion is dynamically adjusted according to its individual classification performance. k , giving higher weights to the modalities with better classification performance. Then the classification labels y of each modality k Perform linear weighted summation to obtain the weighted result Σr k y k The weighted result is passed through a fully connected layer with softmax as the operation function, and the final classification label of the psychological state is output as the discrimination result of the patient's psychological state.

[0097] In some embodiments, see Figure 2The data processing module 21 divides the patient's physical state into peripheral fatigue degree and rehabilitation task performance according to the acquired physiological signals (mainly electromyographic signals) and behavioral signals. The peripheral fatigue degree can directly reflect the patient's muscle state, which is also divided into three levels: high, medium and low. The electromyographic signal can be used to obtain peripheral fatigue information. The specific steps are: the acquired electromyographic signal data is preprocessed by removing baseline drift, denoising, and sliding average to obtain the peak electromyographic amplitude A of the current period, and the peak electromyographic amplitude A is compared with the peak electromyographic amplitude A within the first 10 seconds of training. max When A / A max When it is between 80% and 100%, it is considered as low fatigue state. max Between 60% and 80%, it is considered a moderate fatigue state. max When it is at 60% or below, it is judged as a high fatigue state. The performance of rehabilitation tasks is directly related to the type and content of the task. For the lower limb rehabilitation robot of the disclosed embodiment, the task performance may include movement speed, which can reflect the continuity and proficiency of the patient's lower limb movements, and movement symmetry, that is, the patient's ability to control the uniform and stable force of the left and right leg muscles during rehabilitation exercises; the original data of task performance such as movement speed and movement symmetry are provided by the built-in sensor module of the lower limb rehabilitation robot, and a certain value can be calculated in each stage of training. For example, the movement speed can be characterized by the circumferential speed of the mechanism, and the movement symmetry Symmetry can be calculated by the following formula:

[0098]

[0099] Among them, F left represents the average pressure applied to the pedal by the patient's left leg after a full cycle, F right Represents the average pressure applied to the pedal by the patient's right leg after a full cycle.

[0100] In some embodiments, the data processing module 21 finally integrates the identified physical and psychological states of the patient, and outputs the identification results in natural language by writing Python code. The output natural language is written in the following unified format:

[0101] "--Mental State--

[0102] Emotional state: positive / neutral / negative;

[0103] Mental workload: high / medium / low;

[0104] --Physical Condition--

[0105] Peripheral fatigue: high / medium / low;

[0106] Task performance: Movement speed %a, movement symmetry %b. ”

[0107] Wherein, “%a” and “%b” are specific values.

[0108] In some embodiments, see Figure 4 The interactive strategy generation module 22 is based on the generative large language model. The output multi-dimensional interactive strategy includes voice and visual interactive content with positive emotions such as comfort and encouragement, as well as training content for lower limb rehabilitation robots 33. The flexible interactive control method of the multi-dimensional interactive strategy is significantly better than the preset control program used in previous human-computer interactions, and shows great potential in simulating the human thinking decision-making process and real interpersonal emotional interaction. Among them:

[0109] The speech interaction content includes language text and speech audio generated based on the language text. Its characteristics include: the language text content fully combines the patient's actual training performance, and can reinforce the patient's correct or good behavior through positive language and text, and correct wrong or poor behavior. The characteristics of the speech audio are that the timbre is similar to the timbre of the person close to the patient, and the emotion has the style of motivation, encouragement, and comfort. Furthermore, in the speech interaction content, the timbre characteristics of the person close to the patient are extracted before rehabilitation training, and the emotional characteristics of motivation and comfort are preset by this system. The language text content and the final interactive audio that integrates the timbre and emotional characteristics are generated in real time by the generative large language model during the lower limb rehabilitation training task.

[0110] The visual interaction content is a virtual rehabilitation physician or other virtual character image, which has the following characteristics: the virtual character image can communicate with the patient through facial expressions, gestures, and body postures, and establish a positive "interpersonal" relationship with the patient through voice audio, thereby stimulating the patient's positive emotions during the lower limb rehabilitation training. In addition, the characteristics of the virtual character image can be selected as a preset image according to the patient's wishes, or it can be customized according to the patient's needs.

[0111] The training content for the lower limb rehabilitation robot 33 may include information such as training duration, training mode, and task difficulty. The characteristics include: the training content should be reasonably set to maintain a certain degree of challenge, but can give patients a sense of success in achieving their goals, thereby stimulating patients' confidence and motivation in training.

[0112] Furthermore, in order to improve the performance and adaptability of the generative large language model in the lower limb rehabilitation training scenario and make the multi-dimensional interaction strategy it outputs better meet the actual needs of the patient group, the large language model needs to be pre-trained.

[0113] In some embodiments, the generative large language model adopts the first large language model, which is a lightweight large language model, such as the Llama3-8B large language model. Since the parameters of the lightweight large language model are smaller in magnitude, the computing power requirements for its pre-training process are easier to meet, and the real-time nature of the data generated by the lightweight large language model is higher, meeting the real-time requirements for interacting with patients in lower limb rehabilitation training. The steps of pre-training the first large language model include:

[0114] Step S100, chain-of-thought training, i.e., performing the first fine-tuning training on the chain-of-thought (CoT) of the first language model, by inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, to train the first language model to gradually think and reason about the content of lower limb rehabilitation training and the ability to communicate content and attitude;

[0115] Step S200, interactive text training, collects aging-friendly corpus and sports training motivational corpus to perform secondary fine-tuning training on the interactive text content output by the first language model, so that the interactive text output by the model is more suitable for the lower limb rehabilitation training scenario and meets the patient's expectations.

[0116] Furthermore, the specific steps of step S100, thinking chain training, include:

[0117] Step S110: Collection of existing cases.

[0118] Collect existing decision-making cases and decision-making ideas of licensed professional rehabilitation physicians when performing lower limb training on patients. Optionally, an existing case i should at least record: (1) the patient's condition s i , such as injury time, lesion area, movement assessment, etc.; (2) initial artificial rehabilitation training program, such as movement training method, training speed, training duration, etc., which needs to be approximately equivalent to the initial rehabilitation training program p adapted to the lower limb rehabilitation robot 33 in the embodiment of the present disclosure. i , and combined with the characteristics of the lower limb rehabilitation robot 33 i The normalized representation is a text sequence: "Training content: passive / assisted / active, resistance level, training speed". The equivalent process can refer to the existing experience in the rehabilitation field and the advice of professional rehabilitation physicians; (3) The jth real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and reasons for adjustment ij , where strategy a will be adjusted ij The text normalization of the training content is: "Training content adjustment: passive / assisted / active, resistance level, training speed; communication content adjustment: communication attitude, communication text", further, the communication attitude can be "comfort" or "encouragement" two labels; adjustment reason r ijIt may include changes in the patient's mental state, physical state, and further include professional theoretical knowledge that the physician refers to when making adjustments.

[0119] Step S120: Retrieve professional knowledge using the second language model, and expand the real-time adjustment reasons for the training content and communication content in the existing cases through reasoning analysis. ij .

[0120] By designing prompts, the second language model is prompted to search a wide range of professional rehabilitation theories and disease psychology theories. Based on the zero-sample thinking chain training method, instructions are added to the prompt to make the second language model think step by step, so that the second language model can combine the retrieved knowledge to adjust the strategy a in the existing case i. ij The second language model is based on the retrieved knowledge and the original professional rehabilitation physician’s experience and knowledge to adjust the strategy a ij Conduct step-by-step analytical reasoning to supplement and improve the reasons for adjustment ij To improve training efficiency, compared with the single adjustment strategy a ij Corresponding adjustment reason r ij The number of characters in the inference text is limited to not more than the first maximum length Lr1 = 256. Optionally, the second largest language model used in this embodiment is ChatGPT-4o, so as to fully combine external knowledge and existing case information for efficient reasoning and supplementation, so as to expand the sample size required for offline pre-training of the first largest language model.

[0121] Step S130: data standardization.

[0122] Case i all text s i 、p i 、a ij 、r ij Generate text combination w according to the specified format i , and use all case text combinations to generate the first text sequence W = {w1,w2,…,w i ,…,w N}, since the first text sequence can form a causal chain relationship, the w in the first text sequence W i It can be expressed in the format (s i -->p i ,j:r ij -->a ij ).

[0123] Step S140: first model fine-tuning training.

[0124] The standardized first text sequence W is organized into the first fine-tuning data set, and the first fine-tuning training is performed on the fully open source Llama3-8B large language model. Specifically, the first fine-tuning process adopts the adapter fine-tuning method, that is, a lightweight adapter module is introduced in each layer of the pre-trained first large language model, and the weight parameters of these adapter modules are adjusted through training, while keeping the original parameters of the first large language model unchanged, thereby achieving efficient fine-tuning of the first large language model. Adapter Tuning not only greatly reduces the amount of parameters and resource consumption required in the training process, but also has good task adaptability and scalability, and is suitable for fine-tuning requirements of multi-tasks or specific scenarios. The fine-tuning process of the Llama3-8B large language model is implemented based on the framework of Hugging Face Transformers. HuggingFace provides a convenient TrainerAPI and efficient fine-tuning tools that can quickly integrate Adapter Tuning and optimize the training process. Further, Hugging Face's Accelerate library can be combined to achieve distributed training and automatic mixed precision training, thereby improving fine-tuning efficiency and large language model performance. In a specific embodiment, the fine-tuning process sets the learning rate to 3e-4, the number of training epochs to 20, and uses the first cross entropy loss function L LOSS1 The parameter θ1 of the adapter module is optimized as the objective function, as follows:

[0125]

[0126] In the formula, |r i | represents the total number of training adjustment strategy-adjustment reason groups for case i. Since different cases may have different numbers of training adjustment strategies-adjustment reasons, in order to ensure that the weights of the loss function affected by different cases in the thinking chain training are constant and not affected by the number of adjustment strategies-adjustment reasons, |r is used before the second term of the formula. i |Perform a normalization; P(p i |p i ,s i ; θ1) represents the first language model after training input patient condition s i The output is then used to obtain the initial rehabilitation training program p i The probability of quantifying the ability to infer the initial training plan from the patient's condition; P(a i,j |p i ,r i,≤j ,s i ,a i,<jθ1) is used to quantify the first language model after training. Based on the current patient condition, the initial rehabilitation training program, the rehabilitation training adjustment strategy and adjustment reason before the jth adjustment, and the jth adjustment reason, the jth adjustment strategy a is inferred. i,j The probability of i,≤j represents the adjustment reason for the jth and previous rehabilitation training, a i,<j Represents the rehabilitation training adjustment strategy before the jth adjustment.

[0127] Furthermore, the specific steps of step S200, interactive text training, include:

[0128] Step S210: Corpus collection.

[0129] Aged-friendly corpus and rehabilitation exercise training motivational corpus are collected from the Internet by means of web crawlers. Optionally, the aged-friendly corpus involves the language expressions commonly used by the elderly in daily life, health management related to the elderly, psychological intervention, social support and other contents; its sources include: online articles, forums and comments related to elderly health and psychological support, case corpus of communication with the elderly in social support and companion services, and corpus describing the elderly in rehabilitation medical institutions and professional literature. The rehabilitation exercise training motivational corpus is mainly motivational corpus to enhance the confidence and motivation of patients in rehabilitation training; its sources include: the compilation of motivational sentences in professional rehabilitation training documents; the real dialogue records between doctors and patients in excellent rehabilitation cases; and literature and corpus related to motivation in sports psychology research.

[0130] Step S220: data cleaning and preprocessing.

[0131] The collected corpus is cleaned, including removing duplicate texts, correcting punctuation errors and grammatical confusion, and correcting typos and irregularities in the corpus through automated tools. In some embodiments, a python library named pycorrector can be used to implement Chinese text error correction. After the data cleaning is completed, the corpus is divided into different categories according to semantic functions, such as motivational sentences, comforting sentences, task description sentences, daily communication sentences, etc., so that the subsequent first language model can output appropriate text to patients in rehabilitation training scenarios. To ensure the uniformity and efficient processing of the corpus, the maximum number of characters of a single corpus segment is limited to Lr2 = 128 to avoid interference and influence of ultra-long corpus on the rehabilitation training process. Finally, the obtained corpus is formatted into a second text sequence V = {v1, v2,…, v l ,…,v M}, where v l Represents a single corpus fragment, and there is no causal relationship between the corpus fragments.

[0132] Step S230: secondary model fine-tuning training.

[0133] The standardized second text sequence V is organized into a second fine-tuning dataset, and the fine-tuned Llama3-8B large language model obtained in step S140 is subjected to secondary fine-tuning training. For the specific fine-tuning training process, see step S140. The learning rate used in the secondary fine-tuning process can be 2e-5, the epoch is set to 20, and the second cross entropy loss function L is introduced. LOSS2 As the objective function of training, the model parameters θ2 are fine-tuned as follows:

[0134]

[0135] In the formula, P(v l ; θ2) represents the first language model prediction corpus segment v after secondary fine-tuning l probability.

[0136] After the second fine-tuning, the first language model has a significant performance improvement in the formulation of human-computer interaction strategies in rehabilitation training scenarios. It can fully analyze the patient's multi-dimensional psychological and physical conditions and make interactive decisions in the way of thinking of an excellent rehabilitation physician, including appropriately adjusting the content and difficulty level of the next stage of rehabilitation training tasks to stimulate patients' confidence and motivation to actively participate in rehabilitation training, and timely generating some comforting or encouraging texts to directly provide emotional support to users. By simulating the real doctor-patient and family-patient interpersonal interaction processes in the process of human-computer interaction, it fully integrates and utilizes the positive effects of interpersonal emotional interaction in motor learning, improves patients' mental health level, and amplifies the benefits of rehabilitation training tasks on patients' physical condition.

[0137] In some embodiments, prompt is designed so that the fine-tuned first language model can generate standardized output after receiving standardized input, and the specific output content includes:

[0138] "--Training content--

[0139] Training duration: %c minutes;

[0140] Training mode: active / assisted / passive;

[0141] Training impedance: %dN.

[0142] --Virtual human interaction content--

[0143] Interaction text: %e;

[0144] Interaction attitude: comfort / motivation.

[0145] --Reason for adjustment--

[0146] %f."

[0147] In the above text information, “%c” and “%d” are specific numerical values, and “%e” and “%f” are specific text contents.

[0148] In some embodiments, see Figure 5 , relying on the execution unit 3 to complete the implementation of the multi-dimensional interaction strategy. Specifically, the multi-dimensional interaction strategy generated by the interaction strategy generation module 22 is embodied in two aspects: sports training content and interpersonal interaction simulation content. Among them, the interpersonal interaction simulation is realized by the speaker 31 and the visual interaction device 32, which are responsible for training motivation speech generation and virtual character image generation respectively, forming an auditory and visual interpersonal interaction simulation, which enhances the patient's positive emotions for rehabilitation by providing the patient with the positive effects of interpersonal interaction, thereby enhancing their motivation and involvement.

[0149] In some embodiments, the interactive text generated by the interactive strategy generation module 22 can be further converted into audio, and interacted with the patient through the speaker 31, and the audio can be given the timbre characteristics of the patient's close person and emotional characteristics with positive effects such as comfort and encouragement. The specific steps may include: collecting a speech sample of the target speaker, using a text-to-speech (TTS) model (such as a deep learning model VALL-EX) to extract the timbre characteristics of the collected speech sample, including key parameters such as fundamental frequency, formant and speech duration, and establish a feature vector of the target timbre. Then, the TTS model is used to realize the conversion from text to audio, and the target timbre characteristics extracted in the first step are embedded in the speech generated by the TTS model. Through feature fusion, the audio output by the TTS model has the timbre characteristics of the target speaker; then, the generative adversarial network is used to extract the emotional characteristics of the existing sample voice library with positive motivation and comfort emotions. The sample voice library is composed of a total of 400 audio data generated by recording the audio of 20 adults of different genders and timbre when reading comfort and motivation corpus. Based on the extracted emotional features, the fundamental frequency value, Mel spectrum and energy distribution of the audio output by the TTS model are further adjusted so that the audio output by the speaker 31 presents a soft and positive tone, conveying emotions of comfort and encouragement.

[0150] In some embodiments, the visual interaction device 32 may be a traditional display screen, or an augmented reality or virtual reality device to provide a more immersive visual experience. The generation of a virtual character image based on the visual interaction device 32 includes three aspects: character appearance, facial movements corresponding to the character's voice, and facial expressions.

[0151] In some embodiments, the appearance of the virtual character can come from a preset database or can be customized based on the patient's preferences. The specific implementation method is: the visual interaction device 32 creates a preset 3D character database through the Unity rendering engine, and the character image materials come from various open source or paid platforms MakeHuman, Adobe Mixamo, Renderpeople. Furthermore, the preset virtual character image parameters can be adjusted according to the patient's needs or the desired image can be remodeled and customized. Optionally, the customization process can use a 3D scanning device (such as Artec Eva or Structure Sensor) to collect the real facial and body features of the target person, generate a high-precision 3D mesh model, and then use the Unity engine combined with the scan data to realize the customization of the virtual character model.

[0152] In some embodiments, the facial movements of the virtual character adapted to the speech are mainly to simulate the mouth shape of the character when speaking. The specific implementation method includes: using a preprocessing algorithm (such as MFCCs or Wav2Vec 2.0) to extract features from the speech signal of the training excitation speech generated by the speaker 31, capturing the key rhythm, syllable intensity and speech speed information in the speech; then dividing the extracted speech features into time steps to ensure the synchronization of the audio and subsequent facial animation sequences on the time axis. Based on the VOCA (Voice Operated Character Animation) model, the audio signal is encoded into a low-dimensional speech embedding vector to ensure the correspondence with the dynamic mouth shape and facial movements, and then the extracted speech features are mapped to the expression parameter space of the FLAME (Faces Learned with an Articulated Model and Expressions) model to control the key point movements of the mouth (such as opening, closing, round lips, etc.). The FLAME model then coordinates the reference expression parameters and speech features to generate a facial action mesh that changes dynamically frame by frame, and finally uses the Unity rendering engine to render the generated dynamic facial action mesh into facial movements that adapt to the appearance of the virtual character.

[0153] In some embodiments, the facial expression of the virtual character is determined by the interactive attitude output by the interactive strategy generation module 22, which is mainly the two emotions of comfort and encouragement. On the basis of speech-driven lip generation, the Expression Layer tool of the FLAME model is used to superimpose the 50-dimensional expression control parameters of the virtual character on the basis of the lip adjustment parameters. The expression control parameters can be obtained by fine-tuning the expression parameter library of the FLAME model.

[0154] In some embodiments, the rehabilitation training content generated by the interactive strategy generation module 22 is adjusted by the lower limb rehabilitation robot 33. The training content includes training duration, training mode, and training impedance. With respect to the training impedance, the lower limb rehabilitation robot 33 dynamically adjusts the output torque to change the degree of muscle effort required by the patient when performing the task, so as to achieve the impedance control target at different stages of lower limb rehabilitation. As mentioned above, the interactive strategy generation module 22 establishes reasonable training content by considering the patient's physical and psychological conditions. For example, when the user shows any of the following conditions: "low emotional state", "high mental load", "high peripheral fatigue", and "significant decrease in task performance values", the training intensity and difficulty can be reduced, thereby enhancing the patient's sense of accomplishment and self-efficacy in the process of lower limb rehabilitation training, and stimulating their confidence and motivation for rehabilitation.

[0155] In summary, the lower limb motor rehabilitation system with emotional interaction function provided by the embodiment of the present disclosure collects multi-source physiological and behavioral signals of patients, determines the psychological and physical state of patients, and makes real-time decisions based on this with the help of a large language model to adjust the interaction strategy of the rehabilitation system. The system introduces the positive effects of interpersonal emotional interaction into the traditional human-computer interaction process of sports rehabilitation training dominated by rehabilitation robots, from personalized adjustment of task content to activate patients' confidence and motivation for rehabilitation, to establishing a virtual character image to provide patients with visual feedback and voice motivation for social emotional interaction and support. Therefore, the embodiment of the present disclosure can optimize the functions and effects of traditional rehabilitation systems, improve the effectiveness of rehabilitation training, help patients improve their positive emotions and their compliance and participation in training, and conduct high-quality lower limb motor rehabilitation under the guidance of positive emotions, shorten the rehabilitation cycle, and improve the quality of life of patients.

[0156] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0157] Although embodiments of the present disclosure have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A lower limb exercise rehabilitation system based on emotional interaction, characterized in that: include: A sensing unit, used to obtain physiological and behavioral signals of the patient; A central controller, including a data processing module and an interaction strategy generation module; The data processing module is used to judge the patient's physical state and psychological state in real time according to the physiological signals and behavioral signals acquired by the sensor unit, and output the judgment result as the input of the interaction strategy generation module. The interaction strategy generation module simulates the professional knowledge and thinking mode of the rehabilitation physician based on the large language model to understand the patient's current state and make integrated reasoning decisions, and output a multi-dimensional interaction strategy with the patient, including real-time generated visual and voice interaction content and lower limb rehabilitation training content; Actuator units, including speakers, visual interaction devices, and lower limb rehabilitation robots; The speaker is used to play voice with timbre characteristics familiar to patients and positive emotional style according to the voice interaction content generated by the interaction strategy generation module during training; the visual interaction device is used to generate a virtual character according to the visual interaction content generated by the interaction strategy generation module, as one of the carriers for simulating interpersonal emotional interaction and support, and to make the virtual character produce corresponding facial movements and expressions in conjunction with the voice generated by the speaker; the lower limb rehabilitation robot is used to guide patients to complete lower limb rehabilitation training movements according to the lower limb rehabilitation training content generated by the interaction strategy generation module.

2. The lower limb exercise rehabilitation system according to claim 1, characterized in that: The lower limb rehabilitation robot can guide the patient to complete a full week of cycling exercise, and has multiple training modes and difficulty levels, including passive, assisted, and active. The change in difficulty level is achieved by adjusting the training speed during passive and assisted mode training, and the resistance value during active mode training; the lower limb rehabilitation robot obtains information from its internal pedal pressure sensor and motor encoder, and feeds back information including its own mechanism movement speed and interaction force to the sensing unit as basic data for quantitative evaluation of the patient's training task performance.

3. The lower limb exercise rehabilitation system according to claim 1, characterized in that: The physiological signals acquired by the sensing unit include electroencephalogram, electromyography, electrocardiogram and electrodermal signals, and the behavioral signals acquired include facial expressions, speech, eye movements and motor task performance information; The process of the data processing module determining the patient's psychological state in real time based on the physiological and behavioral signals acquired by the sensing unit includes: preprocessing and preliminary feature extraction of each type of acquired signal; Subsequently, a deep learning-based classification model is used to perform deep feature extraction on the various extracted preliminary features to obtain the patient's psychological state.

4. The lower limb exercise rehabilitation system according to claim 3, characterized in that: Dividing the patient's psychological state into emotional state and mental load level, dividing the emotional state into three levels: positive, neutral and negative, and dividing the mental load level into three levels: high, medium and low; The deep learning-based classification model is used to perform deep feature extraction on various extracted preliminary features to obtain the patient's psychological state, specifically including: Each type of extracted preliminary features is taken as a modality, and a classification model based on deep learning is used to perform single-modal feature fusion and classification to obtain the preliminary classification results of the patient's psychological state; wherein, each modality uses a corresponding classification model, and during the training process of each classification model, a labeled data set is used for supervised learning. The labels are assessed based on the subjective emotion scale and the perceived stress scale. Through the subjects' scores, the scores are divided into high, medium, and low training labels according to the scale standards to classify emotions and mental loads; Based on the rules, all the preliminary psychological state classification results of single modalities are subjected to multimodal decision fusion, wherein the rule is to dynamically adjust the weight of each modality in the decision fusion according to its individual classification performance, and the modality with better classification performance is assigned a higher weight. The weighted result is obtained by linearly weighted summation of the classification labels of each modality, and the weighted result is normalized to obtain the patient's final psychological state classification result as the discrimination result of the patient's psychological state.

5. The lower limb exercise rehabilitation system according to claim 1, characterized in that: The patient's physical condition is divided into peripheral fatigue level and rehabilitation task performance; The peripheral fatigue degree is used to directly reflect the patient's muscle state and is divided into three levels: high, medium and low. The data processing module obtains peripheral fatigue information according to the patient's electromyographic signal, specifically including: performing preprocessing operations on the acquired electromyographic signal data to obtain the peak electromyographic amplitude A of the current period, and comparing the peak electromyographic amplitude A with the peak electromyographic amplitude A within the first 10 seconds of training. max When A / A max When it is between 80% and 100%, it is considered as low fatigue state. max Between 60% and 80%, it is considered a moderate fatigue state. max When it is 60% or below, it is considered as a high fatigue state; The performance of the rehabilitation task is related to the type and content of the lower limb rehabilitation training task. The data processing module uses movement speed and movement symmetry as evaluation indicators. The movement speed is characterized by the circumferential speed of the lower limb rehabilitation robot itself collected by the lower limb rehabilitation robot, and the movement symmetry is characterized by the average positive pressure ratio of the left and right legs applied to the pedals during the current training cycle phase collected by the lower limb rehabilitation robot. The data processing module integrates the identified physical and psychological states of the patient and outputs them in a unified natural language format: "--Psychological state-- Emotional state: positive / neutral / negative; Mental workload: high / medium / low; --Physical Condition-- Peripheral fatigue: high / medium / low; Task performance: movement speed %a, movement symmetry %b" Among them, "%a" and "%b" are specific values.

6. The lower limb exercise rehabilitation system according to claim 1, characterized in that: The large language model used by the interaction strategy generation module is a pre-trained first large language model, and the pre-training process of the first large language model includes: Step S100, thinking chain training: performing the first fine-tuning training on the thinking chain of the first language model, by inputting typical cases of sports rehabilitation and theories of sports rehabilitation and disease psychology, training the first language model to gradually think and reason about the content of lower limb rehabilitation training and the ability to communicate content and attitude, wherein the lower limb rehabilitation training content includes training duration, training mode and task difficulty; Step S200, interactive text training: collect aging-friendly corpus and sports training motivational corpus to perform secondary fine-tuning training on the text content output by the first language model after the first fine-tuning, so that the interactive text output by it is more suitable for the lower limb rehabilitation training scenario and meets the patient's expectations.

7. The lower limb exercise rehabilitation system according to claim 6, characterized in that: Step S100 specifically includes: Step S110: Collection of existing cases: Collect existing decision-making cases and decision-making ideas of licensed professional rehabilitation physicians when performing lower limb training on patients. An existing case should at least record: (1) the patient's condition i , including injury time, lesion area and movement assessment; (2) an initial artificial rehabilitation training program, which is equivalent to an initial rehabilitation training program p adapted for the lower limb rehabilitation robot i and the initial rehabilitation training program p i Normalized representation as a text sequence: "Training content: passive / assisted / active, resistance level, training speed"; (3) The jth real-time adjustment strategy for training content and communication content during lower limb rehabilitation training a ij and reasons for adjustment ij , where strategy a will be adjusted ij The text normalization is represented as: "Training content adjustment: passive / assisted / active, resistance level, training speed; Communication content adjustment: communication attitude, communication text", where the communication attitude has two labels: "comfort" or "encouragement"; Adjustment reason r ij Includes changes in the patient's mental state, physical state, and / or professional theoretical knowledge that the physician refers to when making adjustments; Step S120: Retrieve professional knowledge using the second language model, and expand the real-time adjustment reasons for the training content and communication content in the existing cases through reasoning analysis. ij : By designing prompt words, the second largest language model is prompted to retrieve professional rehabilitation theory and disease psychology theory. Based on the zero-sample thinking chain training method, instructions for the second largest language model to think step by step are added to the prompt words, so that the second largest language model can combine the retrieved knowledge to adjust the strategy a in the existing case i. ij Conduct step-by-step analytical reasoning to supplement and improve the reasons for adjustment ij , and the single adjustment strategy a ij Corresponding adjustment reason r ij The number of characters in the inference text is limited to not exceed the first maximum length Lr1; Step S130: Data standardization: All text s of case i i 、p i 、a ij 、r ij Generate text combination w according to the specified format i , and use all case text combinations to generate the first text sequence W = {w1,w2,…,w i ,…,w N }; Step S140: First fine-tuning training: The standardized first text sequence W is organized into a first fine-tuning dataset, and the first language model is trained using the adapter fine-tuning method and the Hugging Face Transformers framework; wherein the first cross entropy loss function L is used. LOSS1 The adapter parameter θ1 is optimized as the objective function: In the formula, |r i | represents the total number of training adjustment strategy-adjustment reason groups for case i; P(p i |p i ,s i ; θ1) represents the first language model input patient condition s after the first training i The output is then used to obtain the initial rehabilitation training program p i The probability of quantifying the ability to infer the initial training plan from the patient's condition; P(a i,j |p i ,r i,≤j ,s i ,a i,<j θ1) is used to quantify the first language model after the first training. Based on the current patient condition, the initial rehabilitation training program, the rehabilitation training adjustment strategy and adjustment reason before the jth adjustment, and the adjustment reason for the jth adjustment, the jth adjustment strategy a is inferred. i,j The probability of i,≤j represents the adjustment reason for the jth and previous rehabilitation training, a i,<j Represents the rehabilitation training adjustment strategy before the jth adjustment.

8. The lower limb exercise rehabilitation system according to claim 6, characterized in that: Step S200 specifically includes: Step S210: Corpus collection: Collect aging-friendly corpus and rehabilitation exercise training motivational corpus from the Internet through web crawlers; Step S220: Data cleaning and preprocessing: The collected corpus is cleaned, including removing duplicate texts, correcting punctuation errors and grammatical confusion, and correcting typos and language irregularities in the corpus; the cleaned corpus is divided into different categories according to semantic functions, including motivational sentences, comforting sentences, task description sentences and daily communication sentences; the maximum number of characters in a single corpus segment is limited to not more than the second maximum length Lr2; all the obtained corpus segments are formatted into a second text sequence V = {v1, v2, ..., v l ,…,v M },v l Represents a single corpus fragment; Step S230: Second fine-tuning training The standardized second text sequence V is organized into the second fine-tuning dataset, and the first language model after the first fine-tuning is trained using the adapter fine-tuning method and the framework of Hugging Face Transformers; wherein the second cross entropy loss function L is used LOSS2 As the objective function to fine-tune the model parameters θ2: In the formula, P(v l ; θ2) represents the probability of the first language model predicting the corpus segment vl after the second fine-tuning.

9. The lower limb exercise rehabilitation system according to claim 6, characterized in that: By designing prompt words, the pre-trained first language model generates standardized outputs, including: "--Training content-- Training duration: %c minutes; Training mode: active / assisted / passive; Training impedance: %dN. --Virtual human interaction content-- Voice interaction text: %e; Interaction attitude: comfort / motivation. --Reason for adjustment-- %f。” Among them, "%c" and "%d" are specific values, and "%e" and "%f" are specific text contents.

10. The lower limb exercise rehabilitation system according to claim 9, characterized in that: The speech interaction text and interaction attitude generated by the interaction strategy generation module are converted into speech with timbre characteristics and positive emotional style familiar to the patient, and the speaker is used to interact with the patient. The specific steps include: The speech samples of the target speaker are collected, and the timbre features of the collected speech samples are extracted using a text-to-speech model, including fundamental frequency, resonance peaks and speech duration, to establish a feature vector of the target timbre; the text-to-speech model is used to realize text-to-audio conversion, and the extracted target timbre feature vector is embedded into the speech generated by the text-to-speech model, and through feature fusion, the audio output by the text-to-speech model has the timbre features of the target speaker; the generative adversarial network is used to extract the emotional features of the sample speech library with positive motivation and comfort emotions; based on the extracted emotional features, the fundamental frequency value, Mel spectrum and energy distribution of the audio output by the text-to-speech model are adjusted, so that the audio output by the speaker has the timbre features and positive emotional style familiar to the patient.

11. The lower limb exercise rehabilitation system according to claim 9, characterized in that: The visual interaction device is a display screen, a virtual reality device or an augmented reality device; the virtual character image generated by the visual interaction device includes three aspects: the character's appearance, the character's facial movements corresponding to the voice, and the character's facial expressions, wherein: The character appearance is from a preset database or customized based on the patient's preferences; The facial movements corresponding to the speech of the character are generated in the following manner: extracting features from the speech signal output by the speaker to capture key rhythm, syllable strength and speech rate information in the speech; then dividing the extracted speech features into time steps to ensure synchronization of the audio and subsequent facial animation sequences on the time axis; encoding the audio signal into a low-dimensional speech embedding vector to ensure the correspondence with the dynamic mouth shape and facial expression, and mapping the extracted speech features to the action parameter space of the FLAME model to control the key point movement of the mouth, and the FLAME model collaboratively references the facial expression parameters and speech features to generate a facial movement mesh that changes dynamically frame by frame; finally, using a rendering engine to render the generated facial movement mesh into facial movements that match the appearance of the virtual character; The facial expression of the character is determined by the interaction attitude output by the interaction strategy generation module. On the basis of adjusting the lip movement of the extracted speech features, the expression layer tool of the FLAME model is used to superimpose 50-dimensional expression control parameters of the virtual character on the basis of the lip adjustment parameters. The expression control parameters are obtained by fine-tuning the expression parameter library of the FLAME model.

12. The lower limb exercise rehabilitation system according to claim 9, characterized in that: The lower limb rehabilitation robot adjusts its own state in real time according to the training duration, training mode, and training impedance in the training content generated by the interaction strategy generation module.

Citation Information

Patent Citations

  • Exercise rehabilitation robot interactive control method based on emotion perception

    CN104287747A

  • Lower limb rehabilitation training robot system based on virtual reality

    CN107049702A

  • Optimization simulation method and device for satellite communication system

    CN115243296A

  • Video monitoring method and system based on multi-scene recognition and voice interaction

    CN117749995A

  • Multi-modal speech rehabilitation training system based on virtual reality

    CN117789982A

Cited By

  • Rehabilitation training adjustment method and system based on multi-modal feature fusion and storage medium

    CN122091086A