Multi-modal man-machine cooperative interaction method for exercise rehabilitation and cognitive evaluation of old people
By employing a multimodal human-computer collaborative interaction method, combined with offline speech recognition and finite state machine control, the problems of interaction failure and data discontinuity in intelligent rehabilitation assessment for the elderly have been solved. This has enabled efficient, safe, and personalized rehabilitation assessment for the elderly, improving assessment compliance and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTH CHINA UNIV OF TECH
- Filing Date
- 2025-12-24
- Publication Date
- 2026-05-01
AI Technical Summary
Existing intelligent rehabilitation assessment systems for the elderly suffer from problems such as high interaction failure rate, poor user experience, interrupted assessment process, discontinuous data, inability to coordinate and control multimodal tasks, and insufficient data security, making it difficult to meet the home rehabilitation needs of the elderly.
Employing a multimodal human-computer collaborative interaction approach, this system combines offline speech recognition, dual-channel speech broadcasting, finite state machine process control, compliance tracking closed-loop management, and deep learning-optimized reasoning to achieve unified interaction and intelligent guidance of speech, vision, and movement. Through speech feature extraction, semantic reconstruction, finite state machine control, and deep learning models, a personalized rehabilitation assessment system is constructed.
It improves assessment compliance and rehabilitation outcomes among the elderly, achieves continuity and error tolerance in the assessment process, enhances data accuracy and security, provides personalized rehabilitation plan adjustments, and supports intelligent rehabilitation assessment for the elderly in home settings.
Smart Images

Figure CN121964042A_ABST
Abstract
Description
Multimodal human-computer collaborative interaction method for motor rehabilitation and cognitive assessment of the elderly Technical Field
[0001] This invention relates to the fields of intelligent medical care and human-computer interaction technology, specifically to a multimodal human-computer collaborative interaction method and system for elderly motor rehabilitation and cognitive assessment. Background Technology
[0002] As my country's society ages, movement disorders, mild cognitive impairment (MCI), and chronic disease rehabilitation management have become core issues urgently needing to be addressed for the elderly population. In traditional clinical practice, cognitive assessment and movement rehabilitation training primarily rely on face-to-face guidance from healthcare professionals, which suffers from typical bottlenecks such as "high levels of manual intervention, poor assessment continuity, high cognitive load, low compliance, and fragmented data." Furthermore, the process is complex, manpower-dependent, inefficient, and difficult to promote. The elderly population itself exhibits characteristics such as cognitive decline, decreased sensory abilities, deterioration of visual and auditory channels, and easy distractibility, making it difficult to consistently, stably, and accurately complete the process using traditional single-modal interactive assessment models in real-world applications.
[0003] Although existing intelligent rehabilitation assessment systems have achieved digitization and visualization to a certain extent, they still have the following technical bottlenecks: (1) The single interaction mode limits the user group. Elderly users generally have hearing, vision or language expression impairments. Most existing systems rely on screen operation or voice commands, which makes it difficult to take into account the elderly with different ability levels, resulting in a high failure rate of interaction and poor user experience.
[0004] (2) The assessment process lacks continuity and fault tolerance mechanisms. Traditional speech recognition or question-and-answer systems cannot automatically resume when the user pauses, repeats, or makes an operational error, resulting in interruption of the assessment process and loss of results. Elderly people are easily distracted or make accidental touches during the assessment process, and systems lacking automatic rollback and state recovery mechanisms cannot guarantee data integrity.
[0005] (3) Multimodal tasks cannot be controlled collaboratively. Motion and cognitive assessment often involve multimodal tasks such as voice question answering, drawing tasks, option confirmation, and image display. Existing solutions generally implement these functions separately, with voice recognition, TTS broadcasting, interface switching, drawing input, and result display being independent of each other, lacking a unified timing control logic and mutual exclusion mechanism.
[0006] (4) Insufficient data security and offline availability. Many systems rely excessively on cloud connections, and once the network is unstable, voice recognition and data interaction cannot be carried out. For elderly users in home or community settings, this mode poses both privacy risks and seriously affects system availability.
[0007] In the field of medical rehabilitation, clinical scales (such as the MoCA, MMSE, and Berg Balance Scale) are standardized tools for assessing cognitive and motor function in older adults. However, fully digitizing and automating the process of these scales requires precise coordination in areas such as speech recognition, interruption control during playback, task flow state machine management, interactive drawing, and result synchronization. Traditional questionnaire systems or voice assistants can only complete single-turn dialogues and cannot support complex assessment processes involving multiple modalities and tasks.
[0008] Advances in artificial intelligence, multimodal interaction, and localized reasoning technologies have provided new possibilities for solving the aforementioned problems. By introducing offline speech recognition (ASR), dual-channel text-to-speech (TTS), state machine-based workflow control, and digital human motion linkage, a multimodal human-machine collaborative rehabilitation assessment system for the elderly can be realized. This system ensures both the naturalness and safety of the interaction while significantly improving the continuity and fault tolerance of the assessment process.
[0009] A questionnaire survey of 1,025 elderly people in some communities in Guangzhou City on their health status, chronic diseases, types of functional impairments, and rehabilitation service needs revealed that the elderly population generally suffers from functional problems such as hypertension, joint diseases, asthma, and limited mobility. More than 60% of them require community and home-based rehabilitation services, and they also have a significant need for aging-in-home modifications and long-term health management. This illustrates the current reality of huge rehabilitation needs, insufficient service supply, and low adherence to rehabilitation behaviors among the elderly in home settings, providing an important basis for the development of intelligent rehabilitation systems (Luo Xiaoyuan, Yang Qi, Huang Lili, et al. Preliminary survey and analysis of the rehabilitation needs of the elderly in communities and homes in Guangzhou City [J]. Chinese Journal of Rehabilitation Medicine, 2022, 37(04):515-518.).
[0010] However, existing technologies only conduct statistical analysis from the demand side, without addressing any automated rehabilitation assessments or providing technical pathways for automating scale assessments and rehabilitation training in home settings through voice recognition, visual perception, motion analysis, or task flow scheduling. Current methods generally rely on human guidance, subjective questioning, or static questionnaires, lacking real-time collection and analysis of objective indicators such as the quality of elderly individuals' motor performance, language response speed, attention levels, and emotional changes. The assessment process is prone to interruption and difficult to resume, resulting in a lack of continuity and interpretability. Furthermore, traditional electronic questionnaires or voice assistants cannot support the complex, multi-step, multimodal processes of scales such as MMSE, MoCA, and Berg, lacking automatic process control, error recovery, and intelligent prompting capabilities, and cannot dynamically adjust the assessment pace and difficulty based on the elderly individual's performance. These shortcomings make it difficult to standardize and scale up home-based rehabilitation assessments, failing to meet the elderly's needs for safe, convenient, and low-intensity intelligent rehabilitation assessment methods.
[0011] Therefore, there is an urgent need for a system that can achieve adaptive process control under multimodal input conditions, has offline fault tolerance, and can provide intelligent guidance and analysis for elderly rehabilitation assessment, so as to break through the limitations of existing manual assessment methods and build a new paradigm of intelligent rehabilitation assessment that integrates "voice-vision-operation". Summary of the Invention
[0012] The purpose of this invention is to address the core pain points in the current process of exercise rehabilitation and cognitive assessment for the elderly, such as "excessive manual intervention, poor assessment continuity, high cognitive load, low compliance, and fragmented data," by providing a multimodal human-computer collaborative interaction system for elderly exercise rehabilitation and cognitive assessment. This invention improves the continuity and accuracy of the assessment process through technologies such as offline speech recognition, dual-channel voice broadcasting, finite state machine process control, compliance tracking closed-loop management, and deep learning optimized inference. It replaces traditional manual assessment with intelligent voice and visual interaction, improving the fault tolerance and system adaptability of task completion. It enhances assessment compliance among the elderly, and combined with personalized intelligent reminders and behavior tracking, effectively integrates the completion of rehabilitation tasks into the patient's daily activities, improving rehabilitation outcomes. Through real-time collection and analysis of multimodal data, it ensures the accuracy and scientific validity of assessment results, while providing doctors with real-time assessment feedback, supporting personalized adjustments to rehabilitation plans, and ultimately achieving intelligent, precise, and data-driven closed-loop management of elderly exercise rehabilitation and cognitive assessment.
[0013] The present invention is achieved by at least one of the following technical solutions.
[0014] A multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment includes the following steps: S1, preprocessing and extracting speech features from the collected speech signals; S2, decoding the extracted speech features, performing semantic reconstruction and intent reasoning on the decoded text, generating corresponding execution actions and text sequences, and converting the text sequences into speech signals with time sequences, while controlling digital human facial expressions based on the text sequences; S3, using a reentrant multimodal assessment process control model based on finite state machines to perform unified time management, state determination, and process scheduling control for multimodal interaction tasks; S4, generating a comprehensive compliance score based on the degree of completion of rehabilitation tasks according to plan; S5, adjusting the difficulty of tasks based on the compliance score and predicting the compliance trend of the elderly in the future; S6, using a combination of deep learning models and knowledge graphs to recommend personalized intervention measures to users.
[0015] Further, step S1 includes the following steps: S11, performing preprocessing operations on the acquired raw speech, the preprocessed speech is segmented into continuous frames, and the continuity between frames is maintained by a sliding window; S12, extracting speech features of the preprocessed audio signal using a deep learning-based adaptive acoustic modeling method.
[0016] Further, in step S2, the semantic reconstruction and intent reasoning of the decoded text includes the following steps: First, the decoded text is structured using a bidirectional long short-term memory network to extract key semantic entities, and the output results serve as the node basis for subsequent semantic graph construction; then, an intent graph is constructed based on the extracted key semantic entities, and semantic reasoning is performed through a graph neural network to generate execution actions and related text sequences; wherein, the intent graph consists of multiple semantic nodes and associated edges, where nodes represent task intent or action goals, and edges represent logical relationships and contextual dependencies between different tasks or states. By running a graph neural network on the graph for feature propagation and aggregation, the implicit intent of the user can be identified and path prediction can be performed in different contexts.
[0017] Furthermore, in step S2, controlling the digital human's facial expressions based on text sequences involves extracting emotional cues from acoustic feature parameters to generate corresponding facial expressions and emotional states. Combined with real-time data from rehabilitation assessments, the system judges the user's action completion rate and emotional state. When the user accurately completes the action, the digital human provides positive feedback with a smiling expression and encouraging tone. When the user's action deviates or pauses, the digital human provides prompts with a guiding tone to encourage the user to readjust their action posture.
[0018] Further, step S3 includes the following steps: First, based on the multimodal input vectors from speech recognition, visual detection, motion sensing, and touch interaction, a finite set of states containing multiple task nodes is established, with each node corresponding to a specific assessment task or interaction scenario; the current state is determined based on changes in the multimodal input vectors, and the output behavior to be executed under the current state and input is determined; when an abnormal situation is detected, the system automatically reverts to the most recent valid state through a reentrant mechanism, restores the assessment progress, and re-executes the corresponding task node; after restoration, new multimodal input vectors are received, and the state judgment is updated based on the user's real-time behavior, dynamically adjusting the execution path of subsequent tasks, thereby achieving continuous execution and intelligent scheduling of elderly motor rehabilitation and cognitive assessment.
[0019] Furthermore, the compliance score in step S4 The definition is as follows:
[0020] in, The timestamp represents the current moment. , , , These represent the weighting coefficients of each modality in the overall compliance score, used to adjust the contribution ratio of speech, vision, action, and system feedback information to the final score; Indicates input for speech modality The feature evaluation function is used to quantify speech recognition and responsiveness, including voice command recognition accuracy, response timeliness, and semantic coherence. Indicates the input for visual modality The feature evaluation function is used to measure visual attention and recognition ability, including eye focus, image recognition accuracy and interface operation stability. Indicates input for action modality The feature evaluation function is used to evaluate motor response and body coordination, including postural completion, reaction speed and range of motion; Indicates the system interaction state The comprehensive feedback function is used to reflect the system interaction performance during task execution, such as response latency, error recovery rate, and process continuity.
[0021] Furthermore, in step S4, compliance trend is predicted using a deep learning prediction model. The deep learning prediction model uses a recurrent neural network (RNN) or a long short-term memory network (LSTM) to predict the compliance trend of older adults over a future period of time.
[0022] The system for implementing the multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment includes: a voice acquisition and processing module for acquiring voice signals and extracting voice features; a voice recognition module for recognizing voice features and generating corresponding actions; a voice broadcasting and digital human expression module for outputting voice and adjusting the digital human's facial expressions and voice rhythm; a compliance scoring module for generating a comprehensive compliance score, adjusting the difficulty of tasks or reminder strategies based on the compliance score, and predicting the compliance trend of the elderly in the future; and a personalized intervention module for recommending personalized intervention measures to users.
[0023] The system comprises the following modules: a voice acquisition module for acquiring user voice; a preprocessing and feature extraction module for extracting preprocessed voice features and using a user-level fine-tuning module to predict the probability distribution of each voice feature; a semantic reconstruction and intent reasoning module for generating corresponding actions and text sequences based on voice features; a decision control module for matching the intent obtained from the actions with predefined task nodes using cosine similarity calculation to trigger the operation; a voice generation and playback module for restoring time-series voice signals from text sequences; and a digital human expression control module for extracting emotional cues from the acoustic feature parameters output by the voice generation and playback module, generating corresponding expressions and emotional behaviors for the digital human, achieving context-adaptive immersion. Immersive interaction; a time synchronization module for controlling voice and digital human facial expressions to synchronize voice and vision; a finite state machine-based assessment process control module for unified time management, state determination, and process scheduling control of multimodal interactive tasks in elderly motor rehabilitation and cognitive assessment processes to ensure continuity, fault tolerance, and traceability in complex interactive environments; a compliance scoring module for generating a comprehensive compliance score based on the degree of completion of rehabilitation tasks as planned, increasing the reminder frequency for a particular modality if compliance is low; a task adjustment module for adjusting task difficulty or reminder strategies based on compliance scores and predicting compliance trends in the elderly over a future period; and a personalization module for recommending personalized interventions to users.
[0024] A computer device according to the present invention includes a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, which, when executed by the processor, causes the processor to implement the method described herein.
[0025] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor implements the method described herein.
[0026] Compared with existing technical solutions, the technical solution of this invention has the following beneficial effects: To solve the above problems, this invention introduces key technologies such as multimodal fusion perception, graph neural networks, knowledge graphs, and deep learning to construct an intelligent assessment system and method for medical rehabilitation scenarios. The system supports the collaborative acquisition of information from channels such as action, language, gaze, and facial expressions, integrates historical behavioral features, and utilizes knowledge graphs and deep models for automatic reasoning and personalized intervention matching. This realizes the transformation of rehabilitation assessment from "outcome-oriented" to "process-tracking," and constructs a closed-loop human-machine collaborative intervention paradigm throughout the entire process.
[0027] 1. Significantly Improved Speech Recognition Accuracy and Naturalness of Interaction. Compared to traditional rehabilitation systems that rely on fixed voice commands or cloud-based speech recognition, this invention introduces an adaptive acoustic modeling and semantic reconstruction mechanism. This mechanism automatically adjusts recognition parameters based on the speech rate, tone, and pronunciation characteristics of different elderly users, achieving high-precision speech understanding even offline. The system supports noise suppression, semantic error correction, and context awareness. It no longer relies solely on keyword matching but can understand natural language expressions, reducing false triggers and comprehension biases. This allows the elderly to complete assessment questions and tasks through natural speech communication. The system maintains continuous and smooth human-computer dialogue even in environments without network or with weak signals, significantly improving usability and comfort for the elderly.
[0028] 2. Immersive Guidance Experience Achieved Through Voice Broadcasting and Digital Human Facial Expression Linkage. Traditional voice prompts are typically mechanical, lacking visual feedback and emotional guidance, making it difficult to maintain the attention of elderly users. The dual-channel broadcasting control method proposed in this invention synchronously maps the intensity, tone, and rhythm of the voice with the facial expressions and body movements of the digital human. The system can automatically switch between emotional states such as "waiting," "explaining," and "encouraging" based on the task's semantics, achieving dual-modal fusion feedback of voice and vision. Elderly users not only "understand" but also "see clearly and with a sense of familiarity," significantly reducing cognitive burden and psychological alienation, and improving task completion rate and emotional companionship.
[0029] 3. Finite State Machine-based process control significantly improves system fault tolerance and continuity. Compared to traditional script-based evaluation processes, this invention establishes a reentrant multimodal task management mechanism using a finite state machine (FSM), abstracting the entire evaluation process into six transitionable states: "start-guide-response-confirmation-result-feedback". The system can automatically identify anomalies and backtrack to the previous valid state in case of voice interruption, drawing errors, or timeouts, ensuring that the process does not crash or lose tasks. This reentrant mechanism enables the evaluation process to "pause-resume-continue", which is especially suitable for elderly users with easily distracted attention. In complex evaluation scenarios, this mechanism effectively reduces the rate of human intervention, improves task continuity and result integrity, and provides a stable data foundation for subsequent intelligent analysis.
[0030] 4. Compliance Tracking and Personalized Reminders Construct a Closed-Loop Behavioral Management System. Compared to previous assessment systems that only recorded static results of "whether completion was achieved," this invention comprehensively calculates the user compliance index using multimodal signals such as voice response, motion detection, and touch behavior. It analyzes the elderly's task completion rate, response delay, and emotional stability in real time and automatically generates personalized reminder strategies. The system can determine whether the user is distracted or fatigued based on behavioral trends, dynamically adjusting the frequency and content of reminders to achieve a closed-loop management process of "reminder—confirmation—recording—feedback." This mechanism enables the system to track behavior and proactively intervene, shifting from "outcome-oriented" to "process management," significantly improving the continuity, scientific rigor, and individual adaptability of rehabilitation assessments.
[0031] 5. The fusion of knowledge graphs and deep learning enables intelligent reasoning and adaptive intervention. Unlike existing assessment systems that rely on fixed scale scores, this invention constructs a knowledge graph containing the relationships of "symptoms—tasks—scales—interventions—outcomes," and achieves intelligent reasoning through graph neural networks. The system can automatically identify potential motor or cognitive impairment types based on multimodal assessment data and retrieve corresponding individualized training programs from the graph, forming a self-learning loop of "assessment—diagnosis—intervention—feedback." This method not only outputs highly interpretable assessment results but also automatically optimizes recommendation strategies over long-term user use, allowing intervention plans to evolve over time and adjust to individual differences, achieving a truly adaptive rehabilitation agent. Attached Figure Description
[0032] Figure 1 is a flowchart of a multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment in an embodiment.
[0033] Figure 2 is a flowchart of offline speech recognition and semantic understanding in the embodiment.
[0034] Figure 3 is a flowchart of the interaction between the TTS model and the digital human's facial expressions in the embodiment. Detailed Implementation
[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] This embodiment of a multimodal human-computer collaborative interaction system for elderly motor rehabilitation and cognitive assessment includes: a voice acquisition module for acquiring user voice.
[0037] In this embodiment, to achieve efficient voice acquisition for elderly users, the present invention acquires voice signals through a mobile APP. The voice acquisition module is integrated within the APP and records user voice in real time by calling the device's high-sensitivity microphone array. This module includes: (1) a microphone input management unit, used to call the device's microphone array and acquire the original audio signal; (2) an audio buffer and unified sampling unit, used to maintain signal continuity and sampling stability during the acquisition process; and (3) a local preprocessing unit, which sends the buffered audio into the preprocessing flow to perform noise reduction, endpoint detection, and other operations. The APP, as the running carrier, is responsible for scheduling the above-mentioned voice acquisition module, thereby ensuring the real-time performance and reliability of the voice acquisition link.
[0038] The voice acquisition module in this embodiment employs a high-quality microphone array to enhance effective acquisition capabilities under complex background noise. The app's built-in audio recording and data compression mechanisms further reduce transmission latency and improve overall response speed.
[0039] The preprocessing and feature extraction module is used to extract preprocessed speech features and to predict the probability distribution of each speech feature through a user-level fine-tuning module.
[0040] The semantic reconstruction and intent reasoning module is used to generate corresponding execution actions and corresponding text sequences based on speech features.
[0041] The decision control module uses cosine similarity calculation to match the intent derived from the action with predefined task nodes to trigger the operation, i.e., calling the application programming interface (API) or functional module to execute the action. The system-generated action logically determines the type of operation the system should perform (e.g., playing a rehabilitation guidance video, activating the action recognition module, recording assessment results, etc.). This action is then passed to the decision control module, which completes the operation triggering process by calling the specific API or functional module to execute the action. In other words, generating the action belongs to the task decision layer, while triggering the operation belongs to the system execution layer; the two form a logical sequence. For example, when the user says "next step," the system generates the action "enter the next task node," and the decision control module triggers the playback of the next stage of rehabilitation training video and updates the interface display.
[0042] The speech generation and broadcasting module is used to restore a time-series speech signal from a text sequence.
[0043] The digital human facial expression control module is used to extract emotional cues based on the acoustic feature parameters output by the speech generation and broadcasting module, generate corresponding facial expressions and emotional expressions for the digital human, and realize context-adaptive immersive interaction.
[0044] The time synchronization module is used to control voice and digital human facial expressions, so that voice and vision are synchronized.
[0045] The assessment process control module based on finite state machines is used to perform unified time management, state determination and process scheduling control of multimodal interactive tasks in the process of elderly motor rehabilitation and cognitive assessment, so as to ensure continuity, fault tolerance and traceability in complex interactive environments.
[0046] The compliance scoring module generates a comprehensive compliance score based on the degree to which rehabilitation tasks are completed as planned. If compliance in a certain modality is low, the reminder frequency for that modality is increased.
[0047] The task adjustment module is used to adjust the difficulty of tasks or reminder strategies based on compliance scores and to predict compliance trends in older adults over a future period.
[0048] The personalization module is used to recommend personalized interventions to users.
[0049] As shown in Figures 1-3, this embodiment aims to achieve intelligent and automated tasks in the process of exercise rehabilitation and cognitive assessment for the elderly. Using a multimodal human-computer collaborative interaction system as the core carrier, it proposes a multimodal human-computer collaborative interaction method for elderly exercise rehabilitation and cognitive assessment, including the following steps: S1. Offline speech recognition and understanding based on adaptive acoustic modeling and semantic reconstruction preprocesses and extracts speech features from the collected speech signals, including the following steps: S11. Preprocessing operations are performed on the collected raw speech, including: spectral subtraction or a deep learning-based noise reduction algorithm (DNN-based NR) to improve the signal-to-noise ratio; pre-emphasis processing to enhance high-frequency information, making MFCC features more sensitive to the blurred pronunciation of the elderly; speech endpoint detection (VAD) to accurately identify the start and end points of speech and remove silent areas; amplitude normalization to eliminate differences in sound acquisition from different devices. These preprocessing steps make the subsequently extracted MFCC features more stable and improve the robustness of the model in elderly speech scenarios. The preprocessed speech is segmented into continuous frames, and the continuity between frames is maintained by a sliding window.
[0050] S12. A deep learning-based adaptive acoustic modeling method is used to extract preprocessed speech features.
[0051] Traditional acoustic models primarily rely on handcrafted features such as MFCC (Multiple-Chronic Vocalization) patterns. However, these models are prone to recognition errors when considering the weak vocal intensity, unstable rhythm, and irregular pauses of elderly individuals. To address this, the acoustic model of this invention employs a deep learning-based adaptive acoustic modeling method to perform deep modeling of MFCC features. This acoustic model not only establishes a general acoustic representation but also performs transfer learning and feature reparameterization on a small number of personalized vocal samples from the user, adapting to the user's unique vocal intensity, rhythm, pause patterns, and fundamental frequency variations, thereby further improving recognition stability and accuracy.
[0052] The deep learning-based adaptive acoustic modeling method includes the following steps: S121, in the feature extraction stage, each frame of audio signal after preprocessing is subjected to short-time Fourier transform (STFT) to generate a spectrum, and then the energy features that conform to the hearing characteristics of the human ear are extracted by the Mel filter bank. Finally, the Mel frequency cepstral coefficients (MFCC) are calculated to form a feature vector sequence reflecting the time-frequency structure of speech. The MFCC features can effectively encode the core input features of speech phoneme structure, pitch pattern and formant information.
[0053] Setting the first The audio signal of the frame is The frequency is obtained by short-time Fourier transform (STFT). Spectral characteristics Then, the Mel frequency cepstral coefficients (MFCCs) are calculated using the Mel filter:
[0054] in, Indicates the first after preprocessing Frame audio signal; For signal The complex spectrum obtained by performing a short-time Fourier transform; The amplitude spectrum is the result of the short-time Fourier transform; Mel is the filter bank function that maps the spectrum to the Mel frequency scale. For MFCC feature extraction function; It is the first The MFCC feature vector of the frame is used as input for the subsequent acoustic model.
[0055] S122. An adaptive training model is used to predict the probability distribution of each speech feature: Traditional acoustic models are trained using fixed speech data, while this invention addresses this issue through an adaptive training model. Specifically, personalized adjustments are made based on each user's speech characteristics (such as pitch, speech rate, and pronunciation style). The parameters of the adaptive training model are fine-tuned using a small amount of user speech data (5 to 10 samples), enabling the model to adapt to the audio characteristics of different elderly users.
[0056] The adaptive training model employs a Transformer-based network architecture, which captures long-range dependencies in speech sequences through a multi-head self-attention mechanism, thereby significantly improving the understanding of complex audio signals. The specific network computation is as follows:
[0057] in, These represent the query, key, and value matrices, respectively. It is a normalization factor that ensures the computational stability of multi-head self-attention. The output of the adaptively trained model is the predicted probability distribution for each speech feature. , The output label predicted by the model is the text content corresponding to the speech. The input speech feature sequence to the adaptive acoustic modeling model consists of multiple frame-level MFCC feature vectors. This probability distribution serves as the input to the CTC decoding network, which maximizes the path probability to obtain the optimal label sequence π. This sequence is then used as the input to the semantic reconstruction and intent reasoning module in step S2 to generate the final structured text content.
[0058] S2. The extracted speech features of each frame are fed into the neural network for decoding. The decoded text is then used for semantic reconstruction and intent reasoning to generate the corresponding execution action.
[0059] S21. After completing acoustic modeling, the system enters the speech signal recognition stage: For each extracted frame of speech features, it is fed into a neural network for decoding. This decoding process is based on Connectionist Temporal Classification (CTC decoding network). The CTC decoding network can automatically handle unaligned input-output sequences, ensuring accurate recognition even without precise time annotations. Specifically, the CTC loss function is calculated as follows:
[0060] in, Given the input speech feature sequence, Output the target sequence (such as "raise hand", "start", etc.). Represents the set of all possible paths. It is a set of paths mapped to the target label sequence. This is the CTC loss value. Indicates that given input Path under conditions The probability is calculated from the frame-level classification probability matrix output by the CTC decoding network. The CTC decoding network is a mature existing speech recognition decoding structure that uses beam search technology to estimate the optimal path probability. This invention outputs the optimal label sequence through the CTC decoding network to achieve accurate recognition of speech commands. The output probability distribution is used for path prediction, thereby achieving accurate decoding of the audio signal.
[0061] S22. After completing speech recognition, semantic reconstruction and intent reasoning are performed on the recognized text. Because elderly users may use non-standard or ambiguous natural language (e.g., "What should I do?", "Continue"), traditional semantic analysis methods often fail to accurately understand their true intent. To address this problem, this invention combines Named Entity Recognition (NER) with a graph-based semantic reasoning model to perform deep semantic parsing of the recognized text, thereby accurately identifying user intent and generating executable action commands for the system.
[0062] First, a named entity recognition model constructed using a bidirectional long short-term memory network (Bi-LSTM) is used to perform structured analysis on the recognized text, extracting key semantic entities. The output results serve as the node basis for subsequent semantic graph construction. These key semantic entities include drug names, rehabilitation programs, body parts, and time points.
[0063] Subsequently, an IntentGraph (Metapath-guided Heterogeneous Graph Neural Network for Intent Recommendation) is constructed based on the extracted key semantic entities, and semantic reasoning is performed through a Graph Neural Network (GNN) (Point-GNN: Graph Neural Network for 3D Object Detection in a PointCloud) to generate execution actions and related text sequences.
[0064] The intent graph consists of multiple semantic nodes and associated edges. Nodes represent task intents or action goals (e.g., "start exercising," "perform shoulder exercises," "complete assessment"), while edges represent logical relationships and contextual dependencies between different tasks or states. By running a graph neural network on this graph for feature propagation and aggregation, the system can identify the user's implicit intent and predict paths in different contexts. For example, when a user issues the voice command "I'm ready to start exercising" (this voice command originates from the voice signal collected in step S1 and converted into a text command by the speech recognition module), the system locates the corresponding node "start exercising" in the intent graph and determines the subsequent task as "perform shoulder exercises—monitor progress—complete task" through contextual reasoning. The entire process is adaptively adjusted and optimized according to the context.
[0065] S23. After entity recognition and semantic reasoning are completed, the system uses cosine similarity calculation to match the reasoned intent task nodes. Upon successful matching, the system automatically triggers the corresponding operation command, such as "start task," "pause task," or "confirm input," and provides feedback through voice announcements and on-screen prompts to ensure that the user is aware of the system status in a timely manner.
[0066] Throughout the interaction, the system supports a feedback adjustment mechanism. When it detects that a user's instruction is not correctly recognized or the intent is unclear, it automatically enters "confirmation mode," requesting the user to confirm or repeat the instruction through voice prompts or touch interaction to ensure the continuity and accuracy of the task flow. To further improve the system's fault tolerance, this invention designs a rollback and semantic clarification mechanism (Schlobach S. Debugging and semantic clarification by pinpointing European Semantic Web Conference. Berlin, Heidelberg: Springer Berlin Heidelberg, 2005: 226-240.). By combining context and task status to determine the user's true intent, it avoids task interruptions or errors caused by misrecognition. When recognition fails, the system will automatically roll back to the previous operation and provide voice prompts for correction; if the system detects semantic uncertainty or multiple interpretations, it will activate the semantic clarification mechanism, proactively issuing a rhetorical question to the user, such as "Do you want to continue the current task or start a new exercise?", to ensure the correctness of task execution and the stability of the evaluation process.
[0067] S24. A deep neural network speech synthesis model (TTS, Text-To-Speech) is used to convert the input text into a phoneme sequence, then the corresponding spectral features are generated through an acoustic model, and finally the audio signal is restored through a vocoder to achieve natural and fluent speech output.
[0068] Deep neural network speech synthesis models (TTS, Text-To-Speech) take input text sequences as input. This is the foundation. The input text sequence comes from the output of the semantic reconstruction and intent reasoning module in step S2, namely the identified structured instruction text (such as "start task", "perform shoulder movement", "confirm result", etc.). For the first The model generates a natural speech signal through phoneme decomposition, prosody prediction, spectrum generation, and vocoder reconstruction, as follows: First, through the phoneme mapping function... Convert to phoneme vector sequence :
[0069] in, Indicates the first The phoneme features of a frame, including parameters such as place of articulation, duration, speech rate, and acoustic type, are used for subsequent speech feature generation. Each phoneme feature... It includes parameters such as articulation point, duration, speech rate, and initial consonant category, which are used for subsequent feature generation.
[0070] Predict the energy of each frame using a prosodic modulation network. , baseband and speech rate parameters Generate Mel spectrum
[0071]
[0072] in , ∈ (1~4) is a learnable weight matrix, ensuring that the speech rhythm is natural and conforms to human hearing. Parameters Represents a time frame. Indicates frequency index.
[0073] Subsequently, the vocoder utilizes convolution kernels Restore the Mel spectrum to the speech waveform signal. Vocoder function. The definition is as follows:
[0074] in, Indicates the weights of the vocoder convolution kernel; The number of sampling points generated for the speech waveform; Represents time frame Corresponding to the The function calculates the Mel energy values of each frequency component. It reconstructs the Mel spectrum into a speech signal through convolutional weighting and filtering, ultimately generating a speech signal with a time sequence. And record the timestamp sequence , This indicates the output time of the last audio frame, providing a reference for facial expression synchronization.
[0075] S25. To achieve real-time correspondence between voice and facial expressions, after voice generation, this invention extracts key feature parameters from the audio through an acoustic feature mapping module and converts them into control signals to drive the digital human's facial expressions. The aim is to achieve multimodal unification of "voice-emotion-facial expression," enabling the system to possess emotional human-computer interaction capabilities, thereby enhancing the elderly user's immersion and compliance during the rehabilitation process. The digital human in this invention is not simply used to broadcast system operation results, but rather serves as the core interactive subject for rehabilitation guidance, undertaking the functions of "motivation, guidance, feedback, and companionship."
[0076] During system operation, the digital human extracts emotional cues based on the acoustic feature parameters (such as fundamental frequency, sound intensity, speech rate, energy changes, etc.) output by the TTS model, and uses them to generate corresponding facial expressions and emotional expressions. When the user completes the action accurately, the digital human provides positive feedback with a smiling expression and encouraging tone. When the user's action deviates or pauses, the digital human provides prompts with a caring and guiding tone to encourage the user to readjust their action posture.
[0077] The core of this design lies in the fact that the digital human's voice broadcasts and facial expressions do not passively respond to operational commands, but rather actively respond to the user's performance and emotional state during the rehabilitation process, forming a natural emotional interaction loop. Through the dynamic coordination of voice tone, facial expressions, and body posture, the digital human achieves emotional resonance and motivational guidance for elderly users, thereby effectively improving participation, continuity, and rehabilitation outcomes in rehabilitation training.
[0078] The acoustic feature mapping module extracts the following features from the speech waveform in real time: sound intensity: ; Indicates speech intensity (volume). This represents the amplitude of the j-th sampling point in the speech frame. This indicates the number of sampling points contained in each analysis frame.
[0079] Pitch: Fundamental frequency estimated using the autocorrelation method Prosodic gradient: Characterizing intonation fluctuations; pause markers : This represents a pause frame; these parameters form the acoustic feature vector. : .
[0080] Input the acoustic feature vector into the facial expression mapping function:
[0081] in For expression-driven parameter vectors, The weights of the action units (AU). Let be the nonlinear mapping function from acoustic signal to motion drive. For emotion regulation, This represents the number of sub-channels in the emotion mapping.
[0082] When the speech energy increases, the movement of the mouth and eyebrows is automatically amplified; when the speech rate decreases or the tone softens, the digital human displays a soothing expression (such as a slight nod or smile).
[0083] Voice content is often accompanied by different semantic contexts (such as encouragement, confirmation, and prompts). Therefore, this invention introduces an emotion intensity modulation model, which predicts the emotion vector corresponding to the current semantic meaning through a voice emotion decoding network.
[0084] in, For emotion regulation, The semantic context vector is generated by the BERT-based semantic analysis module [Zhang Wen, Zhang Jiantong, Guo Yushan. Sentiment analysis of online medical reviews based on BERT and dual-channel semantic collaboration [J]. Journal of Medical Informatics, 2024, 45(11):30-35.]. This is the Sigmoid normalization function, used to compress the range of sentiment values; These are the emotion mapping parameters. Based on the current semantic intent (such as "encouragement," "guidance," "questioning") and the evaluation process status (such as action execution, waiting for a response), a preset emotion template is dynamically matched: for example, when the user completes an action, the "encouragement" intent is recognized and the user is in a "feedback" state, which triggers the "encouragement" template, causing the digital human to smile and nod, while the voice increases in pitch, speeds up, and emphasizes positive words; when issuing instructions, it switches to "guidance," with a focused expression, slower speech, and clear pauses; when waiting for the user's response, it adopts "waiting for a response," with a gentle expression, soft voice, and a listening posture, thereby achieving context-adaptive immersive interaction.
[0085] The synchronization of voice and facial expression channels is crucial for the interactive experience. To achieve high-precision audio-visual consistency, this invention designs a time synchronization equation. :
[0086] in Indicates the timestamp of the visual channel frame sequence. This is the synchronization balance coefficient. Indicates the first in the voice channel The actual playback timestamp of the frame (in seconds); Indicates the first in the visual channel The original rendering timestamp of the frame; Represents visual frames The corresponding facial expression feature vector; This indicates the visual frames being searched within the candidate window.
[0087] In this embodiment, the time difference is calculated every 10ms and the facial expression rendering speed is adaptively adjusted to keep speech and vision consistent at the perception level.
[0088] When synchronization deviation is detected When this occurs, compensation is triggered: ,in This is a smoothing factor (ranging from 0.3 to 0.6). This mechanism can automatically correct time deviations in the event of GPU rendering jitter or audio latency, ensuring that the digital human's facial expressions and lip movements accurately correspond to the speech.
[0089]
[0090] in Indicates the actual time deviation between the current speech frame and the matching visual frame (unit: seconds); Indicates the synchronization error threshold; Represents the original rendering timestamp of the visual frame; Indicates the actual playback timestamp of the audio frame; Indicates the corrected visual frame timestamp; The smoothing factor is used to control the magnitude of time adjustment.
[0091] In addition, to ensure a natural transition, the emotion decay model is automatically invoked when the voice broadcast ends or the user pauses their response:
[0092] in: This represents a vector of emotional intensity generated during speech; This represents the decayed emotion vector, used to drive smooth changes in facial expressions; Indicates the time since the end of the speech (in seconds); This indicates the rate of emotional decay and controls the speed at which facial expressions recover. The emotion decay rate is used to control the changes in facial expressions as the digital human gradually calms down from a high emotional state, making the dialogue process more like a real person's performance.
[0093] S3. The reentrant multimodal assessment process control model based on finite state machines provides unified time management, state determination and process scheduling control for multimodal interactive tasks in the process of elderly motor rehabilitation and cognitive assessment, so as to ensure continuity, fault tolerance and traceability in complex interactive environments.
[0094] This model abstracts the evaluation process into a finite set of states consisting of several task nodes. Each state corresponds to an evaluation or interaction subtask (such as voice guidance, response input, motion detection, result confirmation, etc.). During operation, the system determines the current task state in real time based on multimodal signals from speech recognition, visual detection, motion sensors, and touch input, and automatically switches to the next logical node according to the state transition function. When abnormal situations such as user pauses, accidental touches, distraction, or sensor signal interruption are detected, or when speech recognition fails, the finite state machine (FSM) automatically backtracks to the previous valid state and re-executes the task through a reentrant mechanism, thereby avoiding process interruption or data loss. Simultaneously, the model can update the state transition probability matrix based on the task execution history, achieving adaptive optimization and dynamic adjustment of the evaluation process, enabling the system to have "pause-resume-continue" interactive capabilities. Through a reentrant multimodal assessment process control model based on finite state machines, the system can achieve unified and coordinated control of voice, vision, action and touch interaction under multimodal input conditions, ensuring that elderly users obtain a stable, smooth, continuous and recoverable intelligent interactive experience during the assessment and rehabilitation process.
[0095] A finite state machine (FSM) is defined as an ordered 7-tuple:
[0096] in: It is a finite set of system states, representing different task nodes in the evaluation process (such as "voice guidance", "response input", "image detection", "result confirmation" etc.). Indicates the first step in the evaluation process One task node; The set of input events includes voice commands, sensor signals, touch operations, and camera recognition results; The output set represents the feedback actions performed by the system (such as voice broadcasting, interface switching, digital human facial expression changes, etc.). This is the state transition function, used to determine the next state of the system based on the input signal; This is the output function, used to define the execution behavior in each state; Indicates the initial state (system standby or ready); This represents the set of termination states (evaluation task completed or abnormal exit). Represents the first in the set of input events A specific input event; Indicates the first element in the output set. A specific output behavior.
[0097] The operation of FSM is described by the following core equations:
[0098] That is, every moment The system determines the current state. and input Determine the next state and the corresponding output. Indicates the system at time... The current status, such as "voice guidance" or "response input"; Indicates at time Triggered input events, such as voice commands, motion detection results, or touch operations; It is a state transition function, which determines the next state based on the current state and the input. It is the output function, which represents the output behavior that the system should perform given the current state and input. Such as voice broadcasting or interface switching.
[0099] To ensure the system can achieve highly robust state determination in complex interactive environments, this invention introduces a multimodal input vector. :
[0100] in: Features of speech signals (phoneme sequence, semantic embedding, command confidence); Visual features (pose key point coordinates, facial expression encoding, drawing stroke trajectory); For motion and acceleration sensor data; This represents the touch and interface interaction state. After time synchronization and noise filtering, the multimodal feature fusion function is defined as:
[0101] in , The dynamic weighting coefficients for each mode ( ), each sub-function These are modal feature extractors, This represents the fused multimodal feature vector. State classification is achieved using a combination of a Gaussian Mixture Model (GMM) and a deep attention mechanism.
[0102] in: Indicates at time The system is in a state The posterior probability; Candidate state; This is the fused multimodal feature vector; The number of Gaussian components; For the first The mixing weights of Gaussian components; For the first The probability density function of Gaussian components, where For state The The mean vector of Gaussian components, Let its covariance matrix be ; This is the output of the attention mechanism, used to enhance key task features. This is the attention query vector.
[0103] The first term captures the probability distribution of continuous actions, and the second term uses an attention vector. Strengthen the characteristics of key tasks. This represents the fusion coefficient. The final state determination result is:
[0104] This approach enables FSM to recognize multimodal spatiotemporal features, dynamically distinguishing differences in the action patterns and speech response rhythms of elderly users. A state transition matrix is defined. Used to characterize the transition probability between any two states:
[0105] in: Indicates the system at time 10:00 The current status, such as "voice guidance" or "response input"; This represents any initial state (i.e., the current state), used to define the transition probability from that state.
[0106] The state transition matrix satisfies the normalization condition:
[0107] The state transition matrix is automatically updated from the historical evaluation log using maximum likelihood estimation:
[0108] in This represents the number of times an event is counted. The matrix is fine-tuned online at regular intervals to achieve adaptive optimization of the state transition probabilities. Furthermore, to capture the state evolution trend over time, a gated recurrent unit (GRU) is embedded in the state transition function.
[0109] in To update the gate, control the degree to which new information is updated; , These are the weight matrices from the input and hidden states to the hidden layer, respectively; To reset the door, decide whether to ignore the previous hidden state; For hidden layer bias terms; Use the Sigmoid activation function; , Both are weight matrices for the update gate and the reset gate. Both are weight matrices for the update gate and the reset gate; This is a hidden state used to store the task context. Through the introduction of the GRU unit, the FSM can not only identify the input state of a single frame, but also understand the dynamic trends of user behavior.
[0110] In multimodal environments, sensing signals may be subject to interference or frame loss. This invention addresses this by defining a time-gating function. Implement reentrancy mechanism:
[0111] The timestamp represents the current moment; Indicates the timestamp of the last valid input event; The maximum allowed time interval threshold (in seconds); This is a time-gated function where the time difference between the current moment and the last valid input is less than [a certain value]. Output 1 if the current process is allowed to continue (reentrancy is allowed), otherwise output 0 to trigger process backtracking or interruption.
[0112] When at the threshold time If a valid input is detected, the system will automatically return to the state it was in during the last interruption. If the timeout occurs, it will enter recovery mode. The voice prompt is then retried. Simultaneously, to ensure the integrity of the evaluation process, a state backtracking function is introduced:
[0113] This indicates the state after recovery at the current moment; This is a state backtracking function used to calculate the state to be restored based on the current state and the sequence of historical states; It is a historical state sequence, recording each state the system went through during the evaluation process and its corresponding timestamp; This is an index variable, representing the position relative to the current moment in the historical sequence. The event number closest in time; For the first A timestamp of a historical event.
[0114] S4. Generate a comprehensive compliance score based on the degree to which rehabilitation tasks are completed as planned.
[0115] Compliance refers to the degree to which older adults follow the system's pre-set rehabilitation task plan and effectively perform various assessment or training tasks. This system automatically generates structured rehabilitation task sequences based on standardized clinical scales (such as the MoCA and Berg Balance Scale), forming an individualized interactive task flow that includes ordered nodes such as voice question-and-answer, movement instructions, and graphical drawing. This method designs a compliance assessment function. It combines multimodal data (such as voice, vision, and movement) to generate a comprehensive compliance score that reflects the behavior of older adults during the rehabilitation process.
[0116] Compliance score The definition is as follows:
[0117] These are the weighting coefficients for the speech modality, reflecting the importance of the speech response in compliance assessment. These are the weighting coefficients for visual features; These are the weighting coefficients for the action modality, reflecting the contribution of action execution quality to compliance. The weighting coefficients for system interaction and feedback; This is a state feature extraction function that represents the semantic encoding of the task stage or interaction type (such as "instruction execution" or "waiting for response") corresponding to the current system state, used to capture behavioral consistency in the task flow. Features of speech signals (phoneme sequence, semantic embedding, command confidence); Visual features (pose key point coordinates, facial expression encoding, drawing stroke trajectory); For motion and acceleration sensor data, This indicates system interaction and feedback. By dynamically updating scores, it provides real-time feedback on compliance, ensuring that the assessment results accurately reflect the elderly person's task performance.
[0118] Through multimodal data fusion analysis, the evaluation results generated from each modality (speech, action, etc.) are utilized. This method tracks the task performance of elderly individuals in real time. To achieve more accurate assessment, it employs a graph neural network (GNN) to model task performance. The GNN connects each task node... It establishes connections with other nodes (such as the sequential relationships between tasks, the associations between modalities, etc.) and learns and infers the state changes of each node through model learning. The specific task execution modeling process can be represented as:
[0119] in: It is the first A representation of each task, indicating the state of the task. It is an adjacency matrix between tasks, representing the dependencies between tasks. It is a node The degree, It is a node The degree, It is a weight matrix. It is an activation function. In this way, this method can capture the temporal relationships between tasks and the dependencies between modalities, thereby accurately predicting the execution status of each task for the elderly.
[0120] Based on compliance score The system dynamically adjusts its reminder strategy. If compliance with a particular task is low, the system will provide personalized reminders through various means (voice, image, digital human). For example, when compliance with voice recognition and understanding is low, the system will increase the frequency of voice reminders and combine them with interactive text and image prompts to highlight the importance of the task.
[0121] Personalized reminder strategies are dynamically adjusted based on a FeedbackLoopControl algorithm. Set reminders. Function. As time The intensity of a constant reminder is expressed as:
[0122] in: These are weighting coefficients, representing the influence of each factor on the intensity of the reminder; For compliance scoring, if the score falls below a set threshold, the reminder intensity will increase. and These represent compliance with the speech and action modalities, respectively. If compliance with a certain modality is low, the frequency of reminders for that modality will be increased.
[0123] S5. Adjust the difficulty of the task or reminder strategies based on the compliance score, and predict the compliance trend of older adults in the future.
[0124] Closed-loop management is one of the core functions of this invention. When the system detects low compliance among elderly individuals in certain tasks, it automatically triggers a closed-loop control mechanism to adjust the task difficulty or reminder strategies to ensure timely task completion. The specific closed-loop adjustment process is implemented through a control strategy function. Implementation, described as:
[0125] in: To provide real-time intensity alerts, For real-time compliance scoring; The rate of change in compliance represents the dynamic adjustment magnitude of the system; and To adjust weights and control the intensity of policy changes, this function automatically adjusts task difficulty or reminder methods by monitoring compliance in real time, thereby ensuring the continuity and stability of task execution.
[0126] By combining historical adherence data with real-time feedback, adherence prediction is made through a deep learning prediction model. The deep learning prediction model uses a recurrent neural network (RNN) or a long short-term memory network (LSTM) to predict the adherence trend of older adults over a period of time and to make early interventions based on the prediction results.
[0127] Compliance prediction function Indicates at time Prediction of compliance scores:
[0128] in: To predict compliance scores; the LSTM model uses historical compliance data. The system is trained to output future compliance trends. Using this model, the system can make adjustments in advance based on predicted compliance trends, ensuring improved task completion.
[0129] S6. Uses deep learning models and knowledge graphs to recommend personalized intervention measures to users.
[0130] Knowledge graphs are built upon multi-source data (such as sensor data, assessment questionnaire data, etc.), where nodes include various diseases, symptoms, assessment tools, treatment plans, etc., and nodes are connected by relationships (such as "cause" and "treatment"). Each triple... Represents a relationship. and For entities, This represents the relationship between entities. The formula is as follows:
[0131] in, It is a collection of entities. It is a set of relations.
[0132] Knowledge graph reasoning can deduce possible diseases from known symptoms and recommend corresponding treatment plans based on those diseases. This reasoning process, based on graph databases and reasoning algorithms, can effectively utilize historical data to make real-time decisions.
[0133] Deep learning models are used to process various types of data, such as gait data, speech data, and cognitive test results. Commonly used deep learning models include Convolutional Neural Networks (CNNs) for image data processing and Long Short-Term Memory Networks (LSTMs) for time-series data processing. Combining the advantages of these deep learning models, we can use deep learning models for health assessment:
[0134] in, The output of the model (health assessment results), For the input multimodal data, These are the parameters of the model.
[0135] Combining the output of deep learning models with knowledge graphs enables more complex reasoning. For example, after a deep learning model assesses "gait instability," the knowledge graph can infer possible causes (such as "muscle weakness") based on this symptom and, combined with past data, derive suitable intervention plans.
[0136] By reasoning through neighboring nodes in a graph neural network (GNN), the knowledge in the knowledge graph is dynamically updated, thereby recommending personalized intervention measures to users.
[0137] The fusion of deep learning models and knowledge graphs can provide personalized intervention programs for the elderly. These programs include exercise, diet, and cognitive training. The system generates corresponding intervention tasks based on the user's health status (such as gait instability, cognitive decline, etc.) and dynamically adjusts the intervention program based on the task completion status.
[0138] The effectiveness of the intervention is evaluated using real-time feedback data, and the intervention strategy is adjusted as needed. The mathematical model is as follows:
[0139] in, Indicates the intervention plan, This indicates the results of the health assessment. For historical health data, This is a function for generating interventions.
[0140] As the intervention progresses, the system continuously monitors the elderly person's reactions and task completion, adjusting the intervention strategy based on feedback. For example, if the system detects that an elderly person is experiencing negative emotions while completing a task, it dynamically adjusts the task's difficulty and adds emotional support guidance. The intervention effect is represented by the following function:
[0141] in, As an intervention strategy, For the evaluation results, For the current intervention task, For feedback data. These are the weighting coefficients.
[0142] After all health data is input and processed through deep learning and knowledge graph reasoning, the system generates personalized intervention strategies. This process not only relies on historical health data but also adjusts intervention methods through real-time feedback to ensure the optimal rehabilitation process for each elderly person. The assessment and intervention results generated by the system have the following characteristics: Real-time: The intervention plan is adjusted based on real-time health assessment results, ensuring the timeliness and effectiveness of the intervention. Personalized: Each intervention task is tailored to the specific health condition of the elderly person, ensuring the targeted nature of the intervention plan. Adaptable: As health status changes, the system can adaptively adjust the intervention strategy to avoid over- or under-intervention.
[0143] Through continuous health assessments and interventions, the system constantly learns and optimizes its reasoning process. After each intervention, the system evaluates the intervention's effectiveness using multi-dimensional data (such as assessment results, user feedback, and health trends) and adjusts the next intervention strategy based on the evaluation results. This process can be mathematically represented as:
[0144] in, Indicates the effect of intervention. Feedback data, This is the feedback processing function.
[0145] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment, characterized in that: Includes the following steps: S1. Preprocess the acquired speech signal and extract speech features; S2. Decode the extracted speech features, perform semantic reconstruction and intent reasoning on the decoded text, generate corresponding execution actions and text sequences, and convert the text sequences into speech signals with time sequences. Simultaneously, control the digital human's facial expressions based on the text sequences. S3. Perform unified time management, state determination, and process scheduling control for multimodal interaction tasks using a reentrant multimodal assessment process control model based on finite state machines. S4. Generate a comprehensive compliance score based on the degree of completion of rehabilitation tasks according to plan. S5. Adjust the difficulty of the tasks based on the compliance score and predict the compliance trend of the elderly in the future. S6. Use a combination of deep learning models and knowledge graphs to recommend personalized intervention measures to users.
2. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, Step S1 includes the following steps: S11, performing preprocessing operations on the acquired raw speech, the preprocessed speech is segmented into continuous frames, and the continuity between frames is maintained by a sliding window; S12, extracting speech features of the preprocessed audio signal using an adaptive acoustic modeling method based on deep learning.
3. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, In step S2, semantic reconstruction and intent reasoning of the decoded text includes the following steps: First, the decoded text is structured using a bidirectional long short-term memory network to extract key semantic entities, and the output results serve as the node basis for subsequent semantic graph construction; then, an intent graph is constructed based on the extracted key semantic entities, and semantic reasoning is performed through a graph neural network to generate execution actions and related text sequences; wherein, the intent graph consists of multiple semantic nodes and related edges, where nodes represent task intent or action goals, and edges represent logical relationships and contextual dependencies between different tasks or states. By running a graph neural network on the graph for feature propagation and aggregation, the implicit intent of the user can be identified and path prediction can be performed in different contexts.
4. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, In step S2, controlling the digital human's facial expressions based on text sequences involves extracting emotional cues from acoustic feature parameters to generate corresponding facial expressions and emotional states. Combined with real-time data from rehabilitation assessments, the system judges the user's action completion rate and emotional state. When the user completes the action accurately, the digital human provides positive feedback with a smiling expression and encouraging tone. When the user's action deviates or pauses, the digital human provides prompts with a guiding tone to encourage the user to readjust their action posture.
5. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, Step S3 includes the following steps: First, based on the multimodal input vectors from speech recognition, visual detection, motion sensing, and touch interaction, a finite set of states containing multiple task nodes is established, with each node corresponding to a specific assessment task or interaction scenario; the current state is determined based on changes in the multimodal input vectors, and the output behavior to be executed under the current state and input is determined; when an abnormal situation is detected, the system automatically reverts to the most recent valid state through a reentrant mechanism, restores the assessment progress, and re-executes the corresponding task node; after restoration, new multimodal input vectors are received, and the state judgment is updated based on the user's real-time behavior, dynamically adjusting the execution path of subsequent tasks, thereby achieving continuous execution and intelligent scheduling of elderly motor rehabilitation and cognitive assessment.
6. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, Compliance scoring in step S4 The definition is as follows: in, The timestamp represents the current moment. 、 、 、 These represent the weighting coefficients of each modality in the overall compliance score, used to adjust the contribution ratio of speech, vision, action, and system feedback information to the final score; Indicates input for speech modality The feature evaluation function is used to quantify speech recognition and responsiveness, including voice command recognition accuracy, response timeliness, and semantic coherence. Indicates the input for visual modality The feature evaluation function is used to measure visual attention and recognition ability, including eye focus, image recognition accuracy and interface operation stability. Indicates input for action modality The feature evaluation function is used to evaluate motor response and body coordination, including postural completion, reaction speed and range of motion; Indicates the system interaction state The comprehensive feedback function is used to reflect the system interaction performance during task execution, such as response latency, error recovery rate, and process continuity.
7. The multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment according to claim 1, characterized in that, In step S4, compliance trend is predicted using a deep learning prediction model. The deep learning prediction model uses a recurrent neural network (RNN) or a long short-term memory network (LSTM) to predict the compliance trend of older adults over a future period of time.
8. A system for implementing the multimodal human-computer collaborative interaction method for elderly motor rehabilitation and cognitive assessment as described in claim 1, characterized in that, include: The voice acquisition and processing module is used to acquire voice signals and extract voice features; The system comprises several modules: a speech recognition module for recognizing speech features and generating corresponding actions; a speech playback and digital human expression module for outputting speech and adjusting the digital human's facial expressions and speech rhythm; a compliance scoring module for generating a comprehensive compliance score, adjusting task difficulty or reminder strategies based on the score, and predicting compliance trends in the elderly over a future period; and a personalized intervention module for recommending personalized interventions to users. The system also includes a speech acquisition module for collecting user speech; a preprocessing and feature extraction module for extracting preprocessed speech features and using a user-level fine-tuning module to predict the probability distribution of each speech feature; a semantic reconstruction and intent reasoning module for generating corresponding actions and text sequences based on speech features; a decision control module for matching the intent obtained from the actions with predefined task nodes using cosine similarity calculation to trigger operations; a speech generation and playback module for restoring time-series speech signals from text sequences; and a digital human expression control module for extracting emotional cues from the acoustic feature parameters output by the speech generation and playback module, generating corresponding expressions and emotional behaviors for the digital human, achieving context-adaptive immersive interaction. The time synchronization module is used to control the voice and digital human's facial expressions, so that the voice and vision are synchronized; The assessment process control module based on finite state machines is used to perform unified time management, state determination and process scheduling control of multimodal interactive tasks in the process of elderly motor rehabilitation and cognitive assessment, so as to ensure continuity, fault tolerance and traceability in complex interactive environments. The compliance scoring module generates a comprehensive compliance score based on the degree to which rehabilitation tasks are completed as planned. If compliance in a certain modality is low, the reminder frequency for that modality is increased. The task adjustment module adjusts the difficulty of tasks or reminder strategies based on the compliance score and predicts the compliance trend of the elderly in the future. The personalization module recommends personalized intervention measures to users.
9. A computer device comprising a memory and a processor, the memory being electrically connected to the processor, the memory storing a computer program, characterized in that: When the computer program is executed by the processor, it causes the processor to implement the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the processor implements the method as described in any one of claims 1 to 7.