AI social interaction and cognitive activation method driven based on affective computing

By using an AI-driven social interaction approach powered by affective computing, combined with multimodal emotional information and a personalized strategy library, personalized and dynamic cognitive training for Alzheimer's patients has been achieved. This addresses the issues of high cost and insufficient assessment in existing technologies, thereby improving intervention effectiveness and patient engagement.

CN122436151APending Publication Date: 2026-07-21SHENZHEN YIQI GUANGGUANG TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN YIQI GUANGGUANG TECHNOLOGY CO LTD
Filing Date
2026-04-27
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In the current technology, non-pharmacological interventions for Alzheimer's patients are costly, professional therapists are scarce, and dynamic adjustments cannot be made based on the patient's real-time emotional state and cognitive load. Furthermore, the lack of objective assessment leads to poor intervention results and decreased patient interest and compliance.

Method used

We employ an AI-driven social interaction approach based on affective computing. By acquiring multimodal affective information, analyzing affective computing models, and building a personalized interaction strategy library, we generate dynamic social interaction responses and combine them with cognitive training task recommendations to achieve personalized intervention.

Benefits of technology

This approach enhances the naturalness and accessibility of the intervention, increases patient acceptance, ensures that cognitive training remains within the patient's tolerance range, enables personalized dynamic adjustments, and improves intervention effectiveness and participation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122436151A_ABST
    Figure CN122436151A_ABST
Patent Text Reader

Abstract

The embodiment of the application relates to an AI social interaction and cognitive activation method based on emotion computing driving, comprising: acquiring facial video stream and voice audio stream in patient and social robot interaction; inputting the same into an emotion computing model, outputting patient current emotion state classification information and interaction intention inference information; calling a personalized interaction strategy library according to the information, generating and executing a multi-modal social interaction response to maintain a social dialogue; executing a task dynamic loading decision based on a dialogue process, an emotion state and an interaction intention; collecting multi-dimensional task performance data in a task execution process and after completion; generating a quantitative evaluation result of cognitive ability, emotion response mode and task preference based on analysis of the data, and updating a personal ability portrait; and based on the updated portrait, closing loop optimization strategy library calling logic, interaction response generation strategy and task loading decision logic. Thus, personalized dynamic intervention of social interaction and cognitive training of Alzheimer's disease patients is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of medical rehabilitation engineering and robotics, and in particular to an AI social interaction and cognitive activation method based on emotion computing. Background Technology

[0002] Alzheimer's disease (AD) is a progressive neurodegenerative disease characterized not only by a persistent decline in cognitive functions (such as memory, orientation, and executive function), but also by behavioral and psychological symptoms such as emotional apathy, social withdrawal, depression, and anxiety. Currently, non-pharmacological interventions for AD patients mainly include cognitive training, nostalgia therapy, and social activities led by caregivers or therapists. However, these traditional methods have the following problems: High-quality one-on-one human intervention is costly, and professional therapists are scarce, making it difficult to meet the long-term, high-frequency intervention needs of a large number of patients.

[0003] Traditional intervention programs often consist of pre-set, fixed content that cannot be dynamically adjusted based on the patient's real-time emotional state, cognitive load, and interaction intentions. This results in inconsistent intervention outcomes and may even lead to feelings of frustration in patients due to inappropriate difficulty.

[0004] The evaluation of intervention effects relies heavily on therapists’ subjective observations and scale assessments with long intervals, lacking objective, continuous, and fine-grained quantitative data support, making it difficult to detect subtle changes in patients in a timely manner and adjust intervention strategies.

[0005] Due to the inattention and lack of motivation caused by the disease, patients are prone to losing interest in repetitive and stereotyped training tasks, leading to decreased participation and compliance.

[0006] In recent years, assistive technologies such as social robots have been introduced into the field of AD care. However, existing technologies are mostly limited to simple dialogues or task execution based on preset scripts. They lack the ability to deeply understand the patient's state and make intelligent and empathetic responses. In essence, they are still a passive tool rather than an intelligent partner that can establish a therapeutic relationship and proactively adapt to the patient.

[0007] Therefore, there is an urgent need for a technical solution that can deeply integrate emotional perception, cognitive computing and personalized interaction to solve the above-mentioned bottlenecks. Summary of the Invention

[0008] The purpose of this invention is to address the shortcomings of existing technologies by providing an AI-based social interaction and cognitive activation method driven by emotion computing, thereby solving the problems existing in the prior art.

[0009] To achieve the above objectives, this invention provides an AI-driven social interaction and cognitive activation method based on affective computing, the method comprising: Acquire multimodal emotional information generated by the patient during interaction with the social robot, the multimodal information including the patient's facial video stream and voice audio stream; The multimodal emotional information is input into a preset emotion computing model. The emotion computing model analyzes the facial expression features extracted from the facial video stream and the speech acoustic features extracted from the speech audio stream, and outputs classification information of the patient's current emotional state and inference information of his / her interaction intention. Based on the emotional state classification information and the interaction intention inference information, a preset personalized interaction strategy library is invoked to generate and execute a multimodal social interaction response in order to maintain social dialogue with the patient. Based on the progress of the social dialogue, the emotional state classification information, and the interaction intent inference information, the task dynamic loading decision is executed; During and after the task is completed, multi-dimensional task performance data of patients are collected; Based on the analysis of the multi-dimensional task performance data, a quantitative assessment of the patient's cognitive ability level, emotional response pattern and task preference is generated, and the individual ability profile is updated accordingly. Based on the updated personal ability profile, the logic for calling the personalized interaction strategy library, the generation strategy for the multimodal social interaction response, and the logic for dynamic task loading decisions are optimized in a closed loop.

[0010] In one possible implementation, the emotion computing model includes a facial expression analysis subnetwork, a voice emotion analysis subnetwork, and a multimodal fusion decision subnetwork connected in sequence. The facial expression analysis subnetwork takes the continuous image frame sequence extracted from the facial video stream as input and outputs facial expression feature vectors and preliminary expression classification confidence. The speech emotion analysis subnetwork takes the acoustic feature sequence extracted from the speech audio stream as input and outputs a speech emotion feature vector and a preliminary emotion tendency score. The multimodal fusion decision subnetwork takes the facial expression feature vector, the preliminary expression classification confidence, the voice emotion feature vector, and the preliminary emotion tendency score as joint inputs. It performs feature alignment and weighted fusion through a cross-modal attention mechanism and outputs the classification information of the current emotion state and the inferred information of the interaction intent based on a comprehensive judgment.

[0011] In one possible implementation, the training method for the sentiment computing model includes: Collect and construct a labeled dataset consisting of multiple training samples; wherein each training sample includes a synchronously acquired patient facial video stream, speech audio stream, and manually labeled real emotional state labels and interaction intent labels; The facial expression analysis subnetwork was pre-trained using a general facial expression dataset. The speech emotion analysis subnetwork was pre-trained using a general speech emotion dataset. The multimodal fusion decision subnetwork was initially trained using the labeled dataset. The pre-trained and initialized facial expression analysis subnetwork, speech emotion analysis subnetwork, and multimodal fusion decision subnetwork are connected; the unlabeled facial video stream and speech audio stream in the labeled dataset are used as input, and the corresponding real emotion state labels and interaction intent labels are used as joint supervision targets. The connected overall emotion computing model is trained end-to-end using a multi-task loss function; the multi-task loss function is composed of a weighted sum of emotion state classification loss and interaction intent recognition loss. Obtain a clinical interaction dataset from an Alzheimer's disease patient population; use the clinical interaction dataset to perform supervised fine-tuning of the parameters of the trained emotion computing model.

[0012] In one possible implementation, the step of invoking a preset personalized interaction strategy library based on the emotional state classification information and the interaction intent inference information to generate and execute a multimodal social interaction response to maintain social dialogue with the patient specifically includes: The classification information of the current emotional state and the inferred information of the interaction intention are concatenated into a decision feature vector, and the decision feature vector is input into a preset dialogue management model. The dialogue management model outputs a structured dialogue behavior instruction. The dialogue behavior instruction includes at least a semantic content identifier, a target emotional label, and a non-verbal behavior code. The semantic content identifier is used to query a natural language template library to obtain the corresponding basic text template; based on the target emotion tag, words are selected from a preset emotion adaptation vocabulary library to fill in the variables and adjust the sentence structure in the basic text template to generate the final text sequence to be played; based on the target emotion tag, a preset speech parameter mapping table is queried to obtain the corresponding fundamental frequency reference value, speech rate value and volume gain value as input control parameters for the speech synthesis engine; based on the non-verbal behavior code, a preset robot behavior animation library is queried to obtain the corresponding facial expression control parameter sequence and / or joint motion trajectory sequence; The final text sequence to be played and the input control parameters of the speech synthesis engine are sent to the speech synthesis module of the social robot; at the same time, the facial expression control parameter sequence and / or joint motion trajectory sequence are sent to the motion control module of the social robot to drive the speech synthesis module and the motion control module to synchronously complete the multimodal social interaction response.

[0013] In one possible implementation, the dialogue management model includes a feature encoding subnetwork, a policy decision subnetwork, and an instruction generation subnetwork connected in sequence. The feature encoding subnetwork takes the concatenated decision feature vector as input, processes it through a fully connected layer, and outputs a high-dimensional encoded feature vector. The policy decision subnetwork takes the encoded feature vector as input, processes it through a softmax classification layer, and outputs a probability distribution of a policy identifier. The instruction generation subnetwork takes the policy identifier corresponding to the highest probability in the probability distribution of the policy identifier as input, and retrieves and outputs the corresponding structured dialogue behavior instruction from a preset policy identifier-instruction element mapping table through a lookup operation. The instruction includes at least a semantic content identifier, a target sentiment label, and a non-verbal behavior code.

[0014] In one possible implementation, the dynamic loading decision for the execution task specifically includes: Based on the interaction intent, infer whether the information includes a preset cognitive activity guidance intent, and combine this with whether the social dialogue process is in a preset stable interaction phase to generate a binary task trigger flag. If the task trigger flag is true, then the following operations are performed: a) The emotional state classification information, the interaction intention inference information, the dialogue topic obtained from the real-time analysis of the current social dialogue, and the cognitive ability assessment information based on the personal ability profile are combined to form a task selection feature vector. b) Input the task selection feature vector into a preset task recommendation model, wherein the task recommendation model calculates a numerical matching score for each task in the pre-stored gamified cognitive training task library; c) Select the task with the highest matching score from the task library, load its content definition and interaction flow script into the task execution engine of the social robot, and announce the start of the task to the patient.

[0015] In one possible implementation, the task recommendation model includes a sequentially connected feature encoding subnetwork, a matching degree calculation subnetwork, and a score output subnetwork; The feature encoding subnetwork takes the task-selected feature vector as input, performs nonlinear transformation and feature fusion through a fully connected layer, and outputs a unified high-dimensional task-user joint feature representation vector. The matching degree calculation subnetwork takes the joint feature representation vector as input and includes a learnable task feature matrix. Each row of the matrix corresponds to the feature representation of a task in the gamified cognitive training task library. The subnetwork calculates the correlation between the joint feature representation vector and the features in each row of the task feature matrix, assigns an attention weight to each task, and finally outputs a comprehensive task matching degree vector through a weighted summation method. The scoring output subnetwork takes the comprehensive task matching degree vector as input, and outputs a standardized numerical matching degree score for each task in the task library through a linear transformation layer and normalization processing.

[0016] In one possible implementation, the collection of the patient's multi-dimensional task performance data specifically includes: The task execution engine generates a structured interaction log, which records the patient's response to the task stimulus, the response delay time, and the result of the correctness judgment of the response by timestamp. Within the time window of task execution, the gaze point coordinate sequence and facial motion unit intensity value sequence are extracted from the facial video stream; and the average speech rate, pause frequency and fundamental frequency profile standard deviation are extracted from the speech audio stream. The interaction log, the gaze point coordinate sequence, the facial motion unit intensity value sequence, the average speech rate value, the pause frequency, and the fundamental frequency profile standard deviation are aligned based on a unified time reference and encapsulated into a data packet with a task identifier as the multi-dimensional task performance data.

[0017] In one possible implementation, generating quantitative assessment results and updating the individual competency profile specifically includes: Based on the multi-dimensional task performance data, cognitive performance indicators, emotional response indicators, and task participation indicators are calculated. The cognitive performance indicators include average response accuracy and average reaction time. The emotional response indicators include the proportion of positive expressions based on facial motion units and arousal based on voice fundamental frequency. The task participation indicator is the ratio of the average reaction time of the current task to the individual's historical average reaction time. The cognitive performance index, emotional response index, and task participation index are input into the corresponding preset evaluation functions to generate standardized cognitive dimension scores, emotional dimension scores, and participation dimension scores. The execution timestamp and task identifier of this task, along with the scores of the cognitive dimension, emotional dimension, and participation dimension, are added as a new record to the patient's personal ability profile database. Based on the new record, the statistical characteristic values ​​of each dimension score within a preset time window are recalculated to update the personal ability profile.

[0018] In one possible implementation, the closed-loop optimization based on the updated personal capability profile specifically includes: Read the updated cognitive dimension score and emotional dimension score from the personal ability profile, and adjust the calling priority weight of each strategy in the personalized interaction strategy library based on the score; Based on the updated emotional dimension score and engagement dimension score in the personal ability profile, the configurable parameters controlling the generation of the multimodal social interaction response are adjusted. The configurable parameters include at least the baseline value in the speech synthesis parameter mapping table and the behavior triggering conditions in the robot behavior animation library. The task selection feature vector generated during the execution of this task and the key indicator set extracted from the corresponding multi-dimensional task performance data are used to form training samples to update the parameters of the logic for implementing the dynamic loading decision of the task.

[0019] By applying the AI-driven social interaction and cognitive activation method based on affective computing provided by this invention, a multimodal affective computing model specifically trained for AD patients is used to accurately identify the patient's current emotional state and interaction intentions by integrating multi-channel signals such as facial micro-expressions and speech prosody. This lays a precise cognitive foundation for subsequent personalized interactions. Furthermore, based on deep understanding, a dialogue management model and strategy library are used to generate and execute highly coordinated multimodal responses in terms of language content, speech tone, facial expressions, and body movements. This response is not a fixed script but is dynamically generated based on the patient's real-time state, enabling empathetic communication similar to that of a human caregiver, significantly improving the naturalness, affinity, and patient acceptance of the interaction. Furthermore, this application can intelligently determine the optimal time to introduce cognitive training tasks and, using a task recommendation model, match the most suitable task from the task library in real time based on the patient's current emotions, intentions, dialogue topics, and historical ability profile. This ensures that cognitive training is always within the patient's zone of proximal development, effectively stimulating brain activity while avoiding frustration or boredom caused by inappropriate difficulty, thereby maximizing the effectiveness of cognitive intervention and patient participation. Furthermore, this application systematically collects multi-dimensional performance data, including behavioral logs, visual attention, facial expressions, and voice features, during the task process to generate objective scores for cognition, emotion, and participation, and dynamically updates the individual's ability profile. Based on this profile, the system can optimize interaction strategies, response parameters, and task recommendation models in a closed loop, enabling the entire intervention system to continuously learn and self-improve. This truly achieves personalized and dynamic adjustments to the intervention plan, providing data-driven precision support for delaying cognitive decline and improving emotional state. Attached Figure Description

[0020] Figure 1 An architecture diagram of an AI social interaction and cognitive activation method based on emotion computing provided in an embodiment of the present invention; Figure 2 A flowchart of the AI ​​social interaction and cognitive activation method based on emotion computing provided by this invention; Figure 3 The structural diagram of the emotion computing model provided by this invention; Figure 4 for Figure 2 Flowchart for step 230; Figure 5 The diagram shows the structure of the dialogue management model provided by this invention. Figure 6 The structure diagram of the task recommendation model provided by this invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0022] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Figure 1 This is a scenario architecture diagram of the AI ​​social interaction and cognitive activation method based on emotion computing provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the system of this invention is logically divided into four collaborative functional layers: the perception layer, the cognition layer, the execution layer, and the storage layer. These layers interact in a closed loop through standardized data interfaces and instruction channels, as detailed below: The perception layer serves as the system's data acquisition and front-end processing unit. Through a high-definition camera, microphone array, and gaze tracking module integrated into the social robot, it acquires real-time facial video and audio streams from the patient. It also performs simultaneous preprocessing of multimodal signals, including face detection and alignment, speech endpoint detection, gaze point extraction, and facial action unit analysis, providing high-quality input features for subsequent cognitive analysis.

[0023] The cognitive layer is the core decision-making and computational hub of the system. This layer deploys three deep learning models: an emotion computing model that integrates facial and vocal features to output a classification of the patient's current emotional state and inference of their interaction intent; a dialogue management model that, based on emotion and intent information, generates structured multimodal dialogue behavior instructions through a policy decision network; and a task recommendation model that combines the patient's real-time state with their historical ability profile to calculate a matching score for each task in the gamified cognitive training task library, enabling dynamic and optimal task loading.

[0024] The execution layer is the physical unit that enables the system to interact with the patient. This layer includes a speech synthesis engine, a motion control module, and a task execution engine, and is uniformly scheduled by a central timing controller. The speech synthesis engine adjusts parameters such as fundamental frequency and speech rate based on the target emotion tag to generate emotionally expressive response speech; the motion control module drives the robot's facial expressions and body movements; the task execution engine is responsible for loading cognitive training task scripts, presenting stimuli, recording responses, and judging right and wrong. The three work together to complete multimodal social interaction and cognitive activation intervention.

[0025] The storage layer serves as the system's persistent memory and knowledge management unit. It maintains structured resources such as a personalized interaction strategy library, a natural language template library, an emotion-adaptive vocabulary library, a speech parameter mapping table, a robot behavior animation library, and a gamified cognitive training task library. Simultaneously, the individual ability profile database stores the patient's standardized three-dimensional scores of cognition, emotion, and participation after each task in time-series format, providing data-driven closed-loop feedback for strategy optimization and model updates in the cognitive layer.

[0026] Figure 2 This is a flowchart of an AI-driven social interaction and cognitive activation method based on affective computing. The execution entity of this method is a control system, which includes servers, processors, etc., with processing and computing capabilities, hereinafter referred to as the system. Figure 2 As shown, the method includes the following steps: Step 210: Obtain multimodal emotional information generated by the patient during interaction with the social robot, wherein the multimodal information includes the patient's facial video stream and voice audio stream; Specifically, the acquisition of the multimodal emotional information relies on an integrated perception system. This perception system typically includes a visual acquisition unit and an auditory acquisition unit.

[0027] The visual acquisition unit includes at least one high-definition camera deployed at an appropriate location on the head of the social robot or in the interaction space. This camera acquires a video stream, including the patient's face and upper body posture, at a preset frame rate, such as 15-30 fps, i.e., the facial video stream. To accommodate patient movement, a camera with autofocus and wide dynamic range (WDR) capabilities can be used to ensure clear facial images are obtained under varying lighting and distance conditions.

[0028] The auditory acquisition unit includes at least a microphone array with directional sound pickup or noise reduction capabilities, integrated into the social robot itself or the interactive environment. This unit is used to acquire the sounds emitted by the patient while conversing with the robot, talking to themselves, or performing tasks, forming the speech audio stream. The microphone array helps to focus on the patient's sound source and suppress environmental noise interference.

[0029] Real-time face detection and tracking can be performed on facial video streams. For example, using a face detector based on Haar features or deep learning, the facial region image (ROI) of the patient can be accurately extracted from each frame. Face alignment is then performed, which involves normalizing facial key points to standard positions, grayscale, or size to eliminate the effects of pose, distance, and some illumination variations, providing standardized input for subsequent feature extraction.

[0030] Endpoint detection (VAD) can also be performed on the audio stream to distinguish between speech segments and non-speech segments, such as silence and ambient noise, allowing only valid speech segments to undergo further processing. Subsequently, normalization processes such as pre-emphasis, framing, and windowing are typically performed to prepare for acoustic feature extraction.

[0031] Step 220: Input the multimodal emotional information into a preset emotion computing model. The emotion computing model analyzes the facial expression features extracted from the facial video stream and the speech acoustic features extracted from the speech audio stream, and outputs classification information of the patient's current emotional state and inference information of his / her interaction intention.

[0032] like Figure 3 As shown, the emotion computing model includes a facial expression analysis subnetwork, a voice emotion analysis subnetwork, and a multimodal fusion decision subnetwork connected in sequence; The facial expression analysis subnetwork takes the continuous image frame sequence extracted from the facial video stream as input and outputs facial expression feature vectors and preliminary expression classification confidence. The speech emotion analysis subnetwork takes the acoustic feature sequence extracted from the speech audio stream as input and outputs a speech emotion feature vector and a preliminary emotion tendency score. The multimodal fusion decision subnetwork takes the facial expression feature vector, the preliminary expression classification confidence, the voice emotion feature vector, and the preliminary emotion tendency score as joint inputs. It performs feature alignment and weighted fusion through a cross-modal attention mechanism and outputs the classification information of the current emotion state and the inferred information of the interaction intent based on a comprehensive judgment.

[0033] Specifically, the facial expression analysis subnetwork is used to analyze facial muscle movement patterns in the video stream and identify basic emotional expressions. It typically uses a convolutional neural network (CNN) as its backbone, such as ResNet or a variant of VGG, followed by a temporal modeling module, such as an LSTM, GRU, or Transformer encoder. Its input is a sequence of consecutive image frames extracted from the preprocessed facial video stream, for example, a sequence of T frames. The CNN is responsible for extracting high-level spatial semantic features from each frame, such as the shape of the eyes and mouth, and brow wrinkles. The temporal modeling module is responsible for capturing the dynamic changes in expressions over time, such as the unfolding of a smile. Its output includes a facial expression feature vector and preliminary expression classification confidence. The facial expression feature vector is a high-dimensional, fixed-length vector that incorporates spatiotemporal information, representing an abstract mathematical representation of facial expressions within the current time period. The preliminary expression classification confidence is a probability distribution vector, representing the model's confidence in classifying a patient's expression as belonging to various preset basic emotional categories, such as happiness, sadness, surprise, anger, disgust, fear, and calmness.

[0034] The speech sentiment analysis subnetwork is used to analyze acoustic features in an audio stream, capturing emotional nuances and intonation variations in speech. It typically employs a recurrent neural network (RNN) or its variants, such as Bi-LSTM, or a 1D convolutional neural network (1D-CNN) combined with a self-attention mechanism. Its input is a sequence of acoustic features extracted from the preprocessed speech audio stream, such as MFCCs, pitch, energy, and spectral centroid of each frame. The speech sentiment analysis subnetwork can model the temporal dependencies of acoustic features, thereby understanding prosodic information such as intonation fluctuations and speech rate. Its output includes a speech sentiment feature vector and a preliminary sentiment score. The speech sentiment feature vector is a high-dimensional, fixed-length vector representing the emotional characteristics of the speech. The preliminary sentiment score is a continuous value or classification confidence level, representing the emotional valence (positive / negative) and arousal (high / low) conveyed by the speech; for example, a valence score between [-1, 1].

[0035] The multimodal fusion decision subnetwork is responsible for integrating visual and auditory cues, combined with context, to make the final emotional state judgment and interaction intent inference. Its input consists of facial expression feature vectors and their confidence scores, and speech emotion feature vectors and their scores, which are used as joint inputs.

[0036] Because facial expression feature vectors and speech emotion feature vectors may not be perfectly synchronized in time (e.g., facial expression changes lag slightly behind tone changes), a multimodal fusion decision sub-network learns the temporal correspondence between the two modal features. This correspondence can be learned by assigning different weights to facial expression and speech emotion feature vectors for different emotion categories; for example, facial expression feature vectors have higher weights when the patient is silent, while speech emotion feature vectors have higher weights when the patient's speech is clear. An attention mechanism dynamically assigns different weights to features at different times and in different modalities. The aligned and weighted facial expression and speech emotion feature vectors are concatenated or fused and then input into a fully connected classifier network. This network typically has two parallel output heads: an emotion state classification head and an interaction intent inference head. The emotion state classification head outputs the final, comprehensive emotion state classification information, such as: calm and pleasant, mild anxiety. The interaction intent inference head outputs inferred interaction intent information, such as: intent label: seeking help, answering a question, initiating a new topic, indicating refusal.

[0037] Furthermore, the training method for the sentiment computing model includes: Collect and construct a labeled dataset consisting of multiple training samples; wherein each training sample includes a synchronously acquired patient facial video stream, speech audio stream, and manually labeled real emotional state labels and interaction intent labels; The facial expression analysis subnetwork was pre-trained using a general facial expression dataset. The speech emotion analysis subnetwork was pre-trained using a general speech emotion dataset. The multimodal fusion decision subnetwork was initially trained using the labeled dataset. The pre-trained and initialized facial expression analysis subnetwork, speech emotion analysis subnetwork, and multimodal fusion decision subnetwork are connected; the unlabeled facial video stream and speech audio stream in the labeled dataset are used as input, and the corresponding real emotion state labels and interaction intent labels are used as joint supervision targets. The connected overall model is trained end-to-end using a multi-task loss function; the multi-task loss function is composed of a weighted sum of emotion state classification loss and interaction intent recognition loss. Obtain a clinical interaction dataset from the Alzheimer's disease patient population; use the clinical interaction dataset to perform supervised fine-tuning of the trained model parameters.

[0038] Specifically, firstly, a dedicated multimodal emotion-intent annotation dataset is constructed. Synchronous facial video streams and audio streams are recorded during standardized or semi-structured social interactions between patients and social robots or healthcare professionals. To ensure data diversity, the interaction scenarios need to cover various situations such as daily greetings, topic discussions, simple question-and-answer sessions, and cognitive game guidance. After independently reviewing the video and audio, trained neuropsychologists or clinicians perform dual-labeling annotation on each interaction segment, assigning both a true emotional state label and a true interaction intent label. The true emotional state label is selected from a pre-defined emotion category dictionary, typically including but not limited to: calm, pleasure, interest, confusion, anxiety, frustration, and indifference. The true interaction intent label is selected from a pre-defined intent category dictionary, such as: seeking information, confirming understanding, expressing needs, social initiation, task response, and no explicit intent.

[0039] A multi-annotator cross-validation approach is employed to ensure annotation consistency by calculating inter-rater reliability. The resulting labeled dataset is D_labeled={(V_i, A_i, E_i, I_i)}_{i=1}^{N}, consisting of N training samples, where V represents video, A represents audio, E represents sentiment labels, and I represents intent labels.

[0040] Secondly, a facial expression analysis sub-network is pre-trained using large-scale general facial expression datasets, such as AffectNet and FER-2013, in a supervised learning manner. The loss function is typically cross-entropy loss. This stage enables the network to learn to extract general visual features highly correlated with basic emotions, such as joy, anger, sadness, and surprise, from facial images, and to acquire preliminary expression classification capabilities.

[0041] A sub-network for speech sentiment analysis is pre-trained using large-scale general-purpose speech sentiment datasets such as IEMOCAP and EMO-DB. The model learns to extract acoustic representations related to speech sentiment, such as valence and arousal, from acoustic features such as Mel spectrum.

[0042] Initial training of the multimodal fusion decision sub-network was conducted. While keeping the parameters of the first two sub-networks frozen (i.e., not updated), the fusion decision sub-network was trained independently using a self-built labeled dataset, D_labeled. At this stage, the face and speech sub-networks served only as fixed feature extractors. This allowed the fusion network to initially learn how to effectively combine extracted facial expression features and speech emotion features to predict emotion and intent.

[0043] Next, the three pre-trained and initialized sub-networks are sequentially connected to form a complete deep learning model capable of end-to-end backpropagation. The unlabeled raw facial video stream and speech audio stream from D_labeled are used as model input, with the corresponding (E_i, I_i) label pairs serving as joint supervision targets. A weighted summation multi-task loss function is used for optimization: L_total = α * L_emotion + β * L_intention. Here, L_emotion is the emotion state classification loss, typically using the classification cross-entropy loss function to measure the difference between the model's emotion prediction and the true emotion label. L_intention is the interaction intent recognition loss, also using the classification cross-entropy loss function to measure the difference in intent prediction. α and β are hyperparameters used to balance the importance of the two tasks, typically tuned through validation set performance.

[0044] Use an optimizer such as Adam to minimize L_total. Through the backpropagation algorithm, the error signal not only updates the parameters of the fusion decision subnetwork, but also backpropagates to the face and speech subnetworks, fine-tuning them so that the features they extract are more conducive to the subsequent joint determination of emotion and intent.

[0045] Finally, a relatively small but high-quality clinical interaction dataset, D_clinical, was collected. This dataset comes from the target patient group, such as Alzheimer's patients. Its collection and annotation criteria are similar to D_labeled, but it is more representative of specific patterns such as weakened emotional expression, asynchrony between facial expression and speech, and ambiguous intention expression. The model trained above was used as pre-training weights for further supervised training on the D_clinical dataset. Typically, a higher learning rate is used for the back-end of the model, i.e., the fusion decision layer, to quickly adapt to new data distributions; a lower learning rate is used for the front-end, i.e., the two feature extraction sub-networks, for fine-tuning to prevent forgetting general features. Due to the limited amount of clinical data, stronger regularization strategies are needed, such as dropout and weight decay, and early stopping can be used to determine the stopping time based on validation set performance. Through this fine-tuning, the model transforms from a general emotional intention recognizer into an emotional intention recognizer specifically for Alzheimer's patients, and its output results are more accurate for subsequent personalized social interaction and cognitive training decisions.

[0046] Step 230: Based on the emotional state classification information and the interaction intent inference information, call the preset personalized interaction strategy library to generate and execute a multimodal social interaction response in order to maintain social dialogue with the patient.

[0047] The personalized interaction strategy library is a predefined data structure containing a series of interaction strategy items. Each strategy defines a type of interaction mode that the social robot can adopt in a specific context, and its attributes include: Strategy identifiers are unique identifiers, such as STRAT_PROMPT_SIMPLE (simple hints), STRAT_REWARD_VERBAL (verbal rewards), STRAT_DISTRACT_POSITIVE (positive topic shifts), and STRAT_CHALLENGE_MILD (moderate challenges).

[0048] The strategy content framework defines the possible dialogue topics, sentence templates, and nonverbal behavior types under this strategy.

[0049] Basic call priority weight: a preset initial weight value, representing the default probability of this strategy being selected without personalized adjustments.

[0050] The strategy is applicable to the following tags, which indicate the cognitive level, emotional state, and participation level that the strategy is primarily suited for. For example, the tags could be: {Cognitive Needs: Low, Emotional Support: High, Participation Motivation: Medium}.

[0051] Specifically, such as Figure 4As shown, step 230 includes the following: Step 2301: The classification information of the current emotional state and the inferred information of the interaction intention are concatenated into a decision feature vector, and the decision feature vector is input into a preset dialogue management model. The dialogue management model outputs a structured dialogue behavior instruction. The dialogue behavior instruction includes at least a semantic content identifier, a target emotional label, and a non-verbal behavior code.

[0052] like Figure 5 As shown, the dialogue management model includes a feature encoding subnetwork, a policy decision subnetwork, and an instruction generation subnetwork connected in sequence; The feature encoding subnetwork takes the concatenated decision feature vector as input, processes it through a fully connected layer, and outputs a high-dimensional encoded feature vector. The policy decision subnetwork takes the encoded feature vector as input, processes it through a softmax classification layer, and outputs a probability distribution of a policy identifier. The instruction generation subnetwork takes the policy identifier corresponding to the highest probability in the probability distribution of the policy identifier as input, and retrieves and outputs the corresponding structured dialogue behavior instruction from a preset policy identifier-instruction element mapping table through a lookup operation. The instruction includes at least a semantic content identifier, a target sentiment label, and a non-verbal behavior code.

[0053] Specifically, the dialogue management model is a lightweight neural network model employing a hierarchical decision-retrieval-generation hybrid architecture. It is used to map abstract interaction contexts, including sentiment and intent, into concrete, executable multimodal behavioral instructions.

[0054] The feature encoding subnetwork typically consists of 2-3 fully connected layers, each followed by a non-linear activation function (such as ReLU) and an optional batch normalization layer. The first layer maps the input vector to a higher dimension (e.g., 256-dimensional). Subsequent layers perform non-linear transformations and feature fusion. Its input is a decision feature vector, which is composed of sentiment state classification information, such as calmness, anxiety, and interaction intent inferences (e.g., seeking help, completing an answer) concatenated through one-hot encoding or embedding representation. Its output is a high-dimensional, such as 128-dimensional, encoded feature vector. This vector provides a deep, dense representation of the current interaction context, offering an informational foundation for subsequent policy decisions.

[0055] The policy decision subnetwork is a multi-class classifier. It typically consists of one or two fully connected layers followed by a softmax layer. The number of neurons in the softmax layer equals the total number of predefined policy identifiers, for example, K=20 policies. Its input is the encoded feature vector from the feature encoding subnetwork, and its output is a K-dimensional policy identifier probability distribution vector. In this K-dimensional vector, each dimension corresponds to the probability of selecting a policy. These policy identifiers are predefined, for example: STRAT_COMFORT: (Comfort Strategy) - Triggered when patient anxiety or sadness is detected. STRAT_PRAISE: (Praise Strategy) - Triggered when patient completes a task or interacts positively. STRAT_REDIRECT: (Guidance Strategy) - Guiding the conversation to a new topic or task when the patient's intentions are unclear or the conversation reaches a stalemate. STRAT_SIMPLE_QA: (Simple Question-and-Answer Strategy) - Used to maintain basic conversational flow. STRAT_RECAP: (Recall Guidance Strategy) - Guides the patient to recall their life story.

[0056] The instruction generation subnetwork is used to transform selected high-level policies into multimodal instruction elements that drive specific interactive behaviors. It includes a pre-generated policy identifier-instruction element mapping table stored in memory or a database, as shown in Table 1: Table 1 Its input is the policy identifier corresponding to the highest probability in the probability distribution output by the policy decision subnetwork, such as STRAT_COMFORT. The output is a structured dialogue behavior instruction, a data structure comprising the following three key elements: A semantic content identifier, which points to a specific sentence template in a natural language template library.

[0057] The target emotion tag defines the emotional tone of the robot's response and is used to index voice parameters and behavioral intensity.

[0058] Non-verbal behavior codes index specific action sequences in the robot behavior animation library.

[0059] Furthermore, the training of this dialogue management model is a supervised learning process with the goal of learning an accurate mapping from sentiment-intent features to policy identifiers.

[0060] First, a large number of real or simulated patient-bot or patient-caregiver social dialogue records are collected. Each record includes multiple rounds of dialogue. Input features are patient emotional states and interaction intentions on which the responses in each round are based. These can be pre-labeled by an emotion computing model and then manually verified. Policy labels are the most appropriate policy identifiers for this context, determined by experts and selected from a predefined policy library. Instruction verification involves experts specifying or confirming the ideal instruction elements corresponding to the policy, such as semantic content, emotion labels, and behavioral codes, used to construct or verify the policy identifier-instruction element mapping table. Finally, a training sample set {(decision feature vector_i, target policy identifier_i)} is formed.

[0061] Secondly, a fixed instruction mapping table is established. Based on expert knowledge or data annotation results, a policy identifier-instruction element mapping table is predefined and fixed. This table serves as the model's knowledge base and remains unchanged during training. The neural network is trained using the prepared labeled dataset, training both the feature encoding sub-network and the policy decision sub-network. The standard cross-entropy loss function is used to calculate the difference between the model's predicted policy probability distribution and the truly labeled policy identifiers (One-Hot vectors). Through backpropagation and gradient descent algorithms (such as using the Adam optimizer), the weight parameters of the two sub-networks are continuously adjusted, enabling the model to accurately predict the corresponding policy labeled by the expert based on the input decision feature vector.

[0062] Finally, based on the results, add, delete, or modify policy identifiers and their definitions. Adjust the specific instruction elements corresponding to a policy identifier, or adjust the network structure, such as adjusting the hidden layer dimensions or adding Dropout to prevent overfitting.

[0063] Step 2302: Query a natural language template library based on the semantic content identifier to obtain the corresponding basic text template; based on the target emotion tag, select words from a preset emotion adaptation vocabulary library to fill in the variables and adjust the sentence structure in the basic text template to generate the final text sequence to be played; based on the target emotion tag, query a preset speech parameter mapping table to obtain the corresponding fundamental frequency reference value, speech rate value, and volume gain value as input control parameters for the speech synthesis engine; based on the non-verbal behavior code, query a preset robot behavior animation library to obtain the corresponding facial expression control parameter sequence and / or joint motion trajectory sequence.

[0064] Specifically, a natural language template library is a structured, predefined database or resource file that stores a large number of reusable basic text templates. Each basic text template is bound to a unique semantic content identifier. A basic text template typically includes a fixed text portion and variable placeholders. The fixed text portion forms the unchanging text of the sentence's core. Variable placeholders, marked with specific tags such as [address], [event name], and [adjective], indicate fillable positions for personalization. To adapt to different emotions or contexts, one semantic element can correspond to multiple templates with different sentence structures.

[0065] In one example, the base text template is "[Greeting] Good morning! A new day has begun, shall we [Activity Suggestion] together?" The semantic content identifier is CONT_MORNING_GREETING. Based on the semantic content identifier in the dialogue behavior instruction, the system queries and retrieves the corresponding base text template from the natural language template library.

[0066] In one example, based on the target emotion tag in the dialogue behavior command, such as WARM (warm), CALM (calm), and ENERGETIC (energetic), the system queries an emotion-adaptive vocabulary library. This vocabulary library defines a list of words applicable to different variable types under each emotion tag. For example, for the target emotion tag WARM and the variable type [praise word], the vocabulary library might provide a list of words such as {“intelligent,” “excellent,” “inspiring”}; while for the target emotion ENERGETIC, the vocabulary list could be {“wonderful,” “energetic,” “wonderful”}.

[0067] In another example, an appropriate word can be selected from an emotion-fitting vocabulary to fill each variable placeholder. The selection strategy can be deterministic, such as rotation, or based on simple rules, such as random selection. Certain base text templates can be fine-tuned in terms of sentence structure based on emotional intensity. For example, for high-intensity praise, an exclamation like "Wow!" might be added before the template. [Titles] These variables are typically derived from the patient's profile, taking their preferred titles, such as "Grandpa Li" or "Aunt Wang." After filling and adjusting, a complete, final text sequence to be played, conforming to the target emotional tone, is generated.

[0068] The speech parameter mapping table is a predefined lookup table that maps abstract target emotion tags to a set of specific, quantifiable acoustic parameter values. These parameters are those directly accessible through industry-standard speech synthesis markup languages ​​such as SSML or engine APIs. Based on the target emotion tag in the instruction, the system queries this mapping table to obtain the fundamental frequency reference value, speech rate value, and volume gain value in real time.

[0069] The fundamental frequency (FFM) is the pitch of the sound. For example, JOYFUL corresponds to a higher FFM, and SAD corresponds to a lower FFM. The speech rate (PR) is the speed at which the speech is delivered. For example, CALM corresponds to a moderately slow speech rate, and EXCITED might correspond to a faster speech rate. The volume gain is the loudness of the sound. For example, the volume might be slightly increased to attract attention or express affirmation. These parameters are encapsulated as a set of input control parameters for the speech synthesis engine, which are sent directly to the speech synthesis module to guide it in synthesizing speech with specific emotional intonation.

[0070] To achieve human-like nonverbal communication, abstract behavioral codes need to be translated into precise motion instructions for robot hardware. A robot behavior animation library serves as a database storing robot motion skills. It maps each nonverbal behavioral code to a specific, executable robot motion description file. For each behavioral code, the data stored in the library may include facial expression control parameter sequences, joint motion trajectory sequences, and synchronization markers. Facial expression control parameter sequences, for screen displays or mechanical faces, can be a series of image frames, vector graphics instructions, or time-series data of servo motor angles, used to generate expressions such as smiling, blinking, and confusion. Joint motion trajectory sequences, for joints such as the head, arms, and body, are a series of trajectory points related to the target angle, velocity, and acceleration of the joint, usually indexed by time. For example, NOD_SLOW, the code for a slow nod, corresponds to the neck joint completing a specific angle up-and-down motion trajectory within 2 seconds. Synchronization markers are synchronization markers between key motion points and the speech timeline, ensuring that the action and speech emphasis are synchronized, such as nodding when saying "yes."

[0071] Based on the non-verbal behavior codes in the instructions, the system retrieves and calls the corresponding control parameter sequences from the behavior animation library. These facial expression control parameter sequences and / or joint motion trajectory sequences are sent to the robot's underlying motion control module.

[0072] Step 2303: The final text sequence to be played and the input control parameters of the speech synthesis engine are sent to the speech synthesis module of the social robot; at the same time, the facial expression control parameter sequence and / or joint motion trajectory sequence are sent to the motion control module of the social robot to drive the speech synthesis module and the motion control module to synchronously complete the multimodal social interaction response.

[0073] Specifically, the input to the speech synthesis module is the final text sequence to be played and the input control parameters of the speech synthesis engine.

[0074] The system maintains a high-precision central timing controller. Before issuing an execution command, the central controller compares and calibrates the estimated duration of the speech synthesis with the duration of the action sequence. Synchronization start signal: The central controller simultaneously sends a synchronization start command with a precise future execution time (e.g., t = current time + 50ms) to both modules. Each module buffers data internally and begins execution uniformly at the agreed time t.

[0075] The motion control module monitors the real-time status of the speech synthesis module. For example, when the speech engine plays a specific phoneme or a preset emphasis point, it emits an event signal. Upon receiving this signal, the motion control module immediately triggers the corresponding key action, such as executing the vertex frame of a nodding action when saying "yes." In scenarios requiring strict synchronization, such as "speaking while pointing," the motion control module, after completing the preparatory stage of a certain action, will send a ready signal back to the speech module, which will then trigger the corresponding speech segment.

[0076] The speech synthesis module first loads the corresponding acoustic model and speaker data based on the identifiers in the control parameters. It then performs word segmentation, prosody prediction, and phoneme conversion on the input text. Combining the control parameters, including adjusting the fundamental frequency and duration, it generates the final speech waveform data through a vocoder. The audio stream is then played through the robot's speaker system. The playback process itself can be streaming to support real-time feedback for long sentences.

[0077] The motion control module receives a sparse sequence of key points from the trajectory. High-frequency interpolation calculations, such as fifth-order polynomial interpolation, are performed between these points to generate target angle commands for the servo motor every millisecond, ensuring smooth, jitter-free motion. The controller sends the target angle commands to the servo motor and reads the encoder feedback from the motor in real time, forming a closed-loop control to ensure precise movement. Simultaneously, current and position are monitored to prevent overload or collisions; any abnormality immediately triggers a protection mode.

[0078] Simultaneously with or immediately after the execution of the multimodal social interaction response generated in step 230, the system asynchronously executes the subsequent task dynamic loading decision, i.e., step 240. The execution of the task dynamic loading decision and the multimodal social interaction response adopts a parallel or pipelined approach. That is, while the robot is broadcasting speech or making facial expressions, the background asynchronously executes the judgment of task triggering conditions and the calculation of the task recommendation model. If the decision result is to load the task, it is loaded immediately after the current multimodal response is completed; otherwise, it returns to step 210 to continue the next round of interaction loop.

[0079] Step 240: Based on the progress of the social dialogue, the emotional state classification information, and the interaction intent inference information, execute the task dynamic loading decision.

[0080] Specifically, the dynamic loading decision for the execution task includes: Based on the interaction intent, infer whether the information includes a preset cognitive activity guidance intent, and combine this with whether the social dialogue process is in a preset stable interaction phase to generate a binary task trigger flag. The preset smooth interaction phase is defined as a time window that simultaneously meets the following conditions: In the most recent N consecutive rounds of dialogue (N is a fixed value between 3 and 5, configured during system initialization), the patient's real-time emotional state classification results all belonged to positive or neutral categories such as {calm, pleasant, interested}, and no negative categories such as {anxious, frustrated, angry, indifferent} appeared. Within the time window, the intensity values ​​of facial action units related to negative emotions extracted from the facial video stream, such as the intensity of AU4 "corrugator supercilii" and the intensity values ​​of the AU1+2 "inner / outer eyebrow lift" combination, are all lower than the preset threshold, such as intensity value <2, with a value range of 0-5. The standard deviation of the fundamental frequency profile extracted from the speech audio stream does not exhibit an abnormally steep increase, meaning the fundamental frequency change rate between two adjacent frames is less than 30%, and no prolonged pauses are detected, such as no speech segment lasting more than 3 seconds or the speech rate suddenly dropping to less than 50% of the normal value. When all of the above conditions are met, it is considered to be in a stable interaction phase.

[0081] The pre-defined cognitive activity guidance intention refers to the presence of any of the following intention labels in the interaction intention inference information: "Seeking cognitive stimulation": The patient actively expresses semantics such as "want to play games", "want some challenges", "bored"; "Accepting guidance": The patient shows a neutral and open attitude in social dialogue, such as the intention inference confidence level > 0.7, and there are no rejection or avoidance signals; "Expressing the willingness to recall": The patient actively mentions past experiences, people or events, such as "I remember before...", "Back then..."; "Not explicitly expressed but implicitly guideable": The emotion computing model judges that the patient is currently in a combined state of "calm and intentional interaction", such as the emotion classification is calm, and the interaction intention inference confidence level > 0.6, and there are no negative intention labels such as "rejection" or "confusion".

[0082] The complete list of intent labels and their judgment logic are predefined and solidified during the training process of the sentiment computing model. The output layer of the sentiment computing model contains classification neurons corresponding to the aforementioned intents.

[0083] If the task trigger flag is true, then the following operations are performed: a) The emotional state classification information, the interaction intent inference information, the dialogue topic obtained from real-time analysis of the current social dialogue, and the cognitive ability assessment information based on the individual ability profile are collectively used to construct a task selection feature vector. The dialogue topic is obtained in real-time through the following method: The system maintains a lightweight topic classification model based on the BERT-mini architecture. This model takes the concatenated text of the most recent three rounds of dialogue as input and outputs a topic label with the highest probability distribution. The categories of the topic label are predefined as {family, past events, health, hobbies, daily life, no specific topic}. If the current text length is insufficient or the confidence level is lower than a preset threshold, such as 0.6, the topic of the previous round of dialogue is used as the current topic, or it is marked as "unknown". b) Input the task selection feature vector into a preset task recommendation model, wherein the task recommendation model calculates a numerical matching score for each task in the pre-stored gamified cognitive training task library; c) Select the task with the highest matching score from the task library, load its content definition and interaction flow script into the task execution engine of the social robot, and announce the start of the task to the patient.

[0084] like Figure 6 As shown, the task recommendation model includes a sequentially connected feature encoding subnetwork, a matching degree calculation subnetwork, and a score output subnetwork; The feature encoding subnetwork takes the task-selected feature vector as input, performs nonlinear transformation and feature fusion through a fully connected layer, and outputs a unified high-dimensional task-user joint feature representation vector. The matching degree calculation subnetwork takes the joint feature representation vector as input and includes a learnable task feature matrix. Each row of the matrix corresponds to the feature representation of a task in the gamified cognitive training task library. The subnetwork calculates the correlation between the joint feature representation vector and the features in each row of the task feature matrix, assigns an attention weight to each task, and finally outputs a comprehensive task matching degree vector through a weighted summation method. The scoring output subnetwork takes the comprehensive task matching degree vector as input, and outputs a standardized numerical matching degree score for each task in the task library through a linear transformation layer and normalization processing.

[0085] Specifically, the task recommendation model is a deep learning model. It compares multi-dimensional, dynamically changing patient state information with static task database features to calculate the most suitable cognitive training task for each patient in their specific interactive context.

[0086] The feature encoding sub-network serves as the model's input, fusing heterogeneous, multi-source task selection feature vectors and mapping them to a unified semantic space. Its input is the task selection feature vector, which is composed of dynamic sentiment and intention states, dynamic contextual information, and a static ability profile. Dynamic sentiment and intention states include current sentiment state classification information, such as encoding "calm" and "pleasure," and interaction intention inference information, such as encoding "willing to participate" and "seeking challenges." Dynamic contextual information consists of dialogue topics extracted in real-time from the current conversation, such as encoding topics like "family" and "past events." The static ability profile is cognitive ability assessment information based on the individual's ability profile, such as numerical representations of recent memory and attention dimension scores.

[0087] The feature encoding subnetwork typically consists of 2-3 fully connected layers, each followed by a non-linear activation function such as ReLU and dropout regularization. The first layer reduces the dimensionality of the high-dimensional sparse concatenated vector to a constant dimension. Subsequent layers perform deep non-linear transformations and feature interactions, allowing different types of information, such as the emotion "pleasure" and the ability "high memory score," to influence and merge with each other. Finally, the output is a unified high-dimensional task-user joint feature representation vector. This vector forms the basis for subsequent matching calculations.

[0088] The input to the matching degree calculation subnetwork is the joint feature representation vector from the feature encoding subnetwork. This includes a learnable task feature matrix, which is also a model parameter matrix. The number of rows in this matrix equals the total number of tasks M in the gamified cognitive training task library. Each row is a vector with the same dimension as the joint feature representation vector, representing the learnable embedding representation of the corresponding task. This matrix is ​​randomly initialized before training and is continuously optimized and updated during model training based on feedback from user interaction data. Ultimately, the vector in each row of the matrix represents the deep-level features of the corresponding task, such as the cognitive functions primarily trained, the approximate difficulty level, the required emotional investment, and the rhythm and style of the interaction.

[0089] The matching score calculation subnetwork computes the joint feature representation vector of the input, representing the relevance score between the user's current state and each row of the task feature matrix. The M calculated raw relevance scores are then input into a Softmax layer for processing. The Softmax layer transforms the raw scores into a probability distribution, assigning an attention weight to each task. Higher weights indicate higher relevance and potential suitability of the task in the current user state. The calculated attention weights are then used to perform a weighted summation across all rows of the task feature matrix. The output is a comprehensive task matching score vector.

[0090] The scoring output subnetwork transforms the internal matching degree representation into a standardized score that can be used for ranking and selection. The input is a comprehensive task matching degree vector from the matching degree calculation subnetwork. First, a linear transformation layer, such as a fully connected layer without an activation function, maps the matching degree vector to a scalar or a low-dimensional vector. This step can be understood as a comprehensive scoring of the matching degree. Then, the linearly transformed score is normalized using a normalization function, for example, using Softmax after parallel computation for all tasks, or using the Sigmoid function to compress the score to the 0-1 range, generating a standardized numerical matching degree score. Finally, the output is a vector of length M, where the i-th element is the matching degree score of the i-th task in the task library. A higher score indicates a better match between the task and the patient's current emotional state, intention, conversation topic, and cognitive ability, resulting in a higher priority for task loading.

[0091] Furthermore, the training method for the task recommendation model is as follows: Construct an offline training dataset D_offline = {(x_i, y_i)} from the system's historical logs. Here, the feature sample x_i is the task selection feature vector at a historical task loading decision moment. It includes the emotional state, interaction intent, dialogue topic, and personal ability profile information recorded at that moment. The label y_i is the multi-dimensional suitability label for the actual task execution. This label is calculated using multi-dimensional task performance data collected after the task execution and is a comprehensive scalar. An example of the calculation method is as follows: Suitability = w1 * Task completion rate + w2 * Positive change in patient's emotion + w3 * (1 - Normalized reaction time) The weights w1, w2, and w3 are set based on empirical values ​​to ensure that the labels reflect the overall effect of the task in terms of cognitive challenge, emotional experience, and engagement.

[0092] The task recommendation model is trained using the pre-constructed dataset D_offline. A list-based ranking loss function is used. For a sample x_i, the model calculates scores for all tasks in the task library, resulting in a score vector s. The goal of the loss function is to maximize the score of the selected task (with a high suitability label y_i) while minimizing the scores of other tasks (especially those with low y_i values).

[0093] In one specific implementation, for sample x_i, the score difference between the actually executed task t_pos and each unexecuted task t_neg is calculated, and this difference is encouraged to be greater than a boundary value m.

[0094] The formula is L = Σ max(0, m - (s(t_pos) - s(t_neg))) The training process optimizes all model parameters by minimizing the aforementioned loss function and using the backpropagation algorithm, including the weights of the feature encoding subnetwork, the learnable task feature matrix in the matching degree calculation subnetwork, and the weights of the scoring output subnetwork.

[0095] During the online operation, the system continuously records the state-action-report sequence, forming trajectory data τ = (s_1,a_1, r_1, s_2, a_2, r_2, ...).

[0096] Wherein, state s_t is the task selection feature vector at time t. Action a_t is the task recommended and executed by the model at time t. Reward r_t is a comprehensive reward function, which, for short-term rewards, can be calculated based on immediate multi-dimensional performance data after task execution, along with the suitability label from supervised learning. For long-term rewards, the improvement trend of the cognitive dimension in the patient's personal ability profile is assessed periodically (e.g., weekly), and the degree of improvement is allocated to the recently executed task sequence in the form of delayed rewards.

[0097] Finally, r_t = R_short + γ * R_long, where γ is the discount factor.

[0098] To avoid the risks associated with online exploration, such as inappropriate recommendation tasks, pre-trained models can be fine-tuned using conservative offline reinforcement learning algorithms (such as Conservative Q-Learning, CQL). ​​Specifically, during training, the algorithm imposes a conservative constraint while maximizing the expected reward, preventing the policy from deviating excessively from the proven safe recommendation behavior in the offline dataset. This ensures the model's safety when exploring better recommendation strategies.

[0099] Step 250: Collect multi-dimensional task performance data of the patient during and after the task execution.

[0100] Specifically, the collection of patients' multi-dimensional task performance data includes: The task execution engine generates a structured interaction log, which records the patient's response to the task stimulus, the response delay time, and the result of the correctness judgment of the response by timestamp. Within the time window of task execution, the gaze point coordinate sequence and facial motion unit intensity value sequence are extracted from the facial video stream; and the average speech rate, pause frequency and fundamental frequency profile standard deviation are extracted from the speech audio stream. The interaction log, the gaze point coordinate sequence, the facial motion unit intensity value sequence, the average speech rate value, the pause frequency, and the fundamental frequency profile standard deviation are aligned based on a unified time reference and encapsulated into a data packet with a task identifier as the multi-dimensional task performance data.

[0101] Specifically, the structured interaction log is a machine-readable file automatically generated by the task execution engine during runtime, recording atomic events of task interactions in strict chronological order. It includes specific information about the patient's responses to task stimuli. For multiple-choice questions, it includes option identifiers; for voice-based question-and-answer sessions, it includes the text converted by the Automatic Speech Recognition (ASR) engine; and for drag-and-drop operations, it includes the operation object and target location.

[0102] The fixation point coordinate sequence is a time series sampled at a fixed frequency, such as 30Hz. Each data point in the sequence represents the two-dimensional coordinate position of the patient's visual attention on the task interface (such as a screen) at a specific moment. It is obtained by processing a real-time facial video stream using a gaze-tracking algorithm. This sequence is fundamental data for analyzing visual attention allocation patterns, fixation stability, and information retrieval efficiency.

[0103] The facial action unit intensity value sequence is a time series, where each data point represents the activation intensity value of one or more predefined facial action units (AUs, as defined by the FACS system) at a specific moment, typically a continuous value from 0 to 5 or a standardized score. It is calculated in real-time from the video stream using a facial expression analysis model. The combination and intensity of specific AUs, such as AU4 (corrugator supercilii) and AU12 (zygomaticus major), have an interpretable correlation with basic emotions such as confusion, pleasure, and cognitive load levels, serving as micro-expression signals that quantify emotional responses and cognitive effort.

[0104] The alignment based on a unified time base specifically employs the following method: The timestamps in the interaction logs generated by the task execution engine are used as the master clock reference, and the time precision of the logs is in the millisecond range. For the gaze point coordinate sequence and facial action unit intensity value sequence extracted from the facial video stream, due to the inconsistency between the video frame rate (e.g., 30 fps) and the log event frequency, a linear interpolation method is used to resample the gaze point coordinates and AU intensity values ​​to equal-interval time points perfectly aligned with the log timestamps. The interval is the smallest time unit of the log, such as 100ms. Missing gaze point data, such as when the patient briefly closes their eyes or turns their head, is filled with the previous valid value. If consecutive missing values ​​exceed 500ms, they are marked as invalid. The average speech rate, pause frequency, and fundamental frequency profile standard deviation extracted from the speech audio stream are calculated at the task stage, such as "stimulus presentation period" and "response waiting period". Frame-level alignment is not performed. The stage identifier is directly used as the time anchor point and associated with the interaction logs and video features within the corresponding time period. The timestamps in the final packaged data packets uniformly adopt the relative time from the start of the task, such as the unit is milliseconds, and are accompanied by the timestamp mapping table of the original acquisition device for backtracking verification.

[0105] Step 260: Based on the analysis of the multi-dimensional task performance data, generate quantitative assessment results of the patient's cognitive ability level, emotional response pattern and task preference, and update the individual ability profile accordingly.

[0106] Specifically, generating quantitative assessment results and updating individual competency profiles includes: Based on the multi-dimensional task performance data, cognitive performance indicators, emotional response indicators, and task participation indicators are calculated. The cognitive performance indicators include average response accuracy and average reaction time. The emotional response indicators include the proportion of positive expressions based on facial motion units and arousal based on voice fundamental frequency. The task participation indicator is the ratio of the average reaction time of the current task to the individual's historical average reaction time. The cognitive performance index, emotional response index, and task participation index are input into the corresponding preset evaluation functions to generate standardized cognitive dimension scores, emotional dimension scores, and participation dimension scores. The execution timestamp and task identifier of this task, along with the scores of the cognitive dimension, emotional dimension, and participation dimension, are added as a new record to the patient's personal ability profile database. Based on the new record, the statistical characteristic values ​​of each dimension score within a preset time window are recalculated to update the personal ability profile.

[0107] Specifically, the individual competency profile is a multidimensional state database stored in a time-series format, with the patient ID as the primary key. The core fields of each record include a timestamp, a task identifier, and scores for three standardized dimensions: cognitive, emotional, and engagement dimensions.

[0108] The cognitive dimension score is a standardized numerical value, for example, between 0 and 1, representing the patient's current overall cognitive ability level. This score is calculated by the assessment function in step 260 based on cognitive performance indicators such as average response accuracy and average reaction time. A higher score indicates a better current cognitive function and a stronger ability to process complex information.

[0109] The affective dimension score is a standardized numerical value reflecting the patient's current emotional state in terms of positivity and stability. It is calculated based on indicators such as the proportion of positive facial expressions and changes in fundamental voice frequency, reflecting the patient's emotional positivity and stability. A decrease in the score signals frustration, anxiety, or apathy. A higher score indicates a more positive and stable emotional state.

[0110] The engagement dimension score is a standardized numerical value reflecting the patient's intrinsic interest and level of engagement with the current interaction and task. This score is calculated by the assessment function in step 260 based on engagement indicators such as the task acceptability index. A higher score indicates a stronger willingness to participate and a lower likelihood of avoidance or withdrawal behaviors.

[0111] Step 270: Based on the updated personal ability profile, optimize the calling logic of the personalized interaction strategy library, the generation strategy of the multimodal social interaction response, and the parameters of the task recommendation model in a closed loop.

[0112] Specifically, based on the updated personal ability profile, the following closed-loop adjustment operations are performed: The priority weights of each strategy in the personalized interaction strategy library are adjusted according to the updated cognitive and emotional dimension scores in the personal ability profile; the baseline values ​​in the voice parameter mapping table used for generating multimodal social interaction responses and the behavior triggering conditions in the robot behavior animation library are adjusted according to the updated emotional and participation dimension scores in the personal ability profile; and the parameters of the task recommendation model on which the task dynamic loading decision depends are updated using the task selection feature vector generated during this task execution process and the key indicator set extracted from the corresponding multidimensional task performance data to form training samples.

[0113] Specifically, the closed-loop optimization based on the updated personal ability profile includes: Read the updated cognitive dimension score and emotional dimension score from the personal ability profile, and adjust the calling priority weight of each strategy in the personalized interaction strategy library based on the score; Based on the updated emotional dimension score and engagement dimension score in the personal ability profile, the configurable parameters controlling the generation of the multimodal social interaction response are adjusted. The configurable parameters include at least the baseline value in the speech synthesis parameter mapping table and the behavior triggering conditions in the robot behavior animation library. The task selection feature vector generated during the execution of this task and the key indicator set extracted from the corresponding multi-dimensional task performance data are used to form training samples to update the parameters of the logic for implementing the dynamic loading decision of the task. Based on the updated scores across the three dimensions, the system adjusts the invocation priority weight of each strategy item in the strategy library in real time using mapping rules or weight adjustment functions. The adjustment principle is to increase the weight of strategies that best match the patient's current state and decrease the weight of mismatched strategies, thereby making subsequent interactive strategy invocations more personalized, supportive, and effective.

[0114] Specifically, the call priority weight of each policy item in the policy library is adjusted in real time through mapping rules or weight adjustment functions. The adjustment process is as follows: A score-weight influence matrix is ​​constructed, and a three-dimensional weight adjustment coefficient vector (Δ_cog, Δ_emo, Δ_eng) is predefined for each strategy in the strategy library. This vector defines how the weights of the strategy should be adjusted as the scores in the three dimensions change. In an example, for the strategy STRAT_CHALLENGE_MILD, its adjustment coefficient vector might be (+0.3, +0.1, +0.2). This means that when the cognitive dimension score increases, the weight of the strategy should increase significantly; when the engagement score increases, the weight should also increase; and the affective dimension score has a slight positive impact.

[0115] The total adjustment is calculated based on the difference between the latest scores in the three dimensions (S_cog, S_emo, S_eng) and the baseline scores of the strategy (B_cog, B_emo, B_eng), combined with the adjustment coefficient vector.

[0116] The calculation formula is: Weight adjustment amount=Δ_cog *(S_cog - B_cog)+Δ_emo*(S_emo - B_emo)+Δ_eng *(S_eng - B_eng) If the patient's current cognitive score S_cog is higher than the baseline B_cog, and Δ_cog is positive, then the weight of this strategy receives a positive increase, and the calculated weight adjustment is added to the current priority weight of the strategy. After adjusting the weights of all strategies, normalization is performed, for example, using the Softmax function, to ensure that the sum of all weights is 1, and that each weight represents the probability of selecting the strategy the next time it is invoked.

[0117] In a specific example, suppose a patient's current multidimensional score status is as follows: Cognitive dimension score S_cog: low (0.3, indicating a temporary decline in cognitive processing ability), Affective dimension score S_emo: low (0.4, indicating some depression or anxiety), and Engagement dimension score S_eng: moderate (0.6). The adjustment coefficient vector for the strategy STRAT_PROMPT_SIMPLE (simple hints) is (-0.2, +0.1, +0.0), meaning that the weight of this strategy should actually increase when the cognitive score is low. The adjustment coefficient vector for the strategy STRAT_REWARD_VERBAL (verbal reward) is (+0.0, +0.3, +0.2), designed to provide encouragement when the patient is depressed. After calculation, its weight will be significantly increased. The adjustment coefficient vector for the strategy STRAT_CHALLENGE_MILD (moderate challenge) is (+0.3, +0.1, +0.2).

[0118] By applying the AI-driven social interaction and cognitive activation method based on affective computing provided by this invention, a multimodal affective computing model specifically trained for AD patients is used to accurately identify the patient's current emotional state and interaction intentions by integrating multi-channel signals such as facial micro-expressions and speech prosody. This lays a precise cognitive foundation for subsequent personalized interactions. Furthermore, based on deep understanding, a dialogue management model and strategy library are used to generate and execute highly coordinated multimodal responses in terms of language content, speech tone, facial expressions, and body movements. This response is not a fixed script but is dynamically generated based on the patient's real-time state, enabling empathetic communication similar to that of a human caregiver, significantly improving the naturalness, affinity, and patient acceptance of the interaction. Furthermore, this application can intelligently determine the optimal time to introduce cognitive training tasks and, using a task recommendation model, match the most suitable task from the task library in real time based on the patient's current emotions, intentions, dialogue topics, and historical ability profile. This ensures that cognitive training is always within the patient's zone of proximal development, effectively stimulating brain activity while avoiding frustration or boredom caused by inappropriate difficulty, thereby maximizing the effectiveness of cognitive intervention and patient participation. Furthermore, this application systematically collects multi-dimensional performance data, including behavioral logs, visual attention, facial expressions, and voice features, during the task process to generate objective scores for cognition, emotion, and participation, and dynamically updates the individual's ability profile. Based on this profile, the system can optimize interaction strategies, response parameters, and task recommendation models in a closed loop, enabling the entire intervention system to continuously learn and self-improve. This truly achieves personalized and dynamic adjustments to the intervention plan, providing data-driven precision support for delaying cognitive decline and improving emotional state.

[0119] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0120] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0121] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for AI-driven social interaction and cognitive activation based on affective computing, characterized in that, The method includes: Acquire multimodal emotional information generated by the patient during interaction with the social robot, the multimodal information including the patient's facial video stream and voice audio stream; The multimodal emotional information is input into a preset emotion computing model. The emotion computing model analyzes the facial expression features extracted from the facial video stream and the speech acoustic features extracted from the speech audio stream, and outputs classification information of the patient's current emotional state and inference information of his / her interaction intention. Based on the emotional state classification information and the interaction intention inference information, a preset personalized interaction strategy library is invoked to generate and execute a multimodal social interaction response in order to maintain social dialogue with the patient. Based on the progress of the social dialogue, the emotional state classification information, and the interaction intent inference information, the task dynamic loading decision is executed; During and after the task is completed, multi-dimensional task performance data of patients are collected; Based on the analysis of the multi-dimensional task performance data, a quantitative assessment of the patient's cognitive ability level, emotional response pattern and task preference is generated, and the individual ability profile is updated accordingly. Based on the updated personal ability profile, the logic for calling the personalized interaction strategy library, the generation strategy for the multimodal social interaction response, and the logic for dynamic task loading decisions are optimized in a closed loop.

2. The method according to claim 1, characterized in that, The emotion computing model includes a facial expression analysis subnetwork, a voice emotion analysis subnetwork, and a multimodal fusion decision subnetwork connected in sequence; The facial expression analysis subnetwork takes the continuous image frame sequence extracted from the facial video stream as input and outputs facial expression feature vectors and preliminary expression classification confidence. The speech emotion analysis subnetwork takes the acoustic feature sequence extracted from the speech audio stream as input and outputs a speech emotion feature vector and a preliminary emotion tendency score. The multimodal fusion decision subnetwork takes the facial expression feature vector, the preliminary expression classification confidence, the voice emotion feature vector, and the preliminary emotion tendency score as joint inputs. It performs feature alignment and weighted fusion through a cross-modal attention mechanism and outputs the classification information of the current emotion state and the inferred information of the interaction intent based on a comprehensive judgment.

3. The method according to claim 2, characterized in that, The training method for the sentiment computing model includes: Collect and construct a labeled dataset consisting of multiple training samples; wherein each training sample includes a synchronously acquired patient facial video stream, speech audio stream, and manually labeled real emotional state labels and interaction intent labels; The facial expression analysis subnetwork was pre-trained using a general facial expression dataset. The speech emotion analysis subnetwork was pre-trained using a general speech emotion dataset. The multimodal fusion decision subnetwork was initially trained using the labeled dataset. The pre-trained and initialized facial expression analysis subnetwork, speech emotion analysis subnetwork, and multimodal fusion decision subnetwork are connected; the unlabeled facial video stream and speech audio stream in the labeled dataset are used as input, and the corresponding real emotion state labels and interaction intent labels are used as joint supervision targets. The connected overall emotion computing model is trained end-to-end using a multi-task loss function; the multi-task loss function is composed of a weighted sum of emotion state classification loss and interaction intent recognition loss. Obtain a clinical interaction dataset from an Alzheimer's disease patient population; use the clinical interaction dataset to perform supervised fine-tuning of the parameters of the trained emotion computing model.

4. The method according to claim 1, characterized in that, The step of generating and executing a multimodal social interaction response by invoking a preset personalized interaction strategy library based on the emotional state classification information and the interaction intent inference information, in order to maintain social dialogue with the patient, specifically includes: The classification information of the current emotional state and the inferred information of the interaction intention are concatenated into a decision feature vector, and the decision feature vector is input into a preset dialogue management model. The dialogue management model outputs a structured dialogue behavior instruction. The dialogue behavior instruction includes at least a semantic content identifier, a target emotional label, and a non-verbal behavior code. The semantic content identifier is used to query a natural language template library to obtain the corresponding basic text template; based on the target emotion tag, words are selected from a preset emotion adaptation vocabulary library to fill in the variables and adjust the sentence structure in the basic text template to generate the final text sequence to be played; based on the target emotion tag, a preset speech parameter mapping table is queried to obtain the corresponding fundamental frequency reference value, speech rate value and volume gain value as input control parameters for the speech synthesis engine; based on the non-verbal behavior code, a preset robot behavior animation library is queried to obtain the corresponding facial expression control parameter sequence and / or joint motion trajectory sequence; The final text sequence to be played and the input control parameters of the speech synthesis engine are sent to the speech synthesis module of the social robot; at the same time, the facial expression control parameter sequence and / or joint motion trajectory sequence are sent to the motion control module of the social robot to drive the speech synthesis module and the motion control module to synchronously complete the multimodal social interaction response.

5. The method according to claim 4, characterized in that, The dialogue management model includes a feature encoding subnetwork, a policy decision subnetwork, and an instruction generation subnetwork connected in sequence. The feature encoding subnetwork takes the concatenated decision feature vector as input, processes it through a fully connected layer, and outputs a high-dimensional encoded feature vector. The policy decision subnetwork takes the encoded feature vector as input, processes it through a softmax classification layer, and outputs a probability distribution of a policy identifier. The instruction generation subnetwork takes the policy identifier corresponding to the highest probability in the probability distribution of the policy identifier as input, and retrieves and outputs the corresponding structured dialogue behavior instruction from a preset policy identifier-instruction element mapping table through a lookup operation. The instruction includes at least a semantic content identifier, a target sentiment label, and a non-verbal behavior code.

6. The method according to claim 1, characterized in that, The dynamic loading decision for the execution task specifically includes: Based on the interaction intent, infer whether the information includes a preset cognitive activity guidance intent, and combine this with whether the social dialogue process is in a preset stable interaction phase to generate a binary task trigger flag. If the task trigger flag is true, then the following operations are performed: a) The emotional state classification information, the interaction intention inference information, the dialogue topic obtained from the real-time analysis of the current social dialogue, and the cognitive ability assessment information based on the personal ability profile are combined to form a task selection feature vector. b) Input the task selection feature vector into a preset task recommendation model, wherein the task recommendation model calculates a numerical matching score for each task in the pre-stored gamified cognitive training task library; c) Select the task with the highest matching score from the task library, load its content definition and interaction flow script into the task execution engine of the social robot, and announce the start of the task to the patient.

7. The method according to claim 6, characterized in that, The task recommendation model includes a sequentially connected feature encoding subnetwork, a matching degree calculation subnetwork, and a score output subnetwork; The feature encoding subnetwork takes the task-selected feature vector as input, performs nonlinear transformation and feature fusion through a fully connected layer, and outputs a unified high-dimensional task-user joint feature representation vector. The matching degree calculation subnetwork takes the joint feature representation vector as input and includes a learnable task feature matrix. Each row of the matrix corresponds to the feature representation of a task in the gamified cognitive training task library. The subnetwork calculates the correlation between the joint feature representation vector and the features in each row of the task feature matrix, assigns an attention weight to each task, and finally outputs a comprehensive task matching degree vector through a weighted summation method. The scoring output subnetwork takes the comprehensive task matching degree vector as input, and outputs a standardized numerical matching degree score for each task in the task library through a linear transformation layer and normalization processing.

8. The method according to claim 1, characterized in that, The collected multi-dimensional task performance data of patients specifically includes: The task execution engine generates a structured interaction log, which records the patient's response to the task stimulus, the response delay time, and the result of the correctness judgment of the response by timestamp. Within the time window of task execution, the gaze point coordinate sequence and facial motion unit intensity value sequence are extracted from the facial video stream; and the average speech rate, pause frequency and fundamental frequency profile standard deviation are extracted from the speech audio stream. The interaction log, the gaze point coordinate sequence, the facial motion unit intensity value sequence, the average speech rate value, the pause frequency, and the fundamental frequency profile standard deviation are aligned based on a unified time reference and encapsulated into a data packet with a task identifier as the multi-dimensional task performance data.

9. The method according to claim 1, characterized in that, The process of generating quantitative assessment results and updating individual competency profiles specifically includes: Based on the multi-dimensional task performance data, cognitive performance indicators, emotional response indicators, and task participation indicators are calculated. The cognitive performance indicators include average response accuracy and average reaction time. The emotional response indicators include the proportion of positive expressions based on facial motion units and arousal based on voice fundamental frequency. The task participation indicator is the ratio of the average reaction time of the current task to the individual's historical average reaction time. The cognitive performance index, emotional response index, and task participation index are input into the corresponding preset evaluation functions to generate standardized cognitive dimension scores, emotional dimension scores, and participation dimension scores. The execution timestamp and task identifier of this task, along with the scores of the cognitive dimension, emotional dimension, and participation dimension, are added as a new record to the patient's personal ability profile database. Based on the new record, the statistical characteristic values ​​of each dimension score within a preset time window are recalculated to update the personal ability profile.

10. The method according to claim 1, characterized in that, The closed-loop optimization based on the updated personal capability profile specifically includes: Read the updated cognitive dimension score and emotional dimension score from the personal ability profile, and adjust the calling priority weight of each strategy in the personalized interaction strategy library based on the score; Based on the updated emotional dimension score and engagement dimension score in the personal ability profile, the configurable parameters controlling the generation of the multimodal social interaction response are adjusted. The configurable parameters include at least the baseline value in the speech synthesis parameter mapping table and the behavior triggering conditions in the robot behavior animation library. The task selection feature vector generated during the execution of this task and the key indicator set extracted from the corresponding multi-dimensional task performance data are used to form training samples to update the parameters of the logic for implementing the dynamic loading decision of the task.