An emotion regulation system based on an elderly care service robot
By constructing a multimodal data synchronous acquisition and causal correlation analysis framework, the problems of data isolation and attribution difficulties in the emotional computing of elderly care service robots were solved, enabling accurate identification of elderly people's emotions and personalized emotional intervention, thus improving the effectiveness of emotional companionship.
Patent Information
- Application Number
- CN202610680801.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2046-05-18
AI Technical Summary
Existing elderly care service robots suffer from data isolation and attribution difficulties in affective computing. They struggle to perform temporal correlation and causal reasoning through multimodal data, resulting in decreased accuracy in assessing emotional states and an inability to precisely meet the immediate psychological needs of the elderly.
The collaborative acquisition module synchronously collects facial image data, heart rate variability (HRV) and skin conductance response (GSR) physiological parameters, and family conversation voice streams to construct an emotion causal graph. Using the causal relationship chain as a supervisory signal, an adaptive analysis model decouples static facial features from dynamic emotional features, generates the current emotional state recognition result, and generates guiding dialogue content or warning information based on the emotion causal graph.
It enables refined tracing and attribution of the source of emotions, improves the accuracy and robustness of emotion recognition, can proactively intervene and generate personalized guidance strategies, and provides robot emotional companionship with deep cognition and safety assurance.
Smart Images

Figure CN122182944B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of human-computer interaction and affective computing technology, and more specifically, to an affective regulation system based on an elderly care service robot. Background Technology
[0002] In today's home-based environment, a significant proportion of elderly people tend to express their emotions in a reserved manner. Their emotional fluctuations are often conveyed through subtle changes in facial expressions or nonverbal behaviors, and most of their emotional needs are ignored, exacerbating feelings of loneliness. Although existing elderly care service robots are gradually incorporating emotional computing functions, their technological implementation focuses mainly on monitoring single physiological indicators or programmed safety companionship. They have limitations in the depth of data processing models and the initiative of interaction, making it difficult for robots to truly understand the immediate psychological needs of the elderly and provide effective emotional intervention.
[0003] Specifically, most systems employ isolated data analysis paths, and their parallel processing architectures fail to construct a unified, multimodal data collaborative analysis framework capable of temporal correlation and causal inference. This results in the emotional regulation function of elderly care service robots remaining at a superficial level. When the system detects changes in the facial expressions of the elderly, it lacks the ability to correlate these changes with the content of simultaneous family conversations, the identities of specific family members, and their emotional tendencies in real time. Furthermore, it cannot perform cross-validation by incorporating real-time fluctuations in physiological parameters such as heart rate variability.
[0004] This makes it impossible for the system to distinguish whether a frown is caused by physical discomfort, an external event, or a family conversation, thus significantly reducing the accuracy of emotional state judgment and attribution, making it difficult for subsequent robot interactions to be accurate and appropriate. Summary of the Invention
[0005] In view of the problems existing in the prior art, the purpose of this invention is to provide an emotion regulation system based on elderly care service robots, which aims to solve the above-mentioned technical problems.
[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, an emotion regulation system based on elderly care service robots includes: The collaborative acquisition module synchronously acquires facial image data of the elderly, obtains physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR) through wearable devices, and captures the voice streams of conversations between multiple people in the home environment. The contextual attribution module is used to perform speaker recognition and semantic analysis on the dialogue speech stream, extract the identities of different family members and their dialogue keywords and emotional tendencies, use the appearance of keywords as event markers, analyze the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event, establish an emotional causal map, and record the causal relationship chain of specific family members and keywords triggering specific emotional state changes in the elderly. The emotion feature decoupling module is used to receive continuous facial image data, physiological parameters, and causal relationship chains from the emotion causal graph. Using the causal relationship chains as supervision signals, it constructs an adaptive analysis model to decouple static features caused by aging and dynamic features representing real emotions from the mixed facial image data. It also verifies the results by combining synchronous physiological parameters to generate the current emotion state recognition result. The emotion guidance module is used to query the emotion causal graph based on the current emotion state recognition results and generate corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide the emotion to a different direction. When the current emotional state identification result is a continuous and intensifying negative emotion, and the guidance of the emotion guidance module is ineffective, the early warning module generates and sends an early warning message containing specific trigger source analysis based on the causal relationship chain in the emotion causal graph.
[0007] Further: This is used for speaker recognition and semantic analysis of dialogue speech streams, extracting the identities of different family members and their dialogue keywords and emotional tendencies. The occurrence of keywords is used as event markers, and the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event are analyzed to establish an emotional causal graph, including: The system receives dialogue voice streams, facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters synchronously collected by the collaborative acquisition module. It then performs speaker identification on the dialogue voice streams to determine the identities of family members and conducts semantic analysis to extract keywords and assess their emotional tendencies. The identified keywords are marked as emotional triggering events. Facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters collected synchronously within a preset time window before and after the event are extracted to form an event association dataset. The expression change trend of facial image data and the fluctuation pattern of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters in the event association dataset are analyzed. Through time alignment, a facial and physiological temporal response pattern triggered by family member identity, keywords, and emotional tendencies is established. Based on the triggering and response patterns of multiple events, an emotional causal graph is constructed with family members as nodes, keywords as edges, emotional tendencies as edge attributes, and temporal response patterns as weights.
[0008] Further: Record the causal chain of specific family members and keywords that trigger changes in the elderly person's specific emotional state, including: From the emotional causal graph, extract multiple associated records containing the same family member node and the same keyword edge. Each record contains the emotional tendency attribute of the edge and its corresponding temporal response pattern. Multiple temporal response patterns extracted from the same trigger source are aggregated and analyzed to identify the time delay norm of facial expression changes relative to the keyword event, as well as the baseline of response intensity of physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR), and to generate an emotional response template for that trigger source. In the emotional causal graph, emotional response templates are bound to corresponding family member nodes and keyword edges. When a new voice event occurs, based on the time delay norm in the emotional response template, targeted analysis of facial image data and physiological parameters is initiated after a preset time window to capture delayed emotional responses. Based on the comparison between the results of the targeted analysis and the baseline of the reaction intensity, the emotional triggering power value of this event is quantified. The family member identity, keywords, emotional tendency attributes, the emotional triggering power value of this event, and the corresponding emotional response template identifier are encapsulated into a causal relationship chain and stored in the dynamic weight database of the emotional causal graph.
[0009] Furthermore: This involves receiving continuous facial image data, physiological parameters, and causal chains from an emotion causal graph, using these causal chains as supervisory signals to construct an adaptive analysis model, including: Using the causal relationship chains stored in the dynamic weight database of the emotional causal graph as the source of supervision signal, the facial image data subsequences within the time period of the emotional triggering event corresponding to each causal relationship chain are extracted as the input sample set for model training. Construct a three-component adaptive analysis model, which includes a feature extraction unit, a dynamic feature separation unit, and a static feature compensation unit; The feature extraction unit is configured to encode a subsequence of input facial image data into a shared feature vector sequence; The dynamic feature separation unit is configured to use the emotional triggering efficacy value recorded in the current training causal relationship chain as the regression target, and the time delay norm in the emotional response template associated with the chain as the temporal attention constraint to extract dynamic feature components related to event triggering from the shared feature vector sequence. The static feature compensation unit is configured to learn and generate a static feature baseline vector representing the long-term aging trend of an individual from facial image features extracted by the feature extraction unit outside the emotional trigger event time window.
[0010] Furthermore: Static features caused by aging and dynamic features representing true emotions are decoupled from the mixed facial image data, and verified using synchronized physiological parameters to generate current emotional state recognition results, including: The shared feature vector sequence generated by the feature extraction unit is simultaneously input into the dynamic feature separation unit and the static feature compensation unit. The dynamic feature separation unit processes the emotion triggering efficacy value and time delay norm in the causal relationship chain and outputs the dynamic emotion feature vector. The static feature compensation unit extracts the static feature baseline vector from the same shared feature vector sequence. By jointly optimizing the feature extraction unit, dynamic feature separation unit, and static feature compensation unit, and introducing a feature decoupling loss function, the correlation between the dynamic emotion feature vector and the static feature baseline vector in the feature space is lower than a preset threshold, thereby achieving feature separation. The temporal changes of dynamic emotion feature vectors are correlated with the temporal changes of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters synchronously recorded within the corresponding time period. If the correlation is higher than the verification threshold, it is used as a reward signal to optimize the dynamic feature separation unit. When applying the model, the same processing is performed on the real-time facial image data stream to obtain a real-time dynamic emotion feature vector. After the physiological parameters are synchronously verified to be valid, the feature values are mapped to a predefined emotion state space to generate the current emotion state recognition result.
[0011] Furthermore: the emotional response template includes: The time delay norm sub-template is used to store typical delay time parameters and their confidence intervals for facial expression changes relative to keyword event triggering, obtained from historical time-series response pattern statistics. The baseline sub-template for response intensity is used to store the mean and fluctuation range of the baseline response intensity of heart rate variability (HRV) and skin conductance response (GSR) calculated based on the statistical characteristics of historical physiological parameters. The emotion evolution pattern sub-template is used to store the typical emotional state evolution path triggered by the trigger source, which is obtained by clustering the temporal change trajectory of dynamic emotion feature vectors in multiple events. Template confidence metrics are used to record the number of historical events on which the emotional response template was based, the consistency measure of the temporal response pattern, and the most recent update time. Among them, the time delay norm sub-template and the response intensity baseline sub-template are used to guide the temporal attention constraints and feature extraction of the dynamic feature separation unit, and the emotion evolution pattern sub-template is used to assist the emotion guidance module in generating guidance strategies that conform to the individual's emotional change patterns.
[0012] Furthermore: This is used to query the emotional causal graph based on the current emotional state recognition results and generate corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide emotional shifts, including: Receive the current emotion state recognition result, which includes the current emotion category and intensity level. Use the current emotion category and intensity level as query conditions to search for causal relationship chains in the emotion causal graph that contain corresponding and similar emotion states as weights. From the causal relationship chain retrieved, extract the associated emotional response templates and read the emotional evolution pattern sub-templates to obtain the historical typical emotional evolution path for the trigger source. Based on the matching degree between the current emotional intensity level and the typical historical emotional evolution path, a guidance strategy is selected from the pre-set guidance strategy library. The strategies in the guidance strategy library are pre-labeled with the expected guidance goals according to the emotional state and evolution path. Combining the expected goals of the guidance strategy, the path patterns in the emotional evolution model sub-templates, and the specific family member identities and keywords in the causal relationship chain, preliminary guiding dialogue content is retrieved and generated from the dialogue corpus template library. Based on the real-time intensity of the current emotion state recognition results, the tone intensity, word intimacy, and guidance urgency parameters of the generated guiding dialogue content are dynamically adjusted. The final generated guided dialogue content will be output through the voice interaction interface and interactive dialog box of the elderly care service robot.
[0013] Furthermore: When the current emotional state identification result is a persistent and intensifying negative emotion, and the guidance from the emotion guidance module is ineffective, a warning message containing specific trigger source analysis is generated and sent based on the causal relationship chain in the emotion causal graph, including: When the current emotional state identification result is negative and the intensity continues to rise within multiple consecutive time windows, it is judged as a persistent and intensifying negative emotion. If the intensity of negative emotions is not effectively reduced within the set time after the most recent guidance session, the emotional guidance is deemed ineffective. An alert is triggered when both of the above criteria are met simultaneously. Based on current and recent negative emotional states, the emotional causal graph is traced back to retrieve all negative causal relationship chains with prominent emotional triggering power values; From the retrieved causal relationship chain, extract the family member identity, keywords, emotional triggering power value, and associated emotional response template of the trigger source; Analyze the degree of deviation between the current emotional evolution trajectory and the historical emotional evolution pattern of the trigger source, integrate the trigger source information, emotional trigger effectiveness value and deviation analysis results, and generate an early warning analysis report that includes the specific trigger source, intensity and abnormal description; The warning level is determined based on the persistence and severity of negative emotions, and the report is sent to the designated terminal.
[0014] In a second aspect, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, enable the one or more processors to implement the system.
[0015] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the system.
[0016] Compared with the prior art, the technical solution provided by the present invention has at least the following beneficial effects: 1. By constructing a multimodal data synchronous acquisition and causal association analysis framework, this study addresses the problems of data isolation and attribution difficulties in the emotional computing of existing elderly care service robots. Using keyword events in family dialogues as the temporal benchmark, it strictly correlates the statements of specific family members with the subsequent facial and physiological temporal changes of the elderly, forming an interpretable emotional causal map. This enables refined tracing and attribution of the source of emotional triggers, providing a reliable basis for subsequent interventions.
[0017] 2. An adaptive decoupling model with causal relationship chains as supervision signals is introduced to overcome the interference of static facial features caused by aging on the recognition of dynamic emotional features. The model uses the emotional triggering efficacy value in the historical causal chain as the learning target and combines it with the temporal regularity constraint feature extraction in the personalized response template. At the same time, physiological parameters are used for synchronous verification and feedback, thereby separating pure dynamic emotional features from complex facial signals and improving the accuracy and robustness of emotion recognition in the elderly population.
[0018] 3. Based on attribution and recognition capabilities, a closed-loop regulation from perception and understanding to proactive intervention is achieved. It can query emotional causal maps based on real-time emotional states, match individual historical reaction patterns, and dynamically generate situation-appropriate guidance strategies. When guidance is ineffective and emotions continue to deteriorate, the system can automatically trace the source, identify high-impact negative triggers, assess their abnormal patterns, and generate a concrete early warning report. This transforms the robot from passive monitoring to proactive emotional companionship with deep cognition, personalized interaction, and safety guarantees, truly meeting the introverted and complex emotional needs of the elderly. Attached Figure Description
[0019] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments of the invention and, together with the specification, further serve to explain the principles of the invention and enable those skilled in the art to practice and use the invention.
[0020] Figure 1 This is a schematic diagram of the module of an emotion regulation system based on an elderly care service robot provided in an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the framework of the emotional response template of the present invention. Detailed Implementation
[0021] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0022] like Figure 1 As shown, this embodiment of the invention provides an emotion regulation system based on an elderly care service robot, comprising: The collaborative acquisition module 101 synchronously acquires facial image data of the elderly, obtains physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR) through wearable devices, and captures the voice streams of conversations between multiple people in the home environment. The contextual attribution module 102 is used to perform speaker recognition and semantic analysis on the dialogue speech stream, extract the identities of different family members and their dialogue keywords and emotional tendencies, use the appearance of keywords as event markers, analyze the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event, establish an emotional causal map, and record the causal relationship chain of specific family members and keywords triggering specific emotional state changes in the elderly. The emotion feature decoupling module 103 is used to receive continuous facial image data, physiological parameters, and causal relationship chains from the emotion causal graph. Using the causal relationship chains as supervision signals, it constructs an adaptive analysis model to decouple static features caused by aging and dynamic features representing real emotions from the mixed facial image data. It also verifies the results by combining synchronous physiological parameters to generate the current emotion state recognition result. The emotion guidance module 104 is used to query the emotion causal graph based on the current emotion state recognition result and generate corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide the emotion to a different direction. When the current emotional state identification result is a continuous and intensifying negative emotion, and the guidance of the emotion guidance module is ineffective, the early warning module 105 generates and sends an early warning message containing specific trigger source analysis based on the causal relationship chain in the emotion causal graph.
[0023] In embodiments of the present invention, such as Figure 1As shown, in order to solve the technical problems in the existing technology where the emotional regulation function of robots remains at a superficial level and is difficult to conduct accurate and effective emotional intervention due to the introverted emotional expression of the elderly, the difficulty in emotional attribution, and the interference of facial features caused by aging, an emotional regulation system based on elderly care service robots is provided to solve the problem. The system includes a collaborative acquisition module 101, a contextual attribution module 102, an emotional feature decoupling module 103, an emotional guidance module 104, and an early warning module 105.
[0024] The collaborative acquisition module 101 serves as the entry point for the system's interaction with the physical world. It is responsible for real-time, synchronous acquisition of three key types of data: continuous acquisition of facial image data streams from cameras deployed in the home environment; real-time acquisition of physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR) from wearable devices worn by the elderly, such as smart bracelets in existing technologies; and acquisition of conversational voice streams from multiple people in the home environment via a microphone array. Furthermore, based on the stability of the service system and the stability of the interactions between modules, high-precision time synchronization units are integrated into the cameras, wearable devices, and microphone arrays. This assigns a unified and accurate timestamp to each image frame, each physiological parameter sampling point, and each segment of voice data, thereby creating a strictly aligned multimodal raw data stream in the time dimension. This lays a solid data foundation for subsequent temporal correlation and causal analysis. Throughout the entire data acquisition and processing process, the system adheres to the principles of privacy protection and data security. During initial deployment, the system requires explicit knowledge and authorization from the elderly and their family members. All collected raw data is processed on local devices or designated edge computing units. Raw voice, image, and physiological data are securely erased after real-time feature extraction and analysis, retaining only anonymized feature vectors and metadata, including information such as emotion categories, temporal patterns, and anonymized family relationship tags, for subsequent model updates and graph construction. This prevents the storage and leakage of personal privacy data. Furthermore, to adapt to the unique language habits and family communication contexts of the elderly, the system incorporates contextual disambiguation rules in the semantic analysis stage. This disambiguation process handles common emotional vocabulary, frequently used local expressions containing specific emotions, reminiscent keywords, and family nicknames. Combined with dialogue history and voiceprint identity, the system accurately infers the emotional tendencies of common ambiguous, indirect, or repetitive expressions used by the elderly. This significantly improves the accuracy and adaptability of emotional trigger event recognition in family settings while protecting privacy.
[0025] The contextual attribution module 102 serves as the causal analysis center of the system. It receives synchronous multimodal data streams from the collaborative acquisition module 101 and performs attribution tasks. It segments and identifies the speaker using a pre-trained voiceprint recognition model, matching the segment with a pre-registered family member voiceprint feature database to determine the speaker's identity for each speech segment. Simultaneously, it performs semantic analysis using a natural language processing model optimized for daily communication scenarios among the elderly. Keywords from a pre-defined sentiment lexicon are extracted from the dialogue, and their sentiment tendency score is calculated by considering the context, intonation, and speaker identity of each keyword. The occurrence of each identified keyword is then marked as a sentiment trigger event. Using the moment of this event as the origin, the module extracts facial image data sequences and HRV and GSR physiological parameter sequences within a preset time window from the synchronous data stream, forming a complete event-related dataset. By analyzing this dataset, the module calculates the changing trends of facial expression features before and after the event, as well as the fluctuation patterns of physiological parameters, and encodes this trigger-response relationship into a temporal response pattern. By accumulating and analyzing a large number of such events, the module constructs and continuously updates an emotional causal graph. This graph uses family members as nodes, keywords as edges, emotional tendencies as edge attributes, and temporal response patterns as edge weights, thereby formally recording and expressing the potential causal relationship that specific family members and their statements trigger specific emotional state changes in the elderly.
[0026] The emotion feature decoupling module 103 serves as the system's perception end, addressing the problem of recognizing the mixture of facial aging features and real emotional dynamic features in the elderly. It receives continuous facial image data streams, physiological parameter streams, and emotional causal maps from the context attribution module 102, particularly the causal relationship chains rich in personalized information. An adaptive analysis model is constructed and applied using the causal relationship chains as supervisory signals. The module trains the model using historical causal relationship chains and facial data samples from corresponding time periods. By introducing a feature decoupling loss function and utilizing synchronous physiological parameter fluctuations to perform correlation verification and feedback optimization on the separated dynamic feature components, the model is forced to learn to effectively separate static aging features from dynamic emotional features. In the application phase, the trained model processes the real-time facial image stream, outputs a clean dynamic emotional feature vector in real time, and performs cross-validation with real-time physiological parameters, ultimately mapping and generating the current emotional state recognition result.
[0027] The emotion guidance module 104 is the system's active interactive decision-making unit. Based on the current emotion state recognition result generated by the emotion feature decoupling module 103, and after querying the emotion causal graph, it generates a corresponding guidance strategy. Upon receiving the recognition result containing emotion category and intensity level, it uses this as a query condition to search for causal relationship chains in the emotion causal graph that have triggered the same or similar emotional states. It then extracts the bound emotion response templates from these chains and reads the emotion evolution pattern sub-templates to understand the typical historical path of emotion evolution under that trigger source. The module's built-in guidance strategy library pre-stores various guidance strategies for different emotional states and evolutionary paths, such as empathy, attention diversion, and positive memory guidance. Based on the matching degree between the current emotion intensity and the historical evolutionary path, the most suitable guidance strategy is selected. Then, combining the selected strategy's goal, the emotion evolution law, and specific triggering family member and keyword information, preliminary guiding dialogue content is retrieved and synthesized from a dialogue corpus customized for the elderly. Finally, based on the real-time intensity of emotions, the module dynamically fine-tunes parameters such as tone intensity and word intimacy of the generated content, and outputs the final natural language guidance content through the voice synthesis and playback interface or interactive dialog box of the elderly care service robot, so as to achieve the intervention purpose of positive emotions or guiding negative emotions to turn around.
[0028] The early warning module 105 serves as a safety net and a backup unit for in-depth intervention. It activates when the system detects ineffective emotional guidance and the elderly person is experiencing progressively worsening negative emotions. It continuously monitors the current emotional state identification results. If a negative emotion category persists for multiple consecutive time windows with an increasing intensity, it is determined to be a persistent and intensifying negative emotion. Simultaneously, it evaluates the intervention effect of the emotional guidance module 104. If the intensity of the negative emotion does not effectively decrease within the set time after guidance, it is determined that the guidance is ineffective. When both conditions are met, the early warning module 105 is triggered. Immediately, based on the current and recent negative emotional state, it backtracks and queries the emotional causal graph, quickly retrieving all negative causal relationship chains with high emotional triggering power values (i.e., strong negative impact). From these chains, it extracts specific trigger sources, such as which family member said what keywords, and uses associated emotional response templates to analyze whether the current emotional evolution deviates from the individual's historical norms. By integrating all the analysis results, a structured early warning analysis report is generated, which clearly identifies the core triggering source, the intensity of the impact, and abnormal emotional reactions. The early warning level is determined according to the severity. Finally, the early warning report, which contains specific and actionable information, is promptly sent to designated family members or caregivers via SMS, application push, or other means, providing key decision support for their manual intervention.
[0029] In a preferred embodiment of the present invention, the method for performing speaker recognition and semantic analysis on the dialogue speech stream, extracting the identities of different family members and their dialogue keywords and emotional tendencies, using the occurrence of keywords as event markers, analyzing the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event, and establishing an emotional causal graph includes: The system receives dialogue voice streams, facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters synchronously collected by the collaborative acquisition module. It then performs speaker identification on the dialogue voice streams to determine the identities of family members and conducts semantic analysis to extract keywords and assess their emotional tendencies. The identified keywords are marked as emotional triggering events. Facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters collected synchronously within a preset time window before and after the event are extracted to form an event association dataset. The expression change trend of facial image data and the fluctuation pattern of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters in the event association dataset are analyzed. Through time alignment, a facial and physiological temporal response pattern triggered by family member identity, keywords, and emotional tendencies is established. Based on the triggering and response patterns of multiple events, an emotional causal graph is constructed with family members as nodes, keywords as edges, emotional tendencies as edge attributes, and temporal response patterns as weights.
[0030] As attached Figure 2As shown, in this embodiment of the invention, the system first receives dialogue voice stream, facial image data stream, and heart rate variability (HRV) and skin conductance response (GSR) physiological parameter streams, all synchronously acquired and timestamped by the collaborative acquisition module 101. For the dialogue voice stream, the module performs speaker recognition based on a pre-registered family member voiceprint feature database. This database is established during system initialization. It extracts acoustic features (Mel-frequency cepstral coefficients, MFCC) from voice samples of each family member reading standard sentences and generates corresponding voiceprint feature vectors, which are stored locally in encrypted form. In actual processing, the module performs endpoint detection and segmentation on the continuous voice stream. It extracts the same acoustic features from each valid voice segment and calculates similarity with the vectors in the feature database. When the similarity exceeds a preset threshold, the speaker's identity for that segment is determined, and the corresponding family member identity label is assigned to that segment. Specifically, the context attribution module 102 first performs endpoint detection and segmentation on the received speech stream to obtain a series of audio segments that may contain valid speech. For each segment, the context attribution module 102 performs acoustic feature extraction, using Mel frequency cepstral coefficients as feature representation. The calculation process includes pre-emphasizing the audio signal to balance the spectrum, then framing and adding a Hamming window to reduce spectral leakage. A fast Fourier transform is performed on each frame of the signal to obtain the spectrum. The spectrum is then passed through a set of Mel-scale triangular filter banks that simulate the characteristics of human hearing. The logarithmic energy of the output of each filter is taken. Finally, a discrete cosine transform is performed on the obtained logarithmic Mel spectrum, and the first few coefficients are taken to form a feature vector. This vector represents the short-time spectral envelope feature of the speech in that frame. The feature vector sequence of all frames of an audio segment constitutes the acoustic feature representation of that segment.
[0031] The core of speaker recognition is to compare the extracted features of unknown speech segments with a pre-built family member voiceprint feature database. This database is established during system initialization. After obtaining explicit authorization, each family member is required to read a standard text aloud in a relatively quiet home environment. The context attribution module 102 uses the exact same feature extraction process to process the registered speech segment, obtaining its feature vector sequence. Then, through voiceprint modeling, a compact and discriminative voiceprint feature model is generated for the member and securely stored locally. During recognition, the context attribution module 102 calculates the similarity between the feature representation of the speech segment to be recognized and the voiceprint model of each family member in the database. Specifically, the context attribution module 102 converts the feature sequence of the speech segment to be recognized into a fixed-dimensional overall representation vector, while the voiceprint model of each registered member in the database also corresponds to a representation vector of the same dimension.
[0032] The similarity calculation is accomplished through the following steps: First, the dot product between the two representation vectors is calculated, which is the sum of the product of the corresponding dimension values. Second, the Euclidean norm of each representation vector is calculated, which is the square root of the sum of the squares of all its dimension values. Finally, the dot product result is divided by the product of the two norms to obtain a value between -1 and +1. This value is the cosine similarity score. The closer the score is to +1, the more consistent the directions of the two vectors are, and the more similar their voice features are. The context attribution module 102 presets a similarity threshold. When the similarity score between the speech to be identified and a member's voiceprint model in the database exceeds the threshold for the first time, the speaker of the speech segment is determined to be a family member, and the corresponding identity label is assigned to the speech segment. If the similarity with all registered models is below the threshold, the segment is marked as an unknown speaker. This identification and labeling process is completed entirely locally. The original speech and intermediate features are safely erased after processing, and only the identity label is retained for subsequent analysis. Regarding the preset similarity threshold of the context attribution module 102, specifically during the registration phase, the context attribution module 102 uses these labeled registered voice samples for cross-validation, calculates the similarity between multiple registered voice samples of each member, and calculates their average value. Next, it calculates the similarity between the registered voice samples of different members and calculates their average value. Due to the effectiveness of voiceprints, intra-class similarity is usually significantly higher than inter-class similarity. The module's final recognition threshold is set at a value between the average intra-class similarity and the average inter-class similarity, for example, the midpoint between the two averages, with a small bias added based on experience. This ensures high-accuracy recognition while maintaining sufficient distinguishability of family members' voices, and avoids misidentifying environmental noise or other unknown sounds as family members. This threshold setting process is automatically completed and fixed during system initialization, ensuring the reliability and adaptability of the subsequent real-time recognition phase.
[0033] Meanwhile, the contextual attribution module 102 performs semantic analysis in parallel to extract keywords and assess their sentiment tendencies. This analysis does not use a general natural language processing model, but relies on a constructed lexicon of daily emotional vocabularies of the elderly and accompanying contextual disambiguation rules.
[0034] The contextual attribution module comprises multiple layers: the first layer consists of general emotional vocabulary, such as happiness, worry, and loneliness; the second layer consists of frequently used local expressions or everyday language by the elderly that convey specific emotions, such as feeling uneasy or comfortable; and the third layer consists of recollective keywords and family nicknames closely related to the family and the elderly person's personal history, such as children's nicknames, deceased spouse's title, and specific place names. Each word is pre-labeled with a basic emotional polarity, specifically divided into positive, negative, neutral, and intensity cardinal values. During semantic analysis, the module first matches the vocabulary in the speech-to-text result to identify keywords. The key emotional tendency assessment step involves the module dynamically correcting and quantifying the basic emotional value of the keywords by combining three layers of contextual information: first, the immediate context of the dialogue, analyzing the sentence structure and conjunctions before and after the keywords; second, the identified speaker's identity, because the same sentence, such as being busy at work, may evoke different emotional responses when spoken by different children; and third, the intonation characteristics of the speech segment, such as changes in pitch and speech rate. The context attribution module 102 integrates three contextual features—instantaneous dialogue context, speaker identity, and intonation features—with keyword-based sentiment through a pre-trained scoring model built on a multilayer perceptron structure. Before deployment, this scoring model underwent supervised training using a labeled dataset of elderly family dialogues to learn the mapping relationship between keyword sentiment tendencies in different contexts. The model ultimately outputs a sentiment tendency score between -1 and 1, quantifying the expected emotional impact of the keyword on the elderly in this specific context. Specifically, the pre-trained scoring model employs a three-input, one-output neural network structure. The input layer receives three types of feature vectors: extracting contextual semantic features, embedding speaker identity encoding vectors, and intonation feature vectors containing statistics such as pitch, speech rate, and energy. These three types of features are then concatenated and non-linearly fused through two fully connected layers. The Tanh activation function is used, and the output sentiment tendency score ranges from -1 to 1. The training objective is to minimize the mean squared error between the predicted score and the labeled score.
[0035] Once a keyword is identified and its sentiment score is calculated, the context attribution module 102 immediately marks the moment when the keyword was identified, provided by a synchronization timestamp, as an emotional trigger event. Then, the context attribution module 102 uses this event moment as the time origin and defines the data analysis window according to the system's preset, phased evolution analysis strategy. Specifically, in the early stages of system deployment or before a reliable personalized emotional response template has been formed, a general fixed time window strategy is adopted. This strategy is based on the general reaction laws of emotional physiology. For example, by default, data is extracted from 5 seconds before the event to 15 seconds after the event. This window period aims to cover most immediate facial reactions, which usually have short delays, as well as autonomic nervous system reactions, such as HRV and GSR changes, which usually have delays of several seconds to more than ten seconds. As the system runs, the contextual attribution module 102 generates personalized emotional response templates for specific triggers. These specific triggers can be configured as combinations such as family member A plus keyword B to pinpoint the trigger. The time delay norm sub-template within the emotional response template tracks the typical delay time of facial expression changes in response to such events by the elderly person. Once an emotional response template for a particular trigger has been established and its confidence level reaches a threshold, the analysis strategy will be automatically optimized. For subsequent new events from the same trigger, the system will dynamically adjust the analysis window based on the time delay norm in the template. For example, if the template shows that the elderly person's reaction to their children mentioning being busy at work typically begins after 8 seconds and lasts for approximately 20 seconds, the system may adjust the analysis window to 5 to 25 seconds after the event to more accurately capture their personalized emotional response trajectory.
[0036] After determining the window based on the above strategy, the module extracts all facial image frame sequences and physiological parameter sampling sequences of HRV and GSR within the window from the synchronized data stream, forming an event-related dataset. During the process, strict time alignment is ensured, that is, each frame image and each physiological parameter sampling point is aligned with the event time through its timestamp, ensuring that the subsequent analysis is of multimodal responses within the same time period.
[0037] Subsequently, the module analyzes the event-related dataset to establish facial and physiological temporal response patterns. For facial image data sequences, the context attribution module 102 uses a pre-trained facial expression recognition network, i.e., a convolutional neural network-based expression feature extractor, to process each frame of the image, outputting a multi-dimensional feature vector representing the expression state. By analyzing the changes in this feature vector sequence within the time window before and after the event, the context attribution module 102 calculates its gradient of change after the event, paying particular attention to changes in emotion-related feature dimensions such as the depth of brow wrinkles, the curvature of the corners of the mouth, and eye muscle activity. For physiological parameter sequences, the module calculates the HRV parameter, which is the root mean square (RMSSD) of the difference between adjacent normal heart rate intervals, to characterize the regulatory activity of the autonomic nervous system, its numerical distribution and trend within the window, and simultaneously calculates the amplitude and frequency of the skin conductance response (SCR) in the GSR signal. The final contextual attribution module 102 encodes the trajectory of facial expression features into a set of feature changes, which are combined with statistical features of HRV and GSR, such as the mean, variance, and changes in specific indicators, to form a unified, structured data vector. This vector is the temporal response pattern corresponding to the event, recording the temporal response of the elderly in facial expressions and physiological indicators when a specific family member A says a specific keyword B with a specific emotional tendency, such as a rating of C.
[0038] Through continuous operation, the module accumulates a large number of such trigger-response event instances. Based on these instances, the module constructs and continuously updates an emotional causal graph, where each node represents a family member, each directed edge points from a family member node to a keyword, and the edge itself is the identifier of the keyword. The attributes stored on the edge include the typical emotional tendency score of the keyword in the context of this family member. Most importantly, each edge is associated with one or more temporal response patterns as its weight or attribute set.
[0039] During system operation, the described sentiment tendency exists on two levels: real-time evaluation values and typical values in the sentiment graph. Emotional tendency evaluation is a real-time perception process. For each specific dialogue event, it calculates an immediate and specific sentiment tendency score, reflecting the emotional tone of the statement in the current context. The typical sentiment tendency score stored on the edges of the sentiment causal graph is a statistically summarized knowledge attribute. It is generated by aggregating and analyzing multiple historical values accumulated under the same trigger source (i.e., the same family member node and the same keyword edge), calculating a weighted average, and evaluating its confidence interval. It represents the long-term, stable sentiment tendency pattern exhibited by the trigger source and is a generalized cognition learned by the system from a large number of specific events. The relationship between the two is that real-time evaluation is the original data source and dynamic update basis for the typical score, while the typical score, as knowledge fixed in the graph, provides a stable reference benchmark for the system to understand personalized sentiment triggering patterns and identify abnormal deviations in real-time scores. The system can be understood as a family observer. The real-time emotional tendency assessment is like the observer, recording on the spot, such as when the son mentioned work and sounded somewhat anxious. This is a specific, contextualized observation note. The typical emotional tendency score in the graph is the pattern summarized by the observer after reviewing notes from the past few months or even years. For example, based on the environment of sounding somewhat anxious when mentioning work, it was found that whenever the son talked about work-related topics, he would show anxiety seven or eight times out of ten. The pattern summarized in this way is the typical score. It does not depend on a specific conversation, but reveals a recurring emotional interaction pattern among family members. By refining countless on-the-spot notes into such family emotional patterns, we can understand the emotional world of the elderly more deeply and intelligently, and provide predictive care and adjustment.
[0040] In essence, the emotional causal graph elevates discrete event data into a structured causal knowledge network, revealing the potential causal relationships that trigger specific physiological and facial expression response patterns in the elderly through the specific remarks of different family members. In other words, nodes are weighted as they are triggered, providing a computable and queryable knowledge base for understanding the personalized sources of emotional fluctuations in the elderly.
[0041] In a preferred embodiment of the present invention, recording the causal chain of specific family members and keywords triggering specific emotional state changes in the elderly person includes: From the emotional causal graph, extract multiple associated records containing the same family member node and the same keyword edge. Each record contains the emotional tendency attribute of the edge and its corresponding temporal response pattern. Multiple temporal response patterns extracted from the same trigger source are aggregated and analyzed to identify the time delay norm of facial expression changes relative to the keyword event, as well as the baseline of response intensity of physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR), and to generate an emotional response template for that trigger source. In the emotional causal graph, emotional response templates are bound to corresponding family member nodes and keyword edges. When a new voice event occurs, based on the time delay norm in the emotional response template, targeted analysis of facial image data and physiological parameters is initiated after a preset time window to capture delayed emotional responses. Based on the comparison between the results of the targeted analysis and the baseline of the reaction intensity, the emotional triggering power value of this event is quantified. The family member identity, keywords, emotional tendency attributes, the emotional triggering power value of this event, and the corresponding emotional response template identifier are encapsulated into a causal relationship chain and stored in the dynamic weight database of the emotional causal graph.
[0042] In this embodiment of the invention, the contextual attribution module 102 first extracts multiple historical association records containing the same family member node and the same keyword edge from the constructed emotional causal graph. For example, for the combination of the family member node "son" and the keyword edge "busy at work", all past event records marked as this trigger source are extracted. Each extracted record not only contains the real-time emotional tendency attribute value evaluated when the event occurs, but more importantly, it contains the complete temporal response pattern corresponding to the event. This pattern is a structured data vector that encodes the trajectory of changes in the facial expression features of the elderly and the fluctuation characteristics of the physiological parameters of heart rate variability (HRV) and skin conductance response (GSR) within a specific time window before and after the event. Subsequently, the contextual attribution module 102 performs convergent analysis on multiple temporal response patterns extracted from the same trigger source. Specifically, for the facial data portion, the module analyzes the offset of the moment when the facial features begin to change significantly relative to the keyword event marker moment in each pattern. By statistically analyzing all historical offsets, the mean and standard deviation are calculated to identify the typical time delay pattern of the elderly person's facial expression changes relative to the trigger of the keyword event for this specific trigger source, i.e., the time delay norm. This norm is expressed as a delay of T seconds, and the corresponding fluctuation range fluctuates and is recorded based on T seconds.
[0043] For the physiological parameters, the contextual attribution module 102 calculates the mean amplitude of the heart rate variability (HRV) index within the response window and the mean amplitude of the effective response peak in the skin conductance response (GSR) signal in all historical time-series response patterns, and statistically analyzes their distribution range. This determines the baseline of the response intensity of HRV and GSR physiological parameters to this trigger source. This baseline is a statistical description including the central value and normal fluctuation range, based on the identified time delay norm and response intensity baseline. Specifically, for the GSR signal, the contextual attribution module 102 first preprocesses the raw GSR signal within the post-event window of each historical time-series response pattern, including using low-pass filtering to eliminate high-frequency noise and performing baseline correction to remove slow drift. Subsequently, the context attribution module 102 employs a peak detection algorithm based on amplitude-to-duration thresholds to identify valid response peaks. For a peak to be considered valid, it must simultaneously satisfy the following conditions: its initial amplitude rises above a pre-set static threshold, and the duration from the initial point to the peak peak is within the typical emotional skin conductance response time range. For each detected valid peak, its amplitude value is calculated, which is the difference between the peak peak amplitude and the peak initial amplitude. Finally, the context attribution module 102 calculates the arithmetic mean of the amplitudes of all valid peaks across all historical patterns, as the mean amplitude of the GSR response under that trigger source. Simultaneously, by calculating the standard deviation and percentiles of these amplitude values, its distribution range is determined, thus comprehensively describing the statistical characteristics of the response intensity. Regarding the heart rate variability (HRV) index, the context attribution module 102 focuses on time-domain indicators that reflect rapid regulation of the autonomic nervous system, specifically the root mean square of the difference between adjacent normal heartbeat intervals. The contextual attribution module 102 extracts the RR interval sequence within the post-event response window from each historical time-series response pattern and calculates the RMSSD value within that window. This value characterizes the level of parasympathetic activity and is a sensitive indicator of emotional arousal. Subsequently, the mean of these RMSSD values across all historical patterns is calculated as the mean amplitude of HRV response variation. Similarly, the standard deviation and range of these values are calculated to describe their distribution. Finally, the contextual attribution module 102 integrates the calculated mean GSR amplitude and its distribution range, and the mean HRV-RMSSD and its distribution range, into a structured response intensity baseline. Therefore, the baseline is not a single value, but a statistical descriptive model that includes the core trend mean and the normal fluctuation range within an individual, such as the mean plus or minus one standard deviation, or the percentile range based on historical data. A complete baseline description could be that for the trigger source son to busy work, the GSR response intensity baseline is an amplitude mean of 0.2 microsiemens with a normal fluctuation range of 0.1 to 0.3 microsiemens, and the HRV response intensity baseline is an RMSSD mean of 35 milliseconds with a normal fluctuation range of 25 to 45 milliseconds.Based on this precisely quantified baseline of reaction intensity, along with the identified time delay norms, the contextual attribution module 102 can construct a personalized emotional response template that can be used for subsequent targeted analysis and comparison and trigger effectiveness calculation.
[0044] The context attribution module 102 generates an emotional response template specific to the trigger source. Structurally, this template includes the aforementioned time delay norm sub-template and response intensity baseline sub-template. Next, the context attribution module 102 binds this newly generated emotional response template to the corresponding family member node and keyword edge in the emotional causal graph, treating it as a new attribute of that edge. When a new voice event occurs in the system, and the speaker recognition and semantic analysis determine that its trigger source matches an edge of a bound emotional response template, the context attribution module 102 will no longer use a fixed analysis window. Instead, it will guide data analysis based on the time delay norm in the template. Specifically, after the keyword event occurs, the module waits for a start time calculated based on the time delay norm, for example, subtracting a preset buffer value from the typical delay time T, before initiating targeted analysis of the facial image data stream and physiological parameter stream. The start point and duration of this analysis window are dynamically adjusted according to the norm to ensure accurate capture of any delayed emotional responses the elderly person may exhibit.
[0045] After the targeted analysis is completed, the contextual attribution module 102 compares the captured facial and physiological reaction data with the baseline reaction intensity stored in the emotional reaction template. The quantification process involves calculating the deviation of the physiological reaction intensity from the baseline mean and combining it with the significance of facial expression changes. This is then synthesized using a predefined normalization formula, ultimately outputting a scalar value as the emotional triggering power value of this event. This value quantifies the actual impact of this specific statement on the elderly person's emotions. Finally, the contextual attribution module 102 encapsulates the family member identity, triggering keywords, the real-time assessed emotional tendency attribute value of this event, the calculated emotional triggering power value, and the unique identifier of the emotional reaction template used in this event into a complete causal chain. The synthesis using a predefined normalization formula is a multi-step quantitative analysis process executed by the contextual attribution module 102, aiming to generate a unitless scalar ranging from zero to one to accurately measure the actual impact of this speech event on the elderly person's emotions. This calculation relies on the results of this targeted analysis and the established emotional reaction template for this specific trigger source.
[0046] First, the contextual attribution module performs directional analysis on the synchronously acquired facial image data stream and physiological parameter stream based on the time delay norm in the template, and then begins to calculate the three core input sub-scores.
[0047] First, the context attribution module 102 calculates the skin conductance response intensity sub-score. It processes the raw skin conductance response signal within the current event's response window, performing low-pass filtering and baseline drift correction. Then, it applies a peak detection algorithm. A peak is considered a valid emotionally relevant response peak if it simultaneously meets two conditions: first, its rise exceeds a preset absolute threshold; second, its duration from start to peak falls within the typical emotional physiological response range of 0.5 to 5 seconds. The module counts all such valid peaks and calculates their average amplitude, recorded as the current GSR response amplitude. Next, the context attribution module 102 reads the historical GSR amplitude mean and standard deviation corresponding to the trigger source from the response intensity baseline sub-template of the emotional response template. The former is the average amplitude of valid peaks calculated based on all similar past events, while the latter describes the dispersion of these historical amplitudes. The GSR response intensity sub-score is the difference between the current GSR response amplitude and the historical GSR amplitude mean. The standardized value obtained by dividing the difference by the historical GSR amplitude standard deviation directly reflects the degree of deviation of the current skin electrophysiology activation level from the individual's historical norm.
[0048] Secondly, the heart rate variability response intensity sub-score is calculated. The contextual attribution module analyzes the heartbeat interval sequence within the current event response window and calculates the root mean square of the difference between adjacent heartbeat intervals in the time domain, denoted as the current HRV-RMSSD value. Similarly, the historical HRV-RMSSD mean and standard deviation corresponding to the trigger source are read from the response intensity baseline sub-template. Since emotional arousal is usually accompanied by the inhibition of parasympathetic activity, which manifests as a decrease in RMSSD value, the deviation is calculated by subtracting the current HRV-RMSSD value from the historical HRV-RMSSD mean, and then dividing this difference by the historical HRV-RMSSD standard deviation to obtain the HRV response intensity sub-score. A positive increase in this score corresponds to a stronger autonomic emotional arousal that the current event may trigger.
[0049] Third, the facial expression change saliency sub-score is calculated. The context attribution module analyzes the facial image sequence within the orientation window and outputs a feature vector sequence representing the expression state through a pre-trained expression feature extraction network. The context attribution module 102 calculates the comprehensive change of this feature vector within the reaction window relative to the resting period before the event. Considering individual static differences caused by aging, this score is not directly compared with historical data. Instead, the change calculated this time is compared with a minimum significant change threshold determined according to a general facial expression recognition model. Through a preset mapping rule, the change is classified and quantified into a level score. For example, no significant change corresponds to zero points, moderate change corresponds to 0.5 points, and drastic change corresponds to one point, representing the visual salience of the facial expression change caused by the event.
[0050] After obtaining the three sub-scores, the contextual attribution module 102 performs weighted fusion. The system pre-assigns fixed weight coefficients to the three sub-scores. These coefficients are set based on prior knowledge in the field of affective computing for the elderly, and their sum is one. Generally, given that physiological parameters are less susceptible to subjective manipulation and have higher reliability, the sum of the weights assigned to the two physiological sub-scores, namely, skin conductance response and heart rate variability, will be greater than the weight assigned to the facial expression sub-score. Weighted fusion involves multiplying each sub-score by its corresponding weight coefficient and then adding the three products to obtain a preliminary raw composite value that may exceed the range of zero to one.
[0051] Finally, the contextual attribution module 102 normalizes the original composite value using a standard S-shaped transformation function and a preset sensitivity parameter controlling the slope of the curve, smoothly mapping the original composite value to a closed interval between zero and one. The final value obtained after this transformation is the emotional triggering power value of this event. The closer this value is to one, the stronger the emotional response triggered by the speech event; the closer it is to 0.5, the closer the response intensity is to the individual's historical baseline level; and the closer it is to zero, the more likely there is abnormal emotional suppression or absence of response. This scalar value, as the core attribute for quantifying the strength of causal relationships, is encapsulated in the causal chain of this event and stored in the dynamic weight database of the emotional causal graph.
[0052] The preset absolute threshold in calculating the sub-fraction of skin conductance response intensity is a fixed threshold value used to identify effective emotion-related peaks from the skin conductance response signal. Based on physiological common sense and experimental experience in the field of emotion computing, it is typically set to 0.05 microsiemens, meaning that the instantaneous increase in skin conductance level under a clear emotional or cognitive event trigger usually exceeds this level. During system initialization or calibration, this threshold is loaded as a configurable basic parameter. In actual processing, only when the detected peak amplitude exceeds this threshold is it considered a valid physiological response triggered by an emotional event, rather than random physiological noise or minor fluctuations. It can be fine-tuned according to the different baseline skin conductance levels of elderly individuals, but its core function is to provide an objective and unified starting point for judgment, ensuring the signal quality of subsequent analysis. Similarly, the pre-defined mapping rule in calculating the saliency sub-score of facial expression changes is a lookup or judgment logic that maps the continuous numerical value of facial expression feature change to a discrete saliency level score. This is a minimum saliency threshold pre-calibrated using a large amount of general facial expression data. This threshold represents the minimum change in features that can be reliably identified as a change in expression at the image level. Specifically, the mapping rule compares the norm of the facial feature change calculated by the system with this minimum saliency threshold. If the norm is below a certain percentage of the threshold, such as 80%, it is mapped to a level score of 0, representing no significant change. If the norm is between a certain percentage of the threshold, such as 80% and the threshold itself, it is mapped to a level score of 0.5, representing a moderate change. If the norm is above the threshold, it is mapped to a level score of 1, representing a drastic change. This transforms complex facial changes coupled with individual aging characteristics into a unified, weighted, and standardized metric that can be integrated with physiological scores to achieve multimodal information synthesis.
[0053] Therefore, the causal chain not only records the source of the trigger, but also quantifies the intensity of the trigger and associates it with personalized pattern knowledge used to explain the reaction. This causal chain is stored in a dynamic weight database dedicated to the emotional causal graph, which is used for the subsequent training of the emotional feature decoupling module 103, the query of the emotional guidance module 104, and the source analysis of the early warning module 105, thereby realizing the continuous evolution of system knowledge from discrete event statistics to personalized causal associations.
[0054] Personalized pattern knowledge refers to a set of quantitative models of emotional response patterns built by the system for specific elderly individuals through long-term learning. The direct technical carrier is the emotional response template. Each template is uniquely bound to a specific edge in the emotional causal graph—a combination of family members and keywords. A complete emotional response template contains multiple sub-templates: a time delay norm sub-template, which quantifies the typical delay time and fluctuation range of the elderly person's facial expression response to the trigger; a response intensity baseline sub-template, which stores the historical mean and normal fluctuation range of heart rate variability and skin conductance response under such events; and an emotional evolution pattern sub-template, which depicts the typical temporal path from the initial emotional state to subsequent states triggered by the trigger. These sub-templates together constitute a set of computable and queryable personalized response patterns. Furthermore, each historically or real-time generated causal relationship chain serves as an empirical case, linking and corroborating the corresponding template. Therefore, personalized pattern knowledge is essentially a summary of causal rules abstracted from discrete events by the system regarding who mentions what, how, and what emotional response will be triggered. It is the core knowledge foundation for the system to achieve accurate emotional understanding and personalized intervention. The dynamic weight database is a dedicated data management component within the system, used to persistently store and dynamically manage personalized pattern knowledge and all related process data. It primarily stores two types of key information: first, a complete record of causal relationship chains encapsulated after each event analysis, with each chain containing details of the trigger source, real-time sentiment scores, calculated emotional trigger efficacy values, and referenced template identifiers; second, a continuously evolving and updated emotional response template library. Its dynamic nature is reflected in two aspects: first, the database content continuously grows with the occurrence of new events, with each new causal relationship chain being stored immediately; second, the system uses the newly stored data to automatically trigger the recalculation and updating of various statistical parameters in the relevant emotional response templates, such as the mean of delay time and physiological response baseline parameters, and accordingly refreshes the weights of the corresponding edges in the emotional causal graph. This weight is a calculable dynamic attribute, defined as the rolling average of all historical emotional trigger efficacy values associated with that edge, used to reflect the current emotional influence intensity of the trigger source in real time. It provides a supervised labeled training sample set for the emotion feature decoupling module, provides query and strategy generation basis based on historical patterns for the emotion guidance module, and provides detailed event chains and trigger source weight information for the early warning module for source tracing analysis, supporting the closed loop of the entire system from data accumulation to knowledge evolution to application.
[0055] In a preferred embodiment of the present invention, a model is constructed to receive continuous facial image data, physiological parameters, and causal chains from an emotion causal graph, using the causal chains as supervisory signals, and to build an adaptive analysis model, including: Using the causal relationship chains stored in the dynamic weight database of the emotional causal graph as the source of supervision signal, the facial image data subsequences within the time period of the emotional triggering event corresponding to each causal relationship chain are extracted as the input sample set for model training. Construct a three-component adaptive analysis model, which includes a feature extraction unit, a dynamic feature separation unit, and a static feature compensation unit; The feature extraction unit is configured to encode a subsequence of input facial image data into a shared feature vector sequence; The dynamic feature separation unit is configured to use the emotional triggering efficacy value recorded in the current training causal relationship chain as the regression target, and the time delay norm in the emotional response template associated with the chain as the temporal attention constraint to extract dynamic feature components related to event triggering from the shared feature vector sequence. The static feature compensation unit is configured to learn and generate a static feature baseline vector representing the long-term aging trend of an individual from facial image features extracted by the feature extraction unit outside the emotional trigger event time window.
[0056] In this embodiment of the invention, the emotion feature decoupling module 103 first extracts each stored causal relationship chain record sequentially from the dynamic weight database. For each extracted causal relationship chain, the emotion feature decoupling module 103, based on the emotional triggering event time period information recorded in the chain (i.e., the timestamp of the specific voice keyword event corresponding to the chain and the duration of the personalized analysis window used), locates and extracts all consecutive facial image frames within that time period from the historically stored, strictly time-synchronized facial image data archive, forming a facial image data subsequence. Each such subsequence, together with its corresponding causal relationship chain, which includes recorded emotional triggering efficacy values and associated emotional response template identifiers, constitutes a model training sample pair.
[0057] By traversing a large number of historical causal relationship chains in the database, the emotion feature decoupling module 103 collects a massive number of such sample pairs to form the input sample set for model training. Subsequently, the emotion feature decoupling module 103 constructs a three-component adaptive analysis model, which includes a feature extraction unit, a dynamic feature separation unit, and a static feature compensation unit in its architecture.
[0058] The feature extraction unit consists of a deep convolutional neural network backbone, configured to receive each subsequence of facial image data from the input sample set. This subsequence is a tensor composed of multiple frames of images ordered by timestamps. The feature extraction unit performs layer-by-layer convolution and pooling operations on each frame to extract multi-level spatial features. It then flattens the processed feature map of each frame and maps it to a high-dimensional vector. Finally, it arranges the feature vectors corresponding to all image frames within a time window in chronological order and encodes them into a shared feature vector sequence. This sequence simultaneously carries long-term static features related to the individual in the facial images, such as wrinkle distribution and skin laxity (signs of aging), as well as dynamic features related to instantaneous events, such as texture changes caused by muscle micro-movements. The dynamic feature separation unit is configured as a network structure with temporal modeling capabilities, namely a long short-term memory network. The supervision signal for this unit comes from the emotional trigger strength value recorded in the causal relationship chain corresponding to the current training sample. This value serves as the regression target, guiding the unit to learn to extract feature components related to the intensity of emotional triggering from the shared feature vector sequence. Simultaneously, this unit accepts another key constraint: by querying the emotional response templates associated with the current causal chain, it obtains the time delay norms within them and converts these norms into a temporal attention mask or weight distribution, such as the mean of the delay time and the confidence interval. This constraint allows the dynamic feature separation unit to focus more attention on the time periods predicted based on the time delay norms that are most likely to contain emotional dynamic features when scanning the shared feature vector sequence. This enables it to more accurately separate the dynamic feature components related to event triggering from the shared sequence and output a dynamic emotional feature vector focused on emotional information.
[0059] The static feature compensation unit is configured as a network designed to learn an individual's inherent facial baseline. Its design goal is to model long-term, stable static features resulting from aging. The learning signal does not originate from a single sample's event window but is achieved through a contrastive learning network mechanism. Specifically, the emotion feature decoupling module 103 provides training samples from multiple different event time periods for the same elderly person. The static feature compensation unit is configured to learn and generate a static feature baseline vector representing the individual's long-term aging trend from facial image features extracted by the feature extraction unit outside the emotional trigger event time window—for example, features extracted from images of the elderly person in a calm state without significant daily events, or stable feature patterns independent of the time delay window from aggregating shared feature vector sequences of a large number of different event samples. This baseline vector aims to capture the inherent facial morphology unaffected by momentary emotions. Through the joint construction and training of this three-component model, the emotion feature decoupling module 103 aims to establish an analytical capability that can automatically decompose mixed facial signals into dynamic emotional features and static aging features.
[0060] Furthermore, regarding the contrastive learning mechanism in the static feature compensation unit, an online contrastive learning mechanism is used to model individual static aging characteristics. A dynamically updated personal static feature baseline pool is maintained for each elderly person. During the training initialization phase, the emotion feature decoupling module 103 first uses facial images collected from the elderly person during calm periods without specific emotional events. Facial features are extracted by the feature extraction unit and aggregated to initialize the baseline pool. In each training iteration, when processing a training sample belonging to the elderly person, the static feature compensation unit extracts feature vectors from the shared feature vector sequence of the current sample for all time steps. Simultaneously, it randomly samples or calculates an aggregated baseline vector from the elderly person's personal static feature baseline pool. Subsequently, a trainable similarity measurement network within the unit calculates the similarity score between the feature vector at each time step in the sequence and the baseline vector. The design goal is for the unit to learn to classify high-similarity features as static features, belonging to the aging baseline, while classifying low-similarity features as deviations that may contain dynamic emotions. Through this comparison, the unit can filter out parts similar to the individual baseline from the shared sequence, and fuse and enhance these similar features to output an updated static feature baseline vector that more accurately represents the individual's long-term aging trend. This baseline vector will also be used to update the individual static feature baseline pool after training, so that static knowledge can evolve gradually with the accumulation of data.
[0061] The generation and application mechanism of temporal attention constraints specifically occurs during model training. When the emotion feature decoupling module 103 processes a training sample, the dynamic feature separation unit first queries the associated emotional response template through the causal relationship chain corresponding to the sample and reads the time delay norm sub-template from it. This sub-template contains two key parameters: the mean and standard deviation of the typical delay time of facial responses, statistically derived from historical data. The emotion feature decoupling module 103 uses the event trigger time as the time origin and constructs a virtual time axis within the model that is perfectly aligned with the time steps of the shared feature vector sequence. Based on the aforementioned mean and standard deviation parameters, the module calculates and generates a soft temporal attention weight distribution. Specifically, the module calculates the weight value of each point on the time axis based on a bell-shaped distribution function centered on the typical delay time and with a width equal to its standard deviation, ensuring that the weight is highest near the typical delay time point and smoothly decays as the distance between the time point and the center point increases. Subsequently, this calculated original weight distribution is normalized to ensure that the sum of all weights is one, thus forming a directly applicable temporal attention mask. The dynamic feature separation unit introduces this mask as prior knowledge in its internal attention calculation layer, so that when the network scans and processes the shared feature vector sequence, it automatically assigns higher attention weight to the peak reaction period features predicted based on the time delay norm, thereby suppressing features far away from the region, thus achieving targeted and refined extraction of emotional dynamic features.
[0062] In a preferred embodiment of the present invention, static features caused by aging and dynamic features representing true emotions are decoupled from mixed facial image data, and verified in conjunction with synchronized physiological parameters to generate a current emotional state recognition result, including: The shared feature vector sequence generated by the feature extraction unit is simultaneously input into the dynamic feature separation unit and the static feature compensation unit. The dynamic feature separation unit processes the emotion triggering efficacy value and time delay norm in the causal relationship chain and outputs the dynamic emotion feature vector. The static feature compensation unit extracts the static feature baseline vector from the same shared feature vector sequence. By jointly optimizing the feature extraction unit, dynamic feature separation unit, and static feature compensation unit, and introducing a feature decoupling loss function, the correlation between the dynamic emotion feature vector and the static feature baseline vector in the feature space is lower than a preset threshold, thereby achieving feature separation. The temporal changes of dynamic emotion feature vectors are correlated with the temporal changes of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters synchronously recorded within the corresponding time period. If the correlation is higher than the verification threshold, it is used as a reward signal to optimize the dynamic feature separation unit. When applying the model, the same processing is performed on the real-time facial image data stream to obtain a real-time dynamic emotion feature vector. After the physiological parameters are synchronously verified to be valid, the feature values are mapped to a predefined emotion state space to generate the current emotion state recognition result.
[0063] In this embodiment of the invention, the emotion feature decoupling module 103 inputs the shared feature vector sequence generated by the feature extraction unit into the dynamic feature separation unit and the static feature compensation unit. The dynamic feature separation unit uses the emotion triggering efficacy value recorded in the causal relationship chain corresponding to the currently processed sample as the regression target, and uses the time delay norm in the emotion response template associated with the chain as the temporal attention constraint to process the shared feature vector sequence.
[0064] The time delay norm provides typical delay parameters and confidence intervals for facial expression changes relative to keyword event triggering. Based on these parameters, the dynamic feature separation unit internally constructs a soft attention weight distribution. This weight distribution is highest near the typical delay time point and gradually decays towards both sides, allowing the network to focus on the time period most likely to contain emotional dynamic responses when processing sequences. Through its internal temporal modeling structure, the dynamic feature separation unit extracts feature components related to the emotional triggering event from the weighted shared feature vector sequence and outputs a dynamic emotional feature vector focused on emotional information.
[0065] Meanwhile, the static feature compensation unit extracts a static feature baseline vector representing an individual's long-term aging trend from the same shared feature vector sequence. This is achieved through an online contrastive learning mechanism, maintaining a dynamically updated personal static feature baseline pool for each elderly person. This pool is initialized by aggregating facial features from the elderly person during resting periods. When processing each training sample, the static feature compensation unit calculates the similarity between the feature vector at each time step in the shared feature vector sequence and the baseline vector obtained from the baseline pool. This aims to identify and strengthen feature patterns that are highly consistent with the individual's long-term aging baseline, thereby outputting an updated and more accurate static feature baseline vector.
[0066] The emotion feature decoupling module achieves separation of dynamic and static features by jointly optimizing the feature extraction unit, dynamic feature separation unit, and static feature compensation unit, and introducing a feature decoupling loss function. The feature decoupling loss function maximizes the incompatibility between the dynamic emotion feature vector and the static feature baseline vector in the feature space, specifically by calculating the estimated mutual information between them and constraining this value to be below a preset threshold. During training, the parameters of the feature extraction unit, dynamic feature separation unit, and static feature compensation unit are updated simultaneously through backpropagation. This enables the feature extraction unit to learn to generate a shared feature representation that facilitates subsequent decoupling, the dynamic feature separation unit to learn to extract dynamic components related to the emotional triggering efficacy value and constrained by temporal attention, and the static feature compensation unit to learn to accurately model the individual's inherent aging baseline.
[0067] Furthermore, the emotion feature decoupling module performs correlation calculations to verify the temporal changes of dynamic emotion feature vectors with the temporal changes of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters synchronously recorded within the corresponding time period. Specifically, for the event time period in the training samples, the module extracts the change curve of the dynamic emotion feature vector in the time dimension within that time period, and simultaneously extracts the synchronous HRV index, i.e., the sliding window sequence of the root mean square (RMSSD) difference between adjacent normal heartbeats, and the GSR index, i.e., the change curve of the response amplitude sequence obtained after filtering and peak detection of the skin conductance response signal. The Pearson correlation coefficient between the dynamic emotion feature curve and the HRV change curve is calculated, and the Pearson correlation coefficient between the dynamic emotion feature curve and the GSR change curve is also calculated. Specifically, when the emotion feature decoupling module 103 performs correlation verification calculations, it first operates on a time period of emotional triggering events clearly defined by a causal chain. The start and end points of this time period are determined by specific data provided by the context attribution module 102, for example, starting 5 seconds before the keyword event occurs and ending at the predicted delayed reaction end period based on the individual elderly person's emotional response template, such as 25 seconds after the event. Within this time period, the emotion feature decoupling module 103 has processed the synchronously acquired facial image sequence through an adaptive analysis model, outputting a series of dynamic emotion feature vectors arranged in chronological order. Each vector is a multidimensional data set. For time-series analysis, the module uses a pre-trained fixed weight matrix to linearly combine these high-dimensional vectors with... Dimensionality reduction transforms the image into a single scalar value representing the overall emotional arousal level. This process is performed sequentially for each frame of the facial image within a time period, thereby generating a sequence of overall emotional arousal values that perfectly corresponds to the time point of the image frame. Specifically, in the emotion feature decoupling module 103, after processing each frame of the facial image, a high-dimensional dynamic emotion feature vector is output, containing data from multiple dimensions that jointly encode the dynamic information related to emotion in the facial expression at the current moment. Then, each dimension value of the dynamic emotion feature vector is multiplied by the corresponding preset weight coefficient in the weight matrix. Finally, these multiplication results across all dimensions are summed to obtain a preliminary weighted sum value. Subsequently, a preset fixed bias value is added to this weighted sum value.
[0068] The adaptive analysis model of the emotion feature decoupling module 103 separates pure dynamic emotion features from facial images. All supervision signals and constraints required for model training are uniquely derived from the emotion causal graph and its causal relationship chain. Specifically, each iteration of the training process relies on a specific causal relationship chain. The causal relationship chain provides two key supervision information: first, the emotion triggering efficacy value, which is a scalar target value quantifying the intensity of the impact of a specific event on the elderly person's emotions; and second, the time delay norm in the emotional response template associated with the causal relationship chain, which provides temporal attention constraints for feature extraction. Driven by the supervision signals, at the beginning of training, the weight coefficients and bias values are randomly initialized. During training, the system selects a causal relationship chain and imports facial image data from the corresponding time period. After processing by the feature extraction unit and the dynamic feature separation unit, a preliminary dynamic emotion feature vector is generated. This vector is obtained by linearly calculating the weight coefficients and bias values at that time, initially with random values, to obtain a predicted arousal scalar. The system then calculates two losses. First, it compares the predicted arousal scalar with the actual recorded emotional triggering power values in the causal chain to calculate the regression error. Then, it calculates the correlation between the temporal change of this predicted arousal scalar and the temporal changes of HRV and GSR physiological parameters collected synchronously during the same time period. If the correlation is insufficient, a reward signal is generated. These together constitute a total training objective. Using the backpropagation algorithm, the above total loss is used to calculate the adjustment direction and magnitude of the current weight coefficients and bias values of the arousal mapper. The Adam optimizer updates these parameters accordingly, enabling the next output to be a scalar that is closer to the actual emotional triggering power value and more synchronized with physiological signals. This process is repeated on a dataset of millions of historical causal chains until the model performance stabilizes.
[0069] In the model training and application process, the high-dimensional dynamic emotion feature vector output by the dynamic feature separation unit is converted into a comprehensive scalar representing the intensity of emotional arousal. This process is entirely implemented by the publicly available module functions and data interaction of the system. Specifically, during the training phase, the dynamic feature separation unit is directly configured to optimize the regression target based on the emotional triggering efficacy value recorded in the causal relationship chain of the emotional causal graph. This allows the internal network parameters of the dynamic feature separation unit to not only learn to separate the dynamic components related to emotions from facial data during training, but also to automatically learn the mathematical mapping relationship from the high-dimensional feature space to the one-dimensional scalar space. After the model training is completed and deployed, the high-dimensional dynamic emotional feature vector generated by the dynamic feature separation unit for each frame of facial image processing can be directly calculated as a scalar of the same scale as the emotional triggering efficacy value. During real-time operation, the emotional feature decoupling module 103 calls this solidified capability to perform the determined mathematical operation on the dynamic emotional feature vector at each time point, instantly outputting the corresponding comprehensive emotional arousal scalar, forming a time series, and using it to perform correlation calculation and verification with the HRV and GSR physiological parameter sequences obtained by the collaborative acquisition module 101, as well as as input to the subsequent classification and regression models in the emotional feature decoupling module 103, directly mapping to generate the final emotional state recognition result. The emotion feature decoupling module calls its internal pre-built emotion classification and intensity regression model to process the dynamic emotion feature vector. During the training phase, the emotion classification and intensity regression model is trained with the causal relationship chain in the historical emotion causal graph as the supervision signal, the emotion tendency attribute recorded in the causal relationship chain as the emotion category label, and the emotion triggering power value as the intensity regression target. The model analyzes the input dynamic emotion feature vector and outputs the corresponding emotion category label and intensity level value. The two together constitute the structured current emotion state recognition result.
[0070] Once the model training is complete and passes validation, the iterative updates of the weight coefficients and bias values cease. At this point, they have converged to a fixed set of optimal values. These values are the best solutions automatically found by the system using its inherent emotional causal graph data through a public supervised learning process. They are stored in the storage space of the emotional feature decoupling module 103 and become an inherent part of the system. Therefore, the preset weight coefficients and preset fixed bias values are direct products of the system's self-learning and optimization using its own core database, namely the emotional causal graph. In application, these parameters are used as pre-trained, fixed parameters in the calculation and remain unchanged. The calculation ultimately yields a single scalar value, the comprehensive emotional arousal score. This value is a quantity with direction and magnitude; its positive or negative sign is typically used to distinguish the valence of emotions. A positive increase is associated with positive emotional arousal, while a negative increase is associated with negative emotional arousal. The absolute value represents the intensity of emotional arousal. Finally, the emotional feature decoupling module 103 performs the complete calculation process for each frame of facial image within the emotional triggering event time period, in chronological order, thereby generating a sequence of comprehensive emotional arousal score scalar values that are completely synchronized with the timestamps of the facial image frames and arranged sequentially. This sequence serves as the direct data basis for subsequent temporal correlation calculations with the physiological parameter sequence and ultimately for mapping and generating specific emotional state recognition results.
[0071] Simultaneously, the module extracts raw physiological signals within the same time period from the synchronous data stream of the collaborative acquisition module 101. For heart rate variability (HRV) data, the module obtains the interval data of each heartbeat with precise timestamps. The emotion feature decoupling module 103 first cleans these raw interval data, removing abnormal data points that clearly exceed the normal physiological range. In actual operation, this is determined according to the actual situation, such as abnormal data points less than 0.3 seconds or greater than 2.0 seconds. Subsequently, the emotion feature decoupling module 103 uses a sliding window analysis method to calculate the RMSSD index of HRV. Specifically, the time length of each analysis window is set to 5 seconds, and the window slides forward once every 1 second. Within each 5-second window, the emotion feature decoupling module 103 first calculates the difference between all adjacent normal heartbeat intervals, then calculates the mean of the squares of these differences, and finally performs a square root operation on the mean to obtain the RMSSD value of that window. This calculated RMSSD value is assigned to the timestamp corresponding to the center point of the analysis window. By performing this sliding window calculation and assignment over the entire event time period, the module obtains a set of RMSSD values that are non-uniformly distributed over time. In order to make this physiological data sequence comparable to the above emotional arousal scalar sequence at the same time point, the emotion feature decoupling module 103 uses a linear interpolation algorithm to transform these non-uniform RMSSD values to uniform time point coordinates that are completely consistent with the emotion data sequence, and finally forms an HRV index sequence that corresponds one-to-one with the time point of the emotion sequence.
[0072] For GSR (Geodermal Response) data, the module obtains the raw skin conductance level signal, with a data acquisition frequency of 10 times per second. The emotion feature decoupling module 103 first filters this continuous signal, using a fourth-order Butterworth digital low-pass filter with a cutoff frequency of 1 Hz to remove high-frequency noise. Next, baseline correction is performed, subtracting the minimum value from the signal values over the entire time period to eliminate individual differences in baseline skin conductance levels. Then, the emotion feature decoupling module 103 performs peak detection on this processed signal to identify specific skin conductance activity peaks associated with emotional responses.
[0073] For a peak to be considered valid, three conditions must be met simultaneously: the rise of the first peak from its starting point to its apex must exceed 0.05 microsiemens; the duration of the second peak from its starting point to its apex must be between 0.5 and 5.0 seconds; and the starting point of the third peak must be at least 2 seconds apart from the starting point of the previous valid peak. For each identified valid peak, its amplitude, i.e., the difference between the apex value and the starting point value, is recorded at the time point when the peak appears. For time points where no valid peak is detected, the amplitude value is recorded as zero. Thus, the emotion feature decoupling module 103 generates an amplitude sequence within the event time period with the same sampling rate as the original but sparse values. Similarly, to align this sequence with the time points of the emotion sequence, the emotion feature decoupling module 103 uses an interpolation method that preserves peak features to resample the sparse sequence onto a uniform time grid of 10 points per second, thereby generating a GSR response amplitude sequence that corresponds one-to-one with the time points of the emotion sequence.
[0074] After obtaining three sequences of identical duration and strictly aligned time points—the emotional arousal sequence, the HRV index sequence, and the GSR amplitude sequence—the emotion feature decoupling module 103 begins calculating the Pearson correlation coefficient. The process of calculating the correlation coefficient between the emotion sequence and the HRV sequence involves first calculating the arithmetic mean of all values in the emotion sequence and the arithmetic mean of all values in the HRV sequence. Next, for each time point in the sequence, the difference between the emotion value at that point and its sequence mean is calculated, along with the difference between the HRV value at that point and its sequence mean. These two differences at each time point are then multiplied, and the products from all time points are summed. Subsequently, the sum of squares of the differences between each value in the emotion sequence and its mean, and the sum of squares of the differences between each value in the HRV sequence and its mean, are calculated, and the square root of these two sums is taken to obtain the numerator of the product of two standard deviations. Finally, the sum of the previously calculated product of differences is divided by the numerator of this product of standard deviations, and the result is the Pearson correlation coefficient between the emotion sequence and the HRV sequence. Using the exact same calculation steps, the emotion feature decoupling module 103 can calculate the Pearson correlation coefficient between the emotion sequence and the GSR sequence.
[0075] Finally, the emotion feature decoupling module 103 compares the two calculated correlation coefficients with a preset verification threshold, which is set to 0.25 in this system. If at least one of the correlation coefficients between the emotion sequence and the HRV sequence, or between the emotion sequence and the GSR sequence, is greater than or equal to 0.25, the module determines that the dynamic emotion features decoupled from the facial image have a significant synchronous change relationship with the physiological response of the autonomic nervous system that is not directly controlled by subjective consciousness. This verifies the effectiveness of the extracted dynamic features, and the high correlation conclusion is used as a positive feedback signal for further optimization of the model.
[0076] If the average value of the calculated correlation coefficient is higher than the preset verification threshold, it is determined that the dynamic features extracted this time are consistent with the physiological arousal signal. The emotion feature decoupling module takes this high correlation as a positive reward signal and feeds it back to the training process of the dynamic feature separation unit through a reinforcement learning framework or directly as an additional optimization target to further optimize the physiological effectiveness of its extracted features.
[0077] When the adaptive analysis model is trained and put into application, the emotion feature decoupling module performs the same processing on the real-time acquired facial image data stream to generate the current emotion state recognition result. For each newly arrived real-time facial image data stream, the module first encodes it into a shared feature vector sequence through the feature extraction unit. Subsequently, this sequence is simultaneously fed into the pre-trained dynamic feature separation unit and static feature compensation unit. The dynamic feature separation unit applies attention constraints based on the time delay norm in the most relevant emotional response template and outputs a real-time dynamic emotion feature vector. The static feature compensation unit outputs the static feature baseline vector of the current individual as a reference. To ensure the reliability of the real-time results, the emotion feature decoupling module 103 simultaneously acquires the real-time heart rate variability (HRV) and skin conductance response (GSR) physiological parameter streams within the corresponding time period and calculates the correlation between the temporal changes of the real-time dynamic emotion feature vector and the temporal changes of these physiological parameters. If the correlation is higher than the verification threshold in the application stage, the dynamic feature extraction is deemed valid. Finally, the emotion feature decoupling module 103 inputs the physiologically validated real-time dynamic emotion feature vector into a predefined emotion classification and intensity regression model. The intensity regression model is a fully connected neural network that maps dynamic emotion feature vectors to a predefined emotion state space. This space consists of multiple discrete emotion categories, such as calm, joy, sadness, and anxiety, along with continuous intensity levels for each category. The final output is a structured current emotion state recognition result, including the identified dominant emotion category and its quantified intensity level. The training and operation of the predefined emotion classification and intensity regression model rely entirely on the system's pre-constructed emotional causal graph. The supervision signals used during the training phase are directly derived from the multimodal analysis results recorded in each causal chain of the emotional causal graph. Specifically, the training data for the emotion classification and intensity regression model consists of triplets from the historical data stream. The input is the dynamic emotion feature vector generated by the emotion feature decoupling module 103 in the corresponding time period. The supervision labels come from the associated causal chains, including two parts: emotional tendency attributes derived from the context attribution module 102's analysis of triggering keywords (positive and negative emotions). These attributes serve as the emotion category, such as joy or sadness, as the classification supervision signal. The other part is the emotion triggering efficacy value recorded in the same chain, quantifying the overall physiological and facial expression intensity. In the design of emotion classification and intensity regression models, the intensity level under the emotion category is not an independent label, but is defined by dividing the emotion triggering power value into intervals.For example, if the emotion is identified as sadness, its emotional trigger power value is mapped to mild, moderate, and severe intensity levels. Therefore, the emotion classification and intensity regression model is essentially trained as a multi-task network. One task is to predict the emotional tendency category based on dynamic features, and the other task is to regress the continuous emotional trigger power value in parallel. In real-time applications, the model infers from the physiologically validated dynamic feature vectors and outputs the predicted emotion category and continuous power value. Then, according to the preset category and intensity mapping rules, it finally generates a structured recognition result, such as sadness plus the keyword "moderate". The training data and mapping rules of the entire model are all derived from the core emotion causal graph of the system.
[0078] In a preferred embodiment of the present invention, the emotional response template includes: The time delay norm sub-template is used to store typical delay time parameters and their confidence intervals for facial expression changes relative to keyword event triggering, obtained from historical time-series response pattern statistics. The baseline sub-template for response intensity is used to store the mean and fluctuation range of the baseline response intensity of heart rate variability (HRV) and skin conductance response (GSR) calculated based on the statistical characteristics of historical physiological parameters. The emotion evolution pattern sub-template is used to store the typical emotional state evolution path triggered by the trigger source, which is obtained by clustering the temporal change trajectory of dynamic emotion feature vectors in multiple events. Template confidence metrics are used to record the number of historical events on which the emotional response template was based, the consistency measure of the temporal response pattern, and the most recent update time. Among them, the time delay norm sub-template and the response intensity baseline sub-template are used to guide the temporal attention constraints and feature extraction of the dynamic feature separation unit, and the emotion evolution pattern sub-template is used to assist the emotion guidance module in generating guidance strategies that conform to the individual's emotional change patterns.
[0079] In this embodiment of the invention, the emotion response template is generated and maintained by the context attribution module 102. Its structure includes four core parts: time delay norm sub-template, response intensity baseline sub-template, emotion evolution pattern sub-template, and template confidence index. The time delay norm sub-template is used to store the typical delay time parameters of facial expression changes relative to the triggering of keyword events, and their confidence intervals, which are obtained from the statistics of historical time-series response patterns.
[0080] The specific generation process involves the context attribution module 102 extracting the offset of the time point when facial expression features begin to change significantly relative to the keyword event marker time for all historical temporal response patterns under the same trigger source, such as the combination of a specific family member node and a specific keyword edge, and forming an offset set.
[0081] The contextual attribution module 102 calculates the arithmetic mean of the set as the typical delay time and calculates its standard deviation. The typical delay time parameter stored in the final time delay norm sub-template is this average value, and its confidence interval is defined as the range formed by plus or minus one standard deviation of this average value. For example, if the calculated average delay time is 8 seconds and the standard deviation is 2 seconds, then the typical delay time parameter is 8 seconds, and the confidence interval is 6 to 10 seconds.
[0082] The baseline response intensity sub-template stores the mean and fluctuation range of the baseline response intensity of heart rate variability (HRV) and skin conductance response (GSR) calculated based on the statistical characteristics of historical physiological parameters. For HRV, the context attribution module 102 extracts the root mean square (RMSSD) value of the difference between adjacent normal heartbeat intervals calculated within the response window after each pattern event from all historical temporal response patterns of the same trigger source, forming a set of RMSSD values. The context attribution module 102 calculates the arithmetic mean of this set as the mean baseline response intensity of HRV and calculates its standard deviation. The fluctuation range of HRV is defined as plus or minus one standard deviation of this mean. For GSR, the context attribution module 102 extracts the amplitude values of all valid skin conductance responses identified by the peak detection algorithm within the response window after each pattern event from all historical temporal response patterns of the same trigger source, forming a set of GSR amplitudes.
[0083] The context attribution module 102 calculates the arithmetic mean of the set as the baseline response intensity mean of the GSR, and calculates its standard deviation. The fluctuation range of the GSR is also defined as the mean plus or minus one standard deviation. For example, the response intensity baseline sub-template under a trigger source is stored as HRV baseline response intensity mean of 40 milliseconds with a fluctuation range of 30 to 50 milliseconds, and GSR baseline response intensity mean of 0.15 micro Siemens with a fluctuation range of 0.05 to 0.25 micro Siemens.
[0084] The emotion evolution pattern sub-template is used to store the typical emotional state evolution path triggered by the trigger source, obtained by clustering the temporal change trajectory of the dynamic emotion feature vectors output by the emotion feature decoupling module 103 in multiple events. Specifically, the context attribution module 102 extracts all historical causal relationship chains corresponding to the same trigger source from the dynamic weight database, and obtains the sequence of dynamic emotion feature vectors generated by the emotion feature decoupling module 103 within the time period of the emotion triggering event associated with each chain. The context attribution module 102 first preprocesses and aligns these sequences in time, taking the moment of the emotional triggering event recorded in each causal relationship chain as the time origin, that is, taking this point as 0 seconds. Then, it extracts a fixed-length window after the event, such as the dynamic emotional feature vector sequence within 0 to 30 seconds. In order to perform clustering, the context attribution module 102 reduces each high-dimensional dynamic emotional feature vector to a two-dimensional or three-dimensional observation point through a fixed projection matrix. This matrix is determined together with the emotional feature decoupling module 103 during training and is used to map the features to a low-dimensional emotional state subspace. All observation points in the sequence within the time window are connected in order to form a trajectory representing a single emotional evolution.
[0085] Subsequently, the context attribution module 102 employs a hierarchical clustering algorithm based on dynamic time warping distance (DTW distance) to perform clustering analysis on these trajectories. First, it calculates the DTW distance between every two trajectories. DTW distance effectively measures the similarity between two sequences that exhibit nonlinear scaling or velocity differences on the time axis. Then, it uses a hierarchical clustering method with average linkage, initially treating each trajectory as a separate cluster. Iteratively, it merges the two clusters with the smallest DTW distance until all trajectories are merged into a preset number of clusters, such as 2 to 4. After iteration, the context attribution module 102 selects the trajectory with the smallest average DTW distance to all other trajectories within each final cluster as the center trajectory of that cluster. Finally, the context attribution module 102 encodes and stores these center trajectories in an emotion evolution pattern sub-template. For example, it records the coordinates and times of key inflection points on the trajectory. Each such center trajectory represents a typical emotional state evolution path that the trigger source may induce, such as a rapid rise to a high level and sustained high level, or a slow rise followed by a rapid fall back to calm.
[0086] The template confidence index is used to record the number of historical events on which the emotional response template is based, the consistency measure of the temporal response pattern, and the most recent update time. The number of historical events is directly recorded as the total number of causal relationship chains encapsulated under the trigger source. The consistency measure of the temporal response pattern is obtained by calculating the variance or coefficient of variation of facial and physiological response features in all historical temporal response patterns under the trigger source. The lower the dispersion, the higher the consistency. The most recent update time is recorded as the system timestamp when the template was last modified or refreshed by the context attribution module 102. The time delay norm sub-template and the response intensity baseline sub-template are used to guide the temporal attention constraints and feature extraction of the dynamic feature separation unit in the emotion feature decoupling module 103. The emotion evolution pattern sub-template is used to assist the emotion guidance module 104 in generating guidance strategies that conform to the individual's emotional change patterns.
[0087] In a preferred embodiment of the present invention, the method for querying an emotional causal graph based on the current emotional state recognition result and generating corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide emotional shift includes: Receive the current emotion state recognition result, which includes the current emotion category and intensity level. Use the current emotion category and intensity level as query conditions to search for causal relationship chains in the emotion causal graph that contain corresponding and similar emotion states as weights. From the causal relationship chain retrieved, extract the associated emotional response templates and read the emotional evolution pattern sub-templates to obtain the historical typical emotional evolution path for the trigger source. Based on the matching degree between the current emotional intensity level and the typical historical emotional evolution path, a guidance strategy is selected from the pre-set guidance strategy library. The strategies in the guidance strategy library are pre-labeled with the expected guidance goals according to the emotional state and evolution path. Combining the expected goals of the guidance strategy, the path patterns in the emotional evolution model sub-templates, and the specific family member identities and keywords in the causal relationship chain, preliminary guiding dialogue content is retrieved and generated from the dialogue corpus template library. Based on the real-time intensity of the current emotion state recognition results, the tone intensity, word intimacy, and guidance urgency parameters of the generated guiding dialogue content are dynamically adjusted. The final generated guided dialogue content will be output through the voice interaction interface and interactive dialog box of the elderly care service robot.
[0088] In this embodiment of the invention, the emotion guidance module 104 is used to query the emotion causal graph based on the current emotion state recognition result and generate corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide the emotion to a different direction. Specifically, the emotion guidance module 104 receives the current emotion state recognition result from the emotion feature decoupling module 103. This result is a structured data object that explicitly includes the current emotion category obtained by mapping a predefined emotion state space, such as calm, joy, sadness, or anxiety, and the current emotion intensity level quantified by an emotion classification and intensity regression model. This intensity level is quantified in this system as a continuous scalar value between 0.0 and 1.0. For example, 0.8 can represent high intensity based on the actual emotional state of the elderly person in the preset process. The emotion guidance module 104 then uses the combination of the current emotion category and the current emotion intensity level as the query condition to search the dynamic weight database of the emotion causal graph constructed and maintained by the context attribution module 102. The goal of the search is to find all historical causal relationship chains whose final emotion state at the time of their generation matches or is similar to the current query condition. This state serves as the weight attribute of the chain. The specific matching logic is as follows: First, the emotion guidance module 104 filters out causal relationship chains whose historical emotion state categories are completely consistent with the current emotion category. Second, to broaden the reference range, the emotion guidance module 104 calculates the semantic similarity between the current emotion category and the historical emotion state categories of other chains in the graph within a predefined emotion space. This similarity is obtained by calculating cosine similarity using a pre-trained word embedding model. If the similarity exceeds a preset threshold of 0.7, the chain is also included in the candidate set. For each chain in the candidate set, the emotion guidance module 104 further calculates the absolute difference between its historical emotion intensity level and its current emotion intensity level. The smaller the difference, the higher the matching score of the chain with the current context. Finally, the emotion guidance module 104 selects the top 3 causal relationship chains with the highest matching scores as the core reference for formulating subsequent guidance strategies.
[0089] A pre-trained word embedding model refers to a static word vector model trained on a large-scale general text corpus before system deployment. This model maps each word to a fixed-dimensional real-valued vector, making semantically similar words closer together in the vector space. An example of such a pre-trained model is integrated into the sentiment guidance module 104 of this system. When calculating the semantic similarity between the current emotion category and the historical emotion state category recorded in a causal relationship chain in the sentiment causal graph, the sentiment guidance module 104 first retrieves the word vectors corresponding to the words "sad" and "dejected" using the lookup table of the word embedding model. Then, the sentiment guidance module 104 calculates the cosine similarity between these two word vectors. The specific formula is the dot product of the two vectors divided by the product of their respective magnitudes. The result is a value between -1 and 1; the closer the value is to 1, the more semantically similar the two words are. The emotion guidance module 104 compares the calculated cosine similarity value with a preset threshold of 0.7. If the cosine similarity value is greater than or equal to 0.7, the two emotion categories are determined to be highly similar in semantics. The corresponding causal relationship chain is then included in the candidate set, ensuring that even if the emotion category descriptions are not completely consistent, historical associations with similar semantics can be effectively utilized. This enhances the breadth of guidance strategies in complex emotional contexts. The process is completed automatically within the emotion guidance module 104 without the need for an external network connection, ensuring real-time processing and privacy security.
[0090] From the optimal causal relationship chains retrieved, the emotion guidance module 104 extracts the unique identifier of the emotional response template associated with each chain. Using this identifier, the emotion guidance module 104 accesses the emotional response template library maintained by the context attribution module 102, obtains the corresponding complete template, and specifically reads the emotional evolution pattern sub-template. The content stored in this emotional evolution pattern sub-template is the historical typical emotional evolution path for the trigger source. The generation process of this path has been described in detail in the context attribution module 102. It is obtained by clustering the temporal change trajectory of the dynamic emotional feature vector output by the emotional feature decoupling module 103 in multiple historical events under the same trigger source, i.e., a specific family member and keyword combination. Specifically, in the query and application phase of the emotion guidance module 104, each historical typical emotion evolution path stored in this sub-template is represented as a trajectory composed of time series points. Specifically, it is encoded as the corresponding emotional state subspace coordinate sequence at time series 0 seconds, 5 seconds, 10 seconds, 20 seconds, and 30 seconds after the event is triggered, which are (0.1, 0.2), (0.3, 0.5), (0.6, 0.7), (0.5, 0.6), and (0.3, 0.4). This trajectory represents the typical pattern of emotion rising rapidly from low arousal to a high point and then slowly falling back. The emotion guidance module 104 obtains these path data to understand the possible future direction of the current emotion.
[0091] Next, the emotion guidance module 104 selects a guidance strategy from its preset guidance strategy library based on the matching degree between the current emotion intensity level and the acquired historical typical emotion evolution path. This preset guidance strategy library is a structured set of rules stored inside the emotion guidance module 104. Each strategy is a data record, which includes a strategy ID, a list of applicable emotion categories, an applicable intensity range, a description of the matched typical evolution path pattern, an expected guidance goal, and a brief identifier of the strategy execution logic. For example, a strategy can be recorded as having a strategy ID of S02, applicable emotion categories of sadness and frustration, an applicable intensity range of 0.4 and 0.9, a matching path pattern of rapid rise followed by a gradual decline, and an expected guidance goal of empathic comfort and positive memory guidance.
[0092] The selection process is a rule-based decision tree. The emotion guidance module 104 first filters out all applicable strategy candidates based on the current emotion category and the current emotion intensity level. Then, it calculates the time-series curve of the current real-time emotion intensity since the most recent suspected triggering event. The time-series curve of the emotion intensity is continuously provided by the emotion feature decoupling module 103. It performs dynamic time-normalization distance calculation with the historical typical emotion evolution path to determine which historical path the current emotion evolution best matches. Finally, the emotion guidance module 104 selects the strategy whose typical evolution path pattern description matches the current best matching path. If the current emotion state is sadness, intensity is 0.75 and the matching path is a rapid rise followed by a slow decline, the system will select strategy S02.
[0093] Once a strategy is selected, the emotion guidance module 104, combining the expected goals of the guidance strategy, the path patterns in the emotion evolution model sub-templates, and the specific family member identities and keywords in the causal relationship chain, retrieves and generates preliminary guiding dialogue content from the dialogue corpus template library. This dialogue corpus template library is a collection of text templates stored locally by the emotion guidance module 104. Each template is a sentence or dialogue flow structure with specific placeholders marked in square brackets, including but not limited to emotion categories, intensity descriptors, family members, keywords, empathy prefixes, positive recall guiding sentences, and distraction suggestions.
[0094] For example, did you feel a certain type of emotion because this was mentioned? Let's not think about this for now, let's try to distract ourselves, how about that? The Emotional Guidance Module 104, based on the expected goals of the selected strategy (such as empathic comfort and positive memory guidance) and the current path pattern (such as being in the early stages of a gradual decline), retrieves matching template groups from the library. It then fills in specific instance data into placeholders, such as "son," "busy with work," and "sadness." It also selects "I understand how you feel" from a preset phrase library and "Let me play you some of your favorite Suzhou storytelling" as a distraction suggestion, thereby generating preliminary text content. Specifically, it might say something like, "I understand how you feel. Are you feeling a little sad because your son mentioned being busy with work? Let's not think about that for now. How about I play you some of your favorite Suzhou storytelling?" Next, the emotion guidance module 104 dynamically adjusts the broadcast parameters of the generated guiding dialogue content based on the real-time intensity of the current emotion state recognition results continuously obtained from the emotion feature decoupling module 103, including tone intensity, word intimacy, and guidance urgency.
[0095] The emotion guidance module 104 maintains a parameter mapping table to define the mapping rules from emotional intensity scalar values to specific speech synthesis parameters. For example, when the real-time intensity value is greater than 0.7, the tone intensity parameter is set to gentle and slow, corresponding to a 20% reduction in the speech rate and a 10% reduction in pitch of the speech synthesis engine. The word intimacy parameter triggers the use of affectionate terms for the elderly, such as "Grandma Wang" instead of "you." The urgency parameter is set to high, corresponding to a 0.5-second reduction in pauses between sentences and a higher pitch at the end of interrogative sentences. After applying these rules, the final broadcast instruction of the above preliminary text is adjusted to say in a gentle and slow tone, "Grandma Wang, I understand how you feel. Are you feeling a little sad because your son mentioned being busy with work? Let's not think about that for now. Let me play you a piece of your favorite Suzhou storytelling. How does that sound?" The question then asks for a slight rise in tone.
[0096] Finally, the emotional guidance module 104 converts the final guided dialogue content, which has undergone strategy selection, content generation, and dynamic parameter adjustment, into an instruction format that the robot can execute. This instruction is then output in natural speech form through the speech synthesis and playback interface of the elderly care service robot. Simultaneously, the text content is sent to the robot's interactive dialog box control and displayed synchronously on the screen, thus completing a complete, data-driven, and personalized emotional guidance intervention.
[0097] In a preferred embodiment of the present invention, when the current emotional state identification result is a persistent and intensifying negative emotion, and the guidance of the emotion guidance module is ineffective, a warning message containing specific trigger source analysis is generated and sent based on the causal relationship chain in the emotion causal graph, including: When the current emotional state identification result is negative and the intensity continues to rise within multiple consecutive time windows, it is judged as a persistent and intensifying negative emotion. If the intensity of negative emotions is not effectively reduced within the set time after the most recent guidance session, the emotional guidance is deemed ineffective. An alert is triggered when both of the above criteria are met simultaneously. Based on current and recent negative emotional states, the emotional causal graph is traced back to retrieve all negative causal relationship chains with prominent emotional triggering power values; From the retrieved causal relationship chain, extract the family member identity, keywords, emotional triggering power value, and associated emotional response template of the trigger source; Analyze the degree of deviation between the current emotional evolution trajectory and the historical emotional evolution pattern of the trigger source, integrate the trigger source information, emotional trigger effectiveness value and deviation analysis results, and generate an early warning analysis report that includes the specific trigger source, intensity and abnormal description; The warning level is determined based on the persistence and severity of negative emotions, and the report is sent to the designated terminal.
[0098] In this embodiment of the invention, the early warning module 105 serves as the system's safety guarantee and deep intervention backup unit. Its core function is to initiate and execute an automated source analysis and emergency notification process when the active intervention of the emotional guidance module 104 fails and the elderly person falls into a continuously deteriorating negative emotion.
[0099] The early warning module 105 continuously monitors the current emotion state recognition results from the emotion feature decoupling module 103 and forms a historical emotion state sequence. This result includes the emotion category obtained by mapping the predefined emotion state space and the continuous intensity level quantified by the emotion classification and intensity regression model.
[0100] The early warning module 105 first determines whether the condition of persistent and intensifying negative emotions is met. Internally, it sets a threshold for the number of consecutive time windows, which can be adjusted based on feedback from elderly users during operation. For example, it could set three consecutive time windows, each lasting 60 seconds. The early warning module 105 checks whether the emotion categories output by the emotion feature decoupling module 103 are all preset negative categories, such as sadness or anxiety, within these three consecutive time windows. Furthermore, the emotion intensity level must show a strictly monotonically increasing trend within these three windows. If the average intensity of the first window is 0.6, the average intensity of the second window is 0.7, and the average intensity of the third window is 0.8, the early warning module 105 determines it as a persistent and intensifying negative emotion only if the emotion category is consistently negative and the intensity value sequence strictly increases.
[0101] Simultaneously, the early warning module 105 evaluates the intervention effect of the emotional guidance module 104. The early warning module 105 records the system timestamp of the most recent output of guided dialogue content by the emotional guidance module 104 and starts an effect evaluation timer with a duration of 300 seconds. During these 300 seconds, the early warning module 105 continuously acquires real-time emotional intensity data output by the emotional feature decoupling module 103. The early warning module 105 calculates the average emotional intensity over the last 60 seconds (240th to 300th seconds after guidance) and compares this average with the average emotional intensity over the 60 seconds before guidance began. If the average emotional intensity after guidance is not lower than 85% of the average emotional intensity before guidance (i.e., the decrease is less than 15%), the early warning module 105 determines that the emotional guidance is ineffective. This 15% decrease threshold is a preset benchmark based on the physiological and psychological response characteristics of elderly people's emotional regulation. When both the persistent and intensifying negative emotions and the ineffective emotional guidance conditions are logically true simultaneously, the early warning module 105 triggers the early warning process.
[0102] Once the warning process is triggered, the warning module 105 immediately queries the dynamic weight database of the emotional causal graph constructed and maintained by the context attribution module 102, based on the current and recent negative emotional state (e.g., within the past 30 minutes). The target of the warning module 105 is all negative causal relationship chains with prominent emotional triggering power values. The prominent quantitative standard is that the warning module 105 retrieves all causal relationship chains generated within the past week from the dynamic weight database and filters out chains with recorded emotional triggering power values greater than or equal to 0.7. This threshold of 0.7 is set based on the statistical distribution of historical data accumulated during the long-term operation of the system and represents triggering events that have a high intensity impact on the emotions of the elderly.
[0103] In addition, the early warning module 105 will also perform aggregate analysis on multiple chains under the same trigger source, namely the same family member and keyword combination. If the average emotional trigger effectiveness value of the trigger source in the past week is greater than or equal to 0.6, and the emotional trigger effectiveness value of the most recent chain is greater than or equal to 0.7, then all recent negative chains associated with the trigger source will also be included in the search scope.
[0104] From these high-impact negative causal relationship chains retrieved, the early warning module 105 extracts key information fields, namely, the identity of the family member of the triggering source, such as son; triggering keywords, such as busy at work; the emotional triggering efficacy value recorded in the chain, and if it is an aggregate analysis, the average value and the latest value are extracted, as well as the unique identifier of the emotional response template associated with the chain.
[0105] Subsequently, the early warning module 105 performs the core deviation analysis, that is, analyzes the degree of deviation between the current emotional evolution trajectory and the historical emotional evolution pattern of the trigger source.
[0106] The early warning module 105 obtains the corresponding emotional response template from the template library maintained by the context attribution module 102 through the emotional response template identifier, and reads the emotional evolution pattern sub-template. The sub-template stores one or more typical emotional state evolution paths obtained by clustering the time-series trajectory of historical dynamic emotional feature vectors. Each path represents a common emotional change pattern. If, in actual work, a path is represented by emotional intensity subspace coordinates of (0.1, 0.2), (0.8, 0.9), and (0.3, 0.4) at 0, 100, and 200 seconds after the triggering event, it depicts a typical pattern of rapid emotional rise to high arousal followed by slow decline.
[0107] The early warning module 105 extracts the real-time emotion intensity sequence from the starting time of the current persistent negative emotion, or the most likely trigger time point analyzed in conjunction with recent voice events, from the historical cache of the emotion feature decoupling module 103 to the current moment. This sequence is obtained by mapping the dynamic emotion feature vector output by the emotion feature decoupling module 103.
[0108] The early warning module 105 uses a dynamic time warping algorithm to calculate the DTW distance between the real-time emotion intensity sequence and each typical path in the emotion evolution pattern sub-template. The DTW distance effectively measures the similarity between two sequences that exhibit non-linear scaling on the time axis. The early warning module 105 selects the typical path with the smallest DTW distance as the best-matching historical pattern. Then, the early warning module 105 calculates the DTW distance between the current real-time sequence and the best-matching historical path, and divides this distance by the sequence length of the historical path to obtain a normalized deviation score. For example, if the historical path duration is 200 seconds and the calculated DTW distance is 80, then the deviation is 0.4. The higher the deviation score, the further the current emotional response deviates from the individual's historical norm in terms of pattern or duration, and the higher the degree of abnormality.
[0109] Finally, the early warning module 105 integrates all the above analysis results to generate a structured early warning analysis report. The report content clearly includes the specific trigger source, and the specific format is as follows: Trigger source: Family member identity, such as son; Keywords: specific vocabulary, such as being busy at work.
[0110] Intensity: Described as the emotional triggering power value of this event, i.e., a specific value such as 0.82, which belongs to high intensity influence; Anomaly Description: This is generated based on the deviation analysis results. Specifically, if the current emotional evolution trajectory deviates from the individual's typical historical reaction pattern, with a deviation of 0.4, it indicates that the rate of decline after the emotional peak is significantly slower than the historical pattern, resulting in persistent distress.
[0111] The early warning module 105 determines the warning level based on the persistence and intensification of negative emotions. The calculation rule involves the early warning module 105 acquiring the latest emotion intensity value and calculating the linear regression slope of the emotion intensity within the three most recent judgment time windows. According to a preset matrix, for example, a current intensity greater than 0.8 and an upward slope greater than 0.001 / second is defined as a red warning; a current intensity between 0.6 and 0.8 and an upward slope greater than 0.0005 / second is defined as an orange warning; and all other cases are defined as yellow warnings. After determining the warning level, the early warning module 105 encapsulates a complete early warning analysis report along with the warning level and sends it to a pre-registered designated terminal through the system's integrated communication interface, such as an application push interface bound to a child's mobile phone or a web interface of the community caregiver's management platform. This completes a comprehensive security process from automatic monitoring and analysis to tiered alerts.
[0112] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.
[0113] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the system as described above. All implementations in the above system embodiments are applicable to this embodiment and can achieve the same technical effects.
[0114] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An emotion regulation system based on elderly care service robots, characterized in that, include: The collaborative acquisition module synchronously acquires facial image data of the elderly, obtains physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR) through wearable devices, and captures the voice streams of conversations between multiple people in the home environment. The contextual attribution module is used to perform speaker recognition and semantic analysis on the dialogue speech stream, extract the identities of different family members and their dialogue keywords and emotional tendencies, use the appearance of keywords as event markers, analyze the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event, establish an emotional causal map, and record the causal relationship chain of specific family members and keywords triggering specific emotional state changes in the elderly. The emotion feature decoupling module is used to receive continuous facial image data, physiological parameters, and causal relationship chains from the emotion causal graph. Using the causal relationship chains as supervision signals, it constructs an adaptive analysis model to decouple static features caused by aging and dynamic features representing real emotions from the mixed facial image data. It also verifies the results by combining synchronous physiological parameters to generate the current emotion state recognition result. The emotion guidance module is used to query the emotion causal graph based on the current emotion state recognition results and generate corresponding guiding dialogue content to strengthen positive emotions or avoid negative factors and guide the emotion to a different direction. When the current emotional state identification result is a continuous and intensifying negative emotion, and the guidance of the emotion guidance module is ineffective, the early warning module generates and sends an early warning message containing specific trigger source analysis based on the causal relationship chain in the emotion causal graph.
2. The emotion regulation system based on an elderly care service robot according to claim 1, characterized in that, This tool is used for speaker recognition and semantic analysis of spoken dialogue streams, extracting the identities of different family members and their dialogue keywords and emotional tendencies. The occurrence of keywords is used as event markers, and the temporal changes of facial image data and physiological parameters acquired synchronously before and after the event are analyzed to establish an emotional causal map, including: The system receives dialogue voice streams, facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters synchronously collected by the collaborative acquisition module. It then performs speaker identification on the dialogue voice streams to determine the identities of family members and conducts semantic analysis to extract keywords and assess their emotional tendencies. The identified keywords are marked as emotional triggering events. Facial image data, heart rate variability (HRV), and skin conductance response (GSR) physiological parameters collected synchronously within a preset time window before and after the event are extracted to form an event association dataset. The expression change trend of facial image data and the fluctuation pattern of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters in the event association dataset are analyzed. Through time alignment, a facial and physiological temporal response pattern triggered by family member identity, keywords, and emotional tendencies is established. Based on the triggering and response patterns of multiple events, an emotional causal graph is constructed with family members as nodes, keywords as edges, emotional tendencies as edge attributes, and temporal response patterns as weights.
3. The emotion regulation system based on an elderly care service robot according to claim 2, characterized in that, Record the causal chain of specific family members and keywords that trigger changes in the elderly person's specific emotional state, including: From the emotional causal graph, extract multiple associated records containing the same family member node and the same keyword edge. Each record contains the emotional tendency attribute of the edge and its corresponding temporal response pattern. Multiple temporal response patterns extracted from the same trigger source are aggregated and analyzed to identify the time delay norm of facial expression changes relative to the keyword event, as well as the baseline of response intensity of physiological parameters such as heart rate variability (HRV) and skin conductance response (GSR), and to generate an emotional response template for that trigger source. In the emotional causal graph, emotional response templates are bound to corresponding family member nodes and keyword edges. When a new voice event occurs, based on the time delay norm in the emotional response template, targeted analysis of facial image data and physiological parameters is initiated after a preset time window to capture delayed emotional responses. Based on the comparison between the results of the targeted analysis and the baseline of the reaction intensity, the emotional triggering power value of this event is quantified. The family member identity, keywords, emotional tendency attributes, the emotional triggering power value of this event, and the corresponding emotional response template identifier are encapsulated into a causal relationship chain and stored in the dynamic weight database of the emotional causal graph.
4. The emotion regulation system based on an elderly care service robot according to claim 3, characterized in that, This is used to receive continuous facial image data, physiological parameters, and causal chains from an emotion causal graph. Using these causal chains as supervisory signals, an adaptive analysis model is constructed, including: Using the causal relationship chains stored in the dynamic weight database of the emotional causal graph as the source of supervision signal, the facial image data subsequences within the time period of the emotional triggering event corresponding to each causal relationship chain are extracted as the input sample set for model training. A three-component adaptive analysis model is constructed, which includes a feature extraction unit, a dynamic feature separation unit, and a static feature compensation unit. The feature extraction unit is configured to encode a subsequence of input facial image data into a shared feature vector sequence; The dynamic feature separation unit is configured to use the emotional triggering efficacy value recorded in the current training causal relationship chain as the regression target, and the time delay norm in the emotional response template associated with the chain as the temporal attention constraint to extract dynamic feature components related to event triggering from the shared feature vector sequence. The static feature compensation unit is configured to learn and generate a static feature baseline vector representing the long-term aging trend of an individual from facial image features extracted by the feature extraction unit outside the emotional trigger event time window.
5. The emotion regulation system based on an elderly care service robot according to claim 4, characterized in that, From mixed facial image data, static features caused by aging and dynamic features representing true emotions are decoupled and validated using synchronized physiological parameters to generate current emotion state recognition results, including: The shared feature vector sequence generated by the feature extraction unit is simultaneously input into the dynamic feature separation unit and the static feature compensation unit. The dynamic feature separation unit processes the emotion triggering efficacy value and time delay norm in the causal relationship chain and outputs the dynamic emotion feature vector. The static feature compensation unit extracts the static feature baseline vector from the same shared feature vector sequence. By jointly optimizing the feature extraction unit, dynamic feature separation unit, and static feature compensation unit, and introducing a feature decoupling loss function, the correlation between the dynamic emotion feature vector and the static feature baseline vector in the feature space is lower than a preset threshold, thereby achieving feature separation. The temporal changes of dynamic emotion feature vectors are correlated with the temporal changes of heart rate variability (HRV) and skin conductance response (GSR) physiological parameters synchronously recorded within the corresponding time period. If the correlation is higher than the verification threshold, it is used as a reward signal to optimize the dynamic feature separation unit. When applying the model, the same processing is performed on the real-time facial image data stream to obtain a real-time dynamic emotion feature vector. After the physiological parameters are synchronously verified to be valid, the feature values are mapped to a predefined emotion state space to generate the current emotion state recognition result.
6. The emotion regulation system based on an elderly care service robot according to claim 5, characterized in that, The emotional response template includes: The time delay norm sub-template is used to store typical delay time parameters and their confidence intervals for facial expression changes relative to keyword event triggering, obtained from historical time-series response pattern statistics. The baseline sub-template for response intensity is used to store the mean and fluctuation range of the baseline response intensity of heart rate variability (HRV) and skin conductance response (GSR) calculated based on the statistical characteristics of historical physiological parameters. The emotion evolution pattern sub-template is used to store the typical emotional state evolution path triggered by the trigger source, which is obtained by clustering the temporal change trajectory of dynamic emotion feature vectors in multiple events. Template confidence metrics are used to record the number of historical events on which the emotional response template was based, the consistency measure of the temporal response pattern, and the most recent update time. Among them, the time delay norm sub-template and the response intensity baseline sub-template are used to guide the temporal attention constraints and feature extraction of the dynamic feature separation unit, and the emotion evolution pattern sub-template is used to assist the emotion guidance module in generating guidance strategies that conform to the individual's emotional change patterns.
7. An emotion regulation system based on an elderly care service robot according to claim 6, characterized in that, This tool is used to query the emotional causal graph based on the current emotional state recognition results and generate corresponding guided dialogue content to strengthen positive emotions or avoid negative factors and guide emotional shifts, including: Receive the current emotion state recognition result, which includes the current emotion category and intensity level. Use the current emotion category and intensity level as query conditions to search for causal relationship chains in the emotion causal graph that contain corresponding and similar emotion states as weights. From the causal relationship chain retrieved, extract the associated emotional response templates and read the emotional evolution pattern sub-templates to obtain the historical typical emotional evolution path for the trigger source. Based on the matching degree between the current emotional intensity level and the typical historical emotional evolution path, a guidance strategy is selected from the pre-set guidance strategy library. The strategies in the guidance strategy library are pre-labeled with the expected guidance goals according to the emotional state and evolution path. Combining the expected goals of the guidance strategy, the path patterns in the emotional evolution model sub-templates, and the specific family member identities and keywords in the causal relationship chain, preliminary guiding dialogue content is retrieved and generated from the dialogue corpus template library. Based on the real-time intensity of the current emotion state recognition results, the tone intensity, word intimacy, and guidance urgency parameters of the generated guiding dialogue content are dynamically adjusted. The final generated guided dialogue content will be output through the voice interaction interface and interactive dialog box of the elderly care service robot.
8. The emotion regulation system based on an elderly care service robot according to claim 7, characterized in that, When the current emotional state is identified as a persistent and intensifying negative emotion, and the guidance from the emotion guidance module is ineffective, a warning message containing specific trigger source analysis is generated and sent based on the causal relationship chain in the emotion causal graph, including: When the current emotional state identification result is negative and the intensity continues to rise within multiple consecutive time windows, it is judged as a persistent and intensifying negative emotion. If the intensity of negative emotions is not effectively reduced within the set time after the most recent guidance session, the emotional guidance is deemed ineffective. An alert is triggered when both of the above criteria are met simultaneously. Based on current and recent negative emotional states, the emotional causal graph is traced back to retrieve all negative causal relationship chains with prominent emotional triggering power values; From the retrieved causal relationship chain, extract the family member identity, keywords, emotional triggering power value, and associated emotional response template of the trigger source; Analyze the degree of deviation between the current emotional evolution trajectory and the historical emotional evolution pattern of the trigger source, integrate the trigger source information, emotional trigger effectiveness value and deviation analysis results, and generate an early warning analysis report that includes the specific trigger source, intensity and abnormal description; The warning level is determined based on the persistence and severity of negative emotions, and the report is sent to the designated terminal.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the system as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the system as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Simulation accompanying robot
CN117283577A
Teenager psychological question recognition and intervention method integrating teaching atlas and grading
CN121460172A