Multi-modal large model and dialect adaptation-based senile cognitive intervention and emotion companion system and method

By using a multimodal large model and dialect-adapted system architecture, the semantic understanding and emotion recognition problems of smart elderly care products when dealing with dialects are solved, enabling personalized and dynamic cognitive intervention and emotional companionship, and improving the interactive experience and health monitoring effect of the elderly.

CN122050360APending Publication Date: 2026-05-15史佩铭
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
史佩铭
Filing Date
2026-03-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing smart elderly care products suffer from semantic comprehension biases and emotional recognition blind spots when dealing with complex Chinese dialects, resulting in mechanical and impersonal interaction processes that lack personalization and dynamic adjustment. Furthermore, traditional cognitive intervention methods are tedious and boring, leading to strong resistance and failing to meet the deeper needs of the elderly.

Method used

The system adopts a multimodal large model and dialect adaptation system architecture, processes dialect speech and emotion through dual adapters, and combines dynamic cognitive load adjustment and implicit memory therapy to achieve natural human-computer interaction, generate speech feedback with regional accents and emotional intonation, and synchronously drive virtual avatar expressions and body movements.

Benefits of technology

It significantly enhances users' trust and sense of belonging, improves intervention compliance and continuity, provides personalized interaction, dynamically adjusts strategies to avoid frustration, supports data-driven health monitoring, and achieves emotional resonance and effective cognitive intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050360A_ABST
    Figure CN122050360A_ABST
Patent Text Reader

Abstract

The invention discloses an elderly cognitive intervention and emotion companion system and method based on a multi-modal large model and dialect adaptation, and the system comprises a multi-modal perception and feature extraction module which is used for collecting multi-modal perception data, processing the multi-modal perception data, and obtaining a multi-modal feature vector flow; the dialect adaptive processing engine module is used for performing dialect acoustic adaptation and emotion semantic adaptation processing on the multi-modal feature vector flow to obtain a semantic vector and an emotion vector; the cognitive state real-time evaluation module is used for performing multi-dimensional linguistic feature extraction on the multi-modal perception data to obtain a cognitive health index of the user; the emotion companion and intervention generation agent module is used for generating a reply text and a multi-mode emotion control label which are adaptive to the user based on the semantic vector, the emotion vector and the cognitive health index; and the multi-modal feedback synthesis module is used for calling a dialect speech synthesis model to generate speech feedback according to the reply text and the multi-modal emotion control label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of artificial intelligence and intelligent elderly care technology, and particularly relates to an elderly cognitive intervention and emotional companionship system and method based on a multimodal large model and dialect adaptation. Background Art

[0002] With the profound transformation of the global population structure, aging has become an irreversible social trend. The two major threats faced by the elderly population are emotional loneliness and cognitive impairment. Emotional loneliness: The fact that children work in different places has led many elderly people to live alone for a long time, lacking effective emotional communication, which is extremely likely to induce senile depression. Cognitive impairment: The incidence of Alzheimer's disease (AD) and related dementias increases exponentially with age. Medical research shows that early cognitive intervention and continuous social interaction are the most effective non - drug means to delay cognitive decline.

[0003] Although various intelligent elderly care products have emerged in the market, such as companion robots, smart speakers, and wearable monitoring devices, there are still significant pain points in the actual application of existing technologies, making it difficult to meet the deep - level needs of the elderly.

[0004] Currently, mainstream basic language models (such as GPT series, ERNIE Bot, etc.) are mainly trained based on standard Mandarin or English. Although some models have certain dialect translation capabilities, when dealing with China's complex dialect systems (such as Cantonese, Wu dialect, Minnan dialect, Southwestern Mandarin, etc.), they have the following serious defects: Semantic understanding deviation: Dialects often contain a large number of slang, inverted sentences, and specific cultural metaphors (such as "ling bu qing" in Hangzhou dialect means unreasonable, rather than the literal meaning). General models often can only perform literal translations, resulting in semantic understanding errors.

[0005] Emotional recognition blind spot: The emotional expression of dialects often depends on specific tone changes, modal particles (such as "lo, lie, gua" in Cantonese), and prosodic features. In the process of transcribing dialects into text by existing ASR (Automatic Speech Recognition) technologies, these paralinguistic features rich in emotional information are often discarded. As a result, the machine "understands the literal meaning but loses the emotion in the tone", leading to a mechanical and cold interaction process and being unable to establish emotional resonance with the elderly.

[0006] Existing digital cognitive intervention means mostly exist in the form of "serious games" or standardized cognitive scales (such as the MMSE scale) tests.

[0007] Boring: Repeated memory tests or arithmetic problems easily make the elderly bored.

[0008] Resistance: Explicit "testing" methods can hurt the self-esteem of the elderly, leading to resistance and extremely low compliance.

[0009] Lack of personalization: Most systems have a static intervention difficulty level, which cannot be dynamically adjusted according to the elderly person's current mental state. Too high a difficulty level can lead to frustration, while too low a difficulty level will be ineffective.

[0010] Currently, there is a lack of a comprehensive system that can organically combine high-precision dialect emotion recognition, dynamic cognitive load regulation, and implicit memory therapy, and achieve natural human-computer interaction through a multimodal large model. Summary of the Invention

[0011] To address the aforementioned technical challenges, this invention proposes an elderly cognitive intervention and emotional support system and method based on a multimodal large model and dialect adaptation. This invention solves the problem of understanding dialect emotions through an innovative dual-adaptor architecture, achieves personalized "cognitive therapy" through a dynamic cognitive load adjustment mechanism, and seamlessly integrates the intervention process into everyday casual conversation through RAG technology.

[0012] To achieve the above objectives, this invention provides an elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation, comprising: a multimodal perception and feature extraction module, a dialect adaptive processing engine module, a real-time cognitive state assessment module, an emotional support and intervention generation agent module, and a multimodal feedback synthesis module; The multimodal perception and feature extraction module is used to collect multimodal perception data, perform signal processing and feature vectorization processing on the multimodal perception data, and obtain a multimodal feature vector stream; The dialect adaptive processing engine module is used to perform dialect acoustic adaptation and sentiment semantic adaptation processing on the multimodal feature vector stream using a basic language model to obtain semantic vectors and sentiment vectors. The real-time cognitive state assessment module is used to extract multi-dimensional linguistic features from the multimodal perception data to obtain the user's cognitive health index. The emotional companionship and intervention generation agent module is used to generate response text and multimodal emotion control tags that are adapted to the user's current dialect habits, emotional state and cognitive level based on the semantic vector, emotional vector and cognitive health index and according to the preset cognitive intervention strategy library. The multimodal feedback synthesis module is used to call a dialect speech synthesis model to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates based on the reply text and multimodal emotion control tags, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

[0013] Optionally, the multimodal sensing data includes: environmental acoustic data, user voice data, facial micro-expression and eye movement data, and physiological indicators of heart rate variability and skin conductance response.

[0014] Optionally, the dialect adaptive processing engine module includes: a dialect acoustic adaptation unit and an emotional semantic adaptation unit; The dialect acoustic adaptation unit is used to map the dialect speech features in the multimodal feature vector stream to the standard semantic embedding space of the basic language model, while preserving the prosody, intonation and paralinguistic sentiment features in the speech, and to obtain the semantic vector. The emotional semantic adaptation unit is used to specifically parse inverted sentence structures, modal particles, slang, and cultural metaphors with implied emotional coloring in specific dialects, and to convert the text features in the multimodal feature vector stream into emotional vectors that can be understood by the basic language model.

[0015] Optionally, the dialect speech synthesis model includes a speaker style encoder and a prosody predictor; The speaker style encoder is used to extract the timbre features of users' relatives or typical speakers from specific dialect regions; The prosody predictor is used to receive emotion control tags from the emotional companionship and intervention generation agent, and dynamically adjust the fundamental frequency, energy and phoneme duration of the output speech to simulate the dialect intonation familiar to the elderly.

[0016] Optionally, the system further includes: a retrieval enhancement-generated memory bank and implicit memory therapy module; The retrieval enhancement generation memory bank and implicit memory therapy module is used to perform vector encoding on the current dialogue context, retrieve memory fragments with a cosine similarity higher than a preset threshold from the vector database, input the retrieved memory fragments as prompt context into the basic language model, guide the model to generate empathetic responses containing details of the user's personal experiences, implicitly trigger the user's long-term memories in natural casual conversation, and realize implicit memory therapy.

[0017] Optionally, the retrieval enhancement generation memory bank and implicit recall therapy module includes: an unstructured data storage area and a vector database; The unstructured data storage area is used to store the user's personal life history data; The vector database is used to store the vectorized representation of the individual's life course data.

[0018] This invention also provides a method for cognitive intervention and emotional support for the elderly based on a multimodal large model and dialect adaptation, including: Collect multimodal sensing data, perform signal processing and feature vectorization processing on the multimodal sensing data, and obtain a multimodal feature vector stream; The multimodal feature vector stream is processed using a basic language model for dialect acoustic adaptation and sentiment semantic adaptation to obtain semantic vectors and sentiment vectors. Multi-dimensional linguistic feature extraction is performed on the multimodal perception data to obtain the user's cognitive health index; Based on the semantic vector, emotional vector, and cognitive health index, and according to the preset cognitive intervention strategy library, a response text and multimodal emotion control tags are generated that are adapted to the user's current dialect habits, emotional state, and cognitive level. Based on the reply text and multimodal emotion control tags, a dialect speech synthesis model is invoked to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

[0019] Optionally, based on the semantic vector, emotion vector, and cognitive health index, and according to a preset cognitive intervention strategy library, a response text and multimodal emotion control tags adapted to the user's current dialect habits, emotional state, and cognitive level are generated, including: When the cognitive health index is higher than the preset first cognitive threshold, the agent enters the cognitive pressure mode, and automatically mixes non-immediate recall questions, logical association tests or rhetorical questions that require multiple rounds of reasoning into the response content generated by combining semantic vectors and emotional vectors. When the cognitive health index is lower than the preset second cognitive threshold, the agent enters a stress-relief mode, which automatically reduces the syntactic depth of the generated content by combining semantic vectors and emotional vectors, reduces the use of long and difficult sentences, and increases confirmatory phrases, reassuring interjections, and local dialect slang with positive emotional induction effects.

[0020] Compared with the prior art, the present invention has the following advantages and technical effects: 1. Strong dialect adaptability and emotional penetration: Through dual adapters, the system can process low-resource dialects and accurately identify subtle emotional changes in dialects, solving the "digital divide" problem for the elderly and significantly improving users' sense of trust and belonging.

[0021] 2. Seamless and routine intervention: Professional cognitive intervention techniques (recollection therapy, cognitive training) are integrated into everyday conversations in local dialects. The elderly do not need to set aside time for tedious training, which greatly improves the compliance and continuity of the intervention.

[0022] 3. High degree of personalization and adaptability: The system can dynamically adjust the interaction strategy according to the elderly person's real-time mental state, avoiding the frustration or ineffectiveness brought about by traditional fixed procedures.

[0023] 4. Data-driven health monitoring: The system's long-term accumulation of linguistic feature data and emotional change curves provides valuable longitudinal data support for doctors to diagnose cognitive impairments such as Alzheimer's disease, realizing the transformation from "post-treatment" to "prevention". Attached Figure Description

[0024] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a schematic diagram of the dual adapter architecture and cross-language sentiment alignment according to an embodiment of the present invention; Figure 2 This is a flowchart of the multimodal emotion consistency verification and abnormal event classification early warning process according to an embodiment of the present invention; Figure 3 This is a dialogue sequence diagram of implicit memory therapy based on RAG according to an embodiment of the present invention; Figure 4 This is a schematic diagram illustrating the changes and pattern switching of the Cognitive Health Index (CHI) over time according to an embodiment of the present invention; Figure 5 This is a structural diagram of an elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation, according to an embodiment of the present invention. Detailed Implementation

[0025] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0026] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0027] This embodiment proposes an elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation, such as... Figure 5 As shown, it specifically includes: a multimodal perception and feature extraction module, a dialect adaptive processing engine module, a real-time cognitive state assessment module, an emotional companionship and intervention generation agent module, and a multimodal feedback synthesis module; The multimodal perception and feature extraction module is used to collect multimodal perception data, perform signal processing and feature vectorization processing on the multimodal perception data, and obtain a multimodal feature vector stream; The dialect adaptive processing engine module is used to perform dialect acoustic adaptation and sentiment semantic adaptation processing on the multimodal feature vector stream using a basic language model to obtain semantic vectors and sentiment vectors. The real-time cognitive state assessment module is used to extract multi-dimensional linguistic features from the multimodal perception data to obtain the user's cognitive health index. The emotional companionship and intervention generation agent module is used to generate response text and multimodal emotion control tags that are adapted to the user's current dialect habits, emotional state and cognitive level based on the semantic vector, emotional vector and cognitive health index and according to the preset cognitive intervention strategy library. The multimodal feedback synthesis module is used to call a dialect speech synthesis model to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates based on the reply text and multimodal emotion control tags, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

[0028] Specifically, the multimodal perception and feature extraction module is configured with a high-sensitivity microphone array for collecting environmental acoustic data and user voice data, a wide-angle visual sensor for capturing facial micro-expressions and eye movement data, and a physiological indicator monitoring interface for monitoring heart rate variability (HRV) and skin conductance response; the perception layer has a built-in front-end signal processing unit for performing blind source separation, echo cancellation, and preliminary feature vectorization of the raw data; Dialect Adaptive Processing Engine Module: As the core inference hub of the system, this engine includes a pre-trained, parameter-frozen basic language model and two parallel adapter branches trained through low-rank adaptation techniques: a dialect acoustic adapter and an emotion semantic adapter. Real-time cognitive state assessment module: used to extract multi-dimensional linguistic features from the user's language output, including at least lexical density, syntactic tree depth, semantic coherence, and speech rate pause ratio; this module combines the user's historical baseline data and current physiological indicators to calculate and output the user's cognitive health index in real time; Emotional support and intervention generation agent module: Based on the basic language model, it receives semantic vectors and emotional vectors from the adapter and cognitive health index from the evaluation module; the agent generates response text and multimodal emotion control tags that are adapted to the user's current dialect habits, emotional state and cognitive level according to the preset cognitive intervention strategy library; Multimodal feedback synthesis module: Based on the reply text and emotion control tags, it calls the dialect speech synthesis model (TTS) to generate speech feedback with corresponding regional accents, specific emotional intonation and speech rate, and simultaneously drives the virtual avatar to generate matching facial expressions and body movements.

[0029] Furthermore, the multimodal perception data includes: environmental acoustic data, user voice data, facial micro-expression and eye movement data, and physiological indicators of heart rate variability and skin conductance response.

[0030] Specifically, the visual sensors in the multimodal perception layer are equipped with micro-expression analysis units based on a combination of optical flow and convolutional neural networks (CNNs): This unit captures instantaneous changes in the user's facial movements in real time; The focus of monitoring is on the combination patterns associated with negative emotions, including AU1 (inner eyebrow raised), AU4 (eyebrow furrowed), AU12 (corner of mouth stretched) and AU15 (corner of mouth drooping). The system has an "emotional consistency check" mechanism: when it detects that the user's language expression is semantically positive (such as "I am fine"), but micro-expression analysis shows that there are sad or painful characteristics (such as AU1+AU4), the system determines it as a "hidden help-seeking" signal, generates an "emotional check" instruction, and drives the agent to initiate an exploratory dialogue or directly mark the abnormal event.

[0031] Furthermore, the dialect adaptive processing engine module includes: a dialect acoustic adaptation unit and an emotional semantic adaptation unit; The dialect acoustic adaptation unit is used to map the dialect speech features in the multimodal feature vector stream to the standard semantic embedding space of the basic language model, while preserving the prosody, intonation and paralinguistic sentiment features in the speech, and to obtain the semantic vector. The emotional semantic adaptation unit is used to specifically parse inverted sentence structures, modal particles, slang, and cultural metaphors with implied emotional coloring in specific dialects, and to convert the text features in the multimodal feature vector stream into emotional vectors that can be understood by the basic language model.

[0032] Specifically, the dialect acoustic adapter is used to receive speech acoustic features and use the mapping matrix obtained by contrastive learning training to map the input dialect speech features to the standard semantic embedding space of the base model, while preserving the prosody, intonation and paralinguistic emotional features in the speech. The emotional semantic adapter is used to receive transcribed text features and specifically parse inverted sentence structures, modal particles, slang, and cultural metaphors with implied emotional coloring in specific dialects, transforming them into emotional context vectors that can be understood by large models. The dialect adaptive processing engine module adopts a "cross-language sentiment alignment" training mechanism: like Figure 1 As shown, this mechanism constructs a parallel corpus containing "standard Mandarin - specific dialects" and an emotion comparison corpus containing "same semantics - different emotions" during the model training phase; By using a contrastive learning algorithm, a loss function is constructed to minimize the feature distance of the same emotional semantics between standard language and dialect corpora, while maximizing the feature distance under different emotional states. By forcibly aligning dialect speech features with standard emotional semantic features in the latent space, the system can achieve high-precision emotion recognition of dialect speech using standard language emotion datasets without the need for large-scale annotation of dialect emotion data.

[0033] Furthermore, the dialect speech synthesis model includes a speaker style encoder and a prosody predictor; The speaker style encoder is used to extract the timbre features of users' relatives or typical speakers from specific dialect regions; The prosody predictor is used to receive emotion control tags from the emotional companionship and intervention generation agent, and dynamically adjust the fundamental frequency, energy and phoneme duration of the output speech to simulate the dialect intonation familiar to the elderly.

[0034] Specifically, the dialect speech synthesis model (TTS) adopts a hybrid architecture based on variational autoencoders (VAE) and generative adversarial networks (GAN): It includes a speaker style encoder to extract the timbre features of users' relatives or typical speakers from specific dialect regions, enabling personalized timbre replication; It includes a prosody predictor that receives emotion control tags from the emotional companionship and intervention generation agent and dynamically adjusts the fundamental frequency (F0), energy, and phoneme duration of the output speech to simulate the dialect intonations familiar to the elderly (such as the soft and gentle intonation of Wu dialect, the intonation of Sichuan and Chongqing dialect, or the nine tones and six pitches of Cantonese).

[0035] Furthermore, the system also includes: a retrieval enhancement-generated memory bank and implicit memory therapy module; The retrieval enhancement generation memory bank and implicit memory therapy module is used to perform vector encoding on the current dialogue context, retrieve memory fragments with a cosine similarity higher than a preset threshold from the vector database, input the retrieved memory fragments as prompt context into the basic language model, guide the model to generate empathetic responses containing details of the user's personal experiences, implicitly trigger the user's long-term memories in natural casual conversation, and realize implicit memory therapy.

[0036] Furthermore, the retrieval enhancement generation memory bank and implicit recall therapy module includes: an unstructured data storage area and a vector database; The unstructured data storage area is used to store the user's personal life history data; The vector database is used to store the vectorized representation of the individual's life course data.

[0037] Specifically, the memory bank includes an unstructured data storage area and a vector database; The unstructured data storage area is used to store users' personal life history data, including descriptions of old photos, digitized text of past diaries, voice samples of family members, historical dialogue records, and cultural and entertainment content that users prefer (such as operas and old movies from specific eras). Vector databases are used to store vectorized representations of the above data; When generating a response, the agent first performs vector encoding on the current dialogue context and retrieves memory fragments with a cosine similarity higher than a preset threshold from the vector database. The retrieved memory fragments are used as contextual prompts to input into the basic language model, guiding the model to generate empathetic responses that include details of the user's personal experiences. This implicitly triggers the user's long-term memories in natural casual conversation, achieving implicit memory therapy.

[0038] This embodiment also provides a method for cognitive intervention and emotional support for the elderly based on a multimodal large model and dialect adaptation, including: Collect multimodal sensing data, perform signal processing and feature vectorization processing on the multimodal sensing data, and obtain a multimodal feature vector stream; The multimodal feature vector stream is processed using a basic language model for dialect acoustic adaptation and sentiment semantic adaptation to obtain semantic vectors and sentiment vectors. Multi-dimensional linguistic feature extraction is performed on the multimodal perception data to obtain the user's cognitive health index; Based on the semantic vector, emotional vector, and cognitive health index, and according to the preset cognitive intervention strategy library, a response text and multimodal emotion control tags are generated that are adapted to the user's current dialect habits, emotional state, and cognitive level. Based on the reply text and multimodal emotion control tags, a dialect speech synthesis model is invoked to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

[0039] Specifically, S1: Baseline Model Construction: In the initial user phase (cold start period), a series of guided dialect dialogues and multimodal data collection are used to establish a personalized dialect language model, emotional expression baseline, and cognitive ability baseline profile for the user. S2: Real-time analysis of multimodal data streams: During natural interaction, speech, visual and physiological data are processed in parallel, and dialect content and emotional features are decoupled using a dual adapter; S3: Cognitive and Emotional Dual-Dimensional Assessment: Real-time cognitive load is calculated using linguistic features, and real-time emotional state is calculated using acoustic and facial features to generate a comprehensive state vector; S4: Dynamic policy matching: Input the comprehensive state vector into the policy network, and determine the current interaction mode (small talk, memory therapy, cognitive training, emotional soothing) based on the dynamic cognitive load adjustment logic. S5: Adaptive Content Generation and Error Correction: Generates dialect responses by calling a large model and adapter, utilizing RAG technology (…). Figure 3 Injecting personal memory context and removing culturally taboo words through a dialect filter ensures the cultural appropriateness of the content. S6: Multimodal Feedback and Iteration: Synthesize dialect speech and virtual avatar actions for output, and update cognitive health records and model parameters based on the user's immediate response to the reply.

[0040] Furthermore, based on the semantic vector, emotion vector, and cognitive health index, and according to a preset cognitive intervention strategy library, a response text and multimodal emotion control tags are generated that are adapted to the user's current dialect habits, emotional state, and cognitive level, specifically including: When the cognitive health index is higher than the preset first cognitive threshold, the agent enters the cognitive pressure mode, and automatically mixes non-immediate recall questions, logical association tests or rhetorical questions that require multiple rounds of reasoning into the response content generated by combining semantic vectors and emotional vectors. When the cognitive health index is lower than the preset second cognitive threshold, the agent enters a stress-relief mode, which automatically reduces the syntactic depth of the generated content by combining semantic vectors and emotional vectors, reduces the use of long and difficult sentences, and increases confirmatory phrases, reassuring interjections, and local dialect slang with positive emotional induction effects.

[0041] like Figure 2 As shown, specifically, Level 1 warning: When the cognitive health index is detected to be 2 standard deviations below the historical average in three consecutive interactions, or when obvious language logic breaks or naming aphasia symptoms appear, the system will automatically generate a "Cognitive Fluctuation Report" and push it to the guardian's terminal. Level 2 warning: When the voice emotion recognition detects extreme sadness or panic (high arousal negative emotion), or when the physiological indicator monitoring interface detects an abnormally high heart rate, the system immediately activates the emergency calming program and automatically dials the preset emergency or care number. Closed-loop feedback: The system records user feedback (interaction delay, degree of emotional improvement) after each intervention, and uses reinforcement learning (RLHF) algorithm to fine-tune the agent's decision-making strategy to achieve personalized iteration of the intervention model.

[0042] The following is a detailed description of this embodiment with reference to the accompanying drawings: I. This embodiment describes a typical product form: "Intelligent Dialect Companion Robot". Considering the privacy and real-time requirements of the elderly in their home environment, this system adopts a cloud-edge collaborative architecture of "edge-side preprocessing + cloud-based deep processing".

[0043] The terminal hardware includes: Main control chip: It adopts a high-performance embedded AI computing platform (such as NVIDIA Jetson Orin Nano) to provide at least 40 TOPS of INT8 computing power to support real-time voice wake-up, VAD (voice activity detection), basic ASR and privacy data anonymization.

[0044] Acoustic Array: Equipped with a 6-microphone circular array, supporting 360-degree sound source localization and 5-meter far-field sound pickup. Built-in deep neural network (DNN) based noise reduction algorithm, specifically designed to filter out TV background noise and renovation noise in home environments.

[0045] Vision module: Equipped with an 8-megapixel, 120-degree wide-angle RGB camera and an infrared ToF (Time of Flight) depth camera. The RGB camera is used for facial recognition and micro-expression capture, while the ToF camera is used for posture monitoring and fall detection in low-light environments.

[0046] Physiological monitoring interface: Connects to the elderly’s smart bracelet or health watch via Bluetooth 5.0 / WiFi to obtain heart rate, blood oxygen saturation and skin conductance (GSR) data in real time.

[0047] Cloud servers: Deploy large-scale multimodal models with extremely large parameters (such as models fine-tuned based on Llama-3 or Qwen-Audio).

[0048] Managed vector databases (such as Milvus or Pinecone) are used to store massive amounts of users' personal memory data.

[0049] It runs complex cognitive assessment algorithms and long-term health trend analysis models.

[0050] Data flow and privacy protection: After facial feature point extraction is completed on the device side, the original image is discarded immediately, and only the desensitized feature vector is uploaded.

[0051] Voice data is encrypted using voiceprint encryption before being uploaded.

[0052] Cloud-based training employs the federated learning approach to ensure that model updates for each user do not reveal specific dialogue content.

[0053] II. Construction and Training of the Dialect Adaptive Processing Engine: This embodiment elaborates on the implementation details of the core technical point, the "dual adapter architecture".

[0054] 1. Data preparation includes: Base corpus: Collect open-source Mandarin speech datasets (such as AISHELL-3).

[0055] Dialect corpus: For specific dialects (such as Hangzhou dialect, Cantonese), collect audio of dialect movies, TV dramas, and radio dramas, and generate annotated texts using ASR tools + manual proofreading.

[0056] Emotion parallel corpus: Record or synthesize dialect speech segments of the same semantics in different emotional states (joy, anger, sorrow, happiness), and construct positive and negative sample pairs.

[0057] 2. The model architecture design includes: In this embodiment, a pre-trained large Transformer architecture model is selected as the base.

[0058] Dialect acoustic adapter (LoRA-Acoustic): Insert a low-rank matrix into the Encoder layer of Transformer. Assuming the pre-trained weights are , we introduce two low-rank matrices , where the rank .

[0059] The updated weights after adaptation are: ; The input of this adapter is the speech mel spectrogram coefficients (MFCC) or Filterbank features processed by the convolutional layer.

[0060] Emotion semantic adapter (LoRA-Semantic): Insert the LoRA module into the Decoder layer of Transformer or a specific Attention layer. This module is specifically fine-tuned for text sequences containing dialect feature words (such as "系咪", "也就是" in Cantonese) so that the output HiddenStates can accurately reflect the emotional color of the dialect.

[0061] 3. Cross-language emotion alignment training (Contrastive Learning): To solve the problem of scarce dialect emotion data, we design a contrastive loss function: ; Where: is the feature vector of the dialect speech.

[0062] is the speech feature vector (positive sample) of the corresponding standard Mandarin with the same emotion label.

[0063] These are feature vectors (negative samples) of other emotions or semantics.

[0064] It is the temperature coefficient.

[0065] By minimizing this loss function, the model is forced to learn that regardless of whether a dialect or Mandarin is used, as long as the emotion is the same, the feature vectors should be as close as possible in the latent space. This allows the model to utilize rich Mandarin emotion data to assist in recognizing dialect emotions.

[0066] III. Dynamic Cognitive Load Regulation and Implicit Intervention Process: This example demonstrates how the system can achieve cognitive intervention in a specific interaction.

[0067] Scenario setting: User Mr. Zhang (78 years old, an early-stage Alzheimer's patient, who is accustomed to speaking Sichuan dialect).

[0068] Initial state perception: Time: 3 PM.

[0069] Sensor data: Mr. Zhang was sitting on the sofa with a stable heart rate (72 bpm), but his facial expression was slightly dull (AU4 slight frown, AU43 droopy eyelids).

[0070] Initiating a conversation: Grandpa Zhang muttered to himself, "It seems like there's nothing to do today..." Cognitive status assessment includes: Speech analysis: The system recognizes that the speech rate is slow (2.5 words / second) with many pauses.

[0071] Linguistic feature extraction includes: Vocabulary density = 0.35 (low).

[0072] Syntactic depth = 1.2 (mainly simple sentences).

[0073] Exponent calculation: such as Figure 4 As shown, based on historical benchmarks, the current Cognitive Health Index (CHI) is calculated to be 65 / 100.

[0074] Strategy decisions include: Set thresholds: Low state threshold = 60, High state threshold = 80.

[0075] CHI (65) is currently in the "low to medium range" but does not require immediate reassurance. The system decides to adopt a "mild cognitive activation" strategy (i.e., gentle pressure) to try to awaken its attention.

[0076] RAG Search and Content Generation: Search: Using "afternoon" and "nothing to do" as queries, search Mr. Zhang's memory bank.

[0077] Results found: Excerpts from Mr. Zhang's diary from when he was a fitter in a factory, as well as a photo of him playing chess with his coworkers.

[0078] The Prompt build includes: Role: Emotional partner (in Sichuan dialect)

[0079] Task: Casual conversation + mild memory recall.

[0080] Constraint: Using Sichuan dialect, mention "playing chess" to construct a simple alternative question (mild cognitive load).

[0081] The generated response includes: Agent outputs: "Uncle Zhang, the weather looks really nice today. Look at the sun over there, doesn't it look like the weather you had when you played chess with Old Li at the factory? Do you prefer to be the 'red general' or the 'black general'?"

[0082] Interactive feedback and dynamic adjustment include: User feedback: When Grandpa Zhang heard "playing chess" and "factory head", his eyes lit up (AU5 levator palpebrae superioris muscle contracted, attention was focused), and he replied: "Oh... Old Li, he's terrible at chess, I'll definitely be the red general, I'll go first!"

[0083] System analysis includes: The user's answer was logically clear and the emotion turned positive.

[0084] The real-time CHI has risen to 75.

[0085] Strategy Adjustment: The system detected that the intervention was effective and decided to gradually increase the load.

[0086] Next round of responses: Agent outputs: "Haha, Old Li is a terrible chess player? Do you remember that time you beat him with a 'post-game cannon' and won a lot of food stamps?" (Introducing specific tactical terms and reward details, requiring a deeper recollection).

[0087] IV. Comparison of System Performance and Effects in Practical Applications: To verify the effectiveness of this invention, we conducted a comparative test between this system and two existing typical products in a simulation experiment. The test results are shown in Table 1. Table 1 Data interpretation includes: In dialect sentiment classification, this invention has a significant advantage (88.7% vs 61.5%). This is directly attributed to the "dialect acoustic adapter" preserving the prosodic features of speech, while traditional systems often only process the transcribed text, losing emotional information.

[0088] In terms of the number of interactive rounds, this system has achieved a doubling of growth. This shows that "implicit memory therapy" can effectively stimulate the elderly's desire to express themselves, turning the conversation from a simple question-and-answer session into a meaningful and memorable "chat."

[0089] V. Abnormal Event Handling and Ethical Safety Mechanisms: Considering the special needs of the elderly, this system incorporates multiple redundancy designs for security.

[0090] Emotion verification mechanism: When multimodal data conflicts occur, for example: Voice content: "I'm fine, I'm doing well." (Semantic: Positive) Voice tone: low and trembling (acoustic: Negative); Facial expression: Downturned corners of the mouth (AU15), even with tears in the eyes.

[0091] Judgment: The system determines it as "masked depression".

[0092] Action: The system will not accept the text content of "I'm fine," but will mark the event as "high risk" and generate a warm, inquiring response: "You sound a little heavy-hearted. Did something sad happen to you? Tell me about it, I'm listening." At the same time, the system will record the event in the background and push it to the family members as appropriate.

[0093] Content filtering and anti-hallucination: Dialect Filter: A blacklist is created for cultural taboos in specific dialect areas (such as certain numbers or specific funeral terms). A secondary validation is performed by the rules engine before generating a response.

[0094] Medical advice filtering: If an elderly person inquires about specific medication dosages, the system recognizes the "medical intent" and forcibly triggers a RAG search of authoritative medical databases. If no definitive source is available, the system strictly outputs: "It's best to ask your doctor or check the instructions on the medicine box for specific instructions on how to take this medication, in case I remember it wrong and cause problems." Large-scale models are strictly prohibited from generating medical advice based on probability.

[0095] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A cognitive intervention and emotional support system for the elderly based on a multimodal large model and dialect adaptation, characterized in that, include: The module includes a multimodal perception and feature extraction module, a dialect adaptive processing engine module, a real-time cognitive state assessment module, an emotional companionship and intervention generation agent module, and a multimodal feedback synthesis module. The multimodal perception and feature extraction module is used to collect multimodal perception data, perform signal processing and feature vectorization processing on the multimodal perception data, and obtain a multimodal feature vector stream; The dialect adaptive processing engine module is used to perform dialect acoustic adaptation and sentiment semantic adaptation processing on the multimodal feature vector stream using a basic language model to obtain semantic vectors and sentiment vectors. The real-time cognitive state assessment module is used to extract multi-dimensional linguistic features from the multimodal perception data to obtain the user's cognitive health index. The emotional companionship and intervention generation agent module is used to generate response text and multimodal emotion control tags that are adapted to the user's current dialect habits, emotional state and cognitive level based on the semantic vector, emotional vector and cognitive health index and according to the preset cognitive intervention strategy library. The multimodal feedback synthesis module is used to call a dialect speech synthesis model to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates based on the reply text and multimodal emotion control tags, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

2. The elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation as described in claim 1, characterized in that, The multimodal perception data includes: environmental acoustic data, user voice data, facial micro-expression and eye movement data, and physiological indicators such as heart rate variability and skin conductance response.

3. The elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation as described in claim 2, characterized in that, The dialect adaptive processing engine module includes: a dialect acoustic adaptation unit and an emotional semantic adaptation unit; The dialect acoustic adaptation unit is used to map the dialect speech features in the multimodal feature vector stream to the standard semantic embedding space of the basic language model, while preserving the prosody, intonation and paralinguistic sentiment features in the speech, and to obtain the semantic vector. The emotional semantic adaptation unit is used to specifically parse inverted sentence structures, modal particles, slang, and cultural metaphors with implied emotional coloring in specific dialects, and to convert the text features in the multimodal feature vector stream into emotional vectors that can be understood by the basic language model.

4. The elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation as described in claim 3, characterized in that, The dialect speech synthesis model includes a speaker style encoder and a prosody predictor; The speaker style encoder is used to extract the timbre features of users' relatives or typical speakers from specific dialect regions; The prosody predictor is used to receive emotion control tags from the emotional companionship and intervention generation agent, and dynamically adjust the fundamental frequency, energy and phoneme duration of the output speech to simulate the dialect intonation familiar to the elderly.

5. The cognitive intervention and emotional support system for the elderly based on a multimodal large model and dialect adaptation as described in claim 1, characterized in that, The system also includes: a retrieval enhancement memory bank generation and implicit memory therapy module; The retrieval enhancement generation memory bank and implicit memory therapy module is used to perform vector encoding on the current dialogue context, retrieve memory fragments with a cosine similarity higher than a preset threshold from the vector database, input the retrieved memory fragments as prompt context into the basic language model, guide the model to generate empathetic responses containing details of the user's personal experiences, implicitly trigger the user's long-term memories in natural casual conversation, and realize implicit memory therapy.

6. The elderly cognitive intervention and emotional support system based on a multimodal large model and dialect adaptation as described in claim 5, characterized in that, The retrieval enhancement generation memory bank and implicit recall therapy module includes: an unstructured data storage area and a vector database; The unstructured data storage area is used to store the user's personal life history data; The vector database is used to store the vectorized representation of the individual's life course data.

7. A method for cognitive intervention and emotional support for the elderly based on a multimodal large model and dialect adaptation, used to implement the method as described in any one of claims 1-6, characterized in that, include: Collect multimodal sensing data, perform signal processing and feature vectorization processing on the multimodal sensing data, and obtain a multimodal feature vector stream; The multimodal feature vector stream is processed using a basic language model for dialect acoustic adaptation and sentiment semantic adaptation to obtain semantic vectors and sentiment vectors. Multi-dimensional linguistic feature extraction is performed on the multimodal perception data to obtain the user's cognitive health index; Based on the semantic vector, emotional vector, and cognitive health index, and according to the preset cognitive intervention strategy library, a response text and multimodal emotion control tags are generated that are adapted to the user's current dialect habits, emotional state, and cognitive level. Based on the reply text and multimodal emotion control tags, a dialect speech synthesis model is invoked to generate speech feedback with corresponding regional accents, specific emotional intonations and speech rates, and simultaneously drive the virtual avatar to generate matching facial expressions and body movements.

8. The method for cognitive intervention and emotional support for the elderly based on a multimodal large model and dialect adaptation as described in claim 7, characterized in that, Based on the semantic vector, emotional vector, and cognitive health index, and according to a preset cognitive intervention strategy library, a response text and multimodal emotion control tags are generated that are adapted to the user's current dialect habits, emotional state, and cognitive level, including: When the cognitive health index is higher than the preset first cognitive threshold, the agent enters the cognitive pressure mode, and automatically mixes non-immediate recall questions, logical association tests or rhetorical questions that require multiple rounds of reasoning into the response content generated by combining semantic vectors and emotional vectors. When the cognitive health index is lower than the preset second cognitive threshold, the agent enters a stress-relief mode, which automatically reduces the syntactic depth of the generated content by combining semantic vectors and emotional vectors, reduces the use of long and difficult sentences, and increases confirmatory phrases, reassuring interjections, and local dialect slang with positive emotional induction effects.