An old person emotion two-way interaction method based on artificial intelligence technology
By using multimodal emotion feature recognition and deep learning models, we have solved the problems of inaccurate emotional interaction and insufficient cultural adaptation in the field of elderly mental health, and achieved more accurate and humane emotional interaction to meet the diverse needs of the elderly.
Patent Information
- Application Number
- CN202511131021.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing technologies lack emotional interaction and empathy capabilities in the field of elderly mental health, failing to meet the diverse emotional needs of the elderly. Furthermore, unimodal emotion analysis is inaccurate and lacks cultural adaptability, leading to inconvenience in product use.
A multimodal emotion feature recognition method is adopted, which combines emotion and context recognition. By collecting audio and video data of the elderly, audio and visual features are extracted, and a deep learning model is used to perform emotion calculation, generate context query versions and empathy vectors, and optimize the response ranking to improve the accuracy of interaction.
It improves the accuracy and humanization of emotional interactions among the elderly, enabling better understanding and response to their emotions and intentions, meeting their diverse needs, and enhancing emotional support and social interaction functions.
Smart Images

Figure CN120632431B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emotional interaction technology for the elderly, and more specifically, to a two-way emotional interaction method for the elderly based on artificial intelligence technology. Background Technology
[0002] As people age and their physical functions decline, and their children are unable to care for them in their old age for various reasons, many elderly people living alone face great challenges in their later years. In particular, they are more prone to specific mental health problems, such as empty nest syndrome, which can lead to emotional disorders such as depression, anxiety, and loneliness. These problems not only reduce their quality of life but may also have a negative impact on their physical health.
[0003] The rapid development of large language models (LLM), inference models, and agent AI approaches coincides with the growing global mental health crisis, where increasing demand has not translated into sufficient professional support, particularly for underserved populations. This presents a unique opportunity for AI to complement human-led interventions, providing scalable and context-aware support while maintaining human connection in this sensitive area.
[0004] In recent years, AI has been increasingly integrated into mental health support. Applications include Woebot, a chatbot for therapy launched by Fitzpatrick et al. in 2017. Research shows that Woebot has high user acceptance and usability ratings, demonstrating good feasibility and user experience among young people. Tutun et al. proposed an AI-based clinical decision support system (DSS) in 2023, providing an efficient and accurate tool for mental health assessment. This system can quickly identify mental disorders and improve clinical efficiency by simplifying the assessment process. van der Schyff et al. proposed an AI-driven self-help tool in 2023 that provides immediate and personalized mental health support, acting as an empathetic companion throughout the mental health journey. These AI companions can perform regular check-ins, provide emotional support, and offer personalized resource recommendations. AI-driven self-tracking insights can support users by identifying patterns and suggesting coping strategies. Li et al. launched an AI-based self-help tool in 2023. AI-driven conversational agents (CAs) can provide empathetic, context-aware support to individuals in distress. Leveraging LLM and multimodal input, CAs can engage in dynamic and coherent conversations, responding to text, voice, and even nonverbal cues, thereby creating a more holistic and natural support experience similar to that of a human supporter. Research has shown that these CAs can alleviate anxiety and depressive symptoms in college students and adults. These AI-driven chatbots use cognitive behavioral therapy-based approaches to provide coping strategies.
[0005] While current technologies play a role in mental health intervention, they suffer from problems such as mechanical interaction methods, lack of empathy, and insufficient cultural adaptability. Furthermore, these applications and products often fail to adequately consider the physiological and psychological characteristics of the elderly, leading to inconvenience in use. Some emotional intelligence agents have limited functionality, only capable of simple tasks like reminding medication or playing music, failing to meet the diverse emotional needs of the elderly. For example, smart mattresses and smart bracelets, while capable of real-time health monitoring, lack emotional design and fail to provide emotional support or social interaction. Smart care robots, despite their powerful functions, lack human-centered design and fail to meet the emotional needs of the elderly, leading to feelings of loneliness and loss during use. Smart home systems, while enabling voice or mobile app control, lack sufficient emotional design and interaction with the elderly. AI digital human social technologies, while capable of real-time communication, interaction, and emotional connection, still have limitations in emotional copywriting, often merely piling up words without delivering truly touching content. Therefore, existing products and applications have not gained widespread use in the field of mental health treatment for the elderly.
[0006] Most existing technologies only adopt a single-modal analysis and modeling approach. However, changes in human facial expressions and tone of voice are, to some extent, ways of expressing emotions. Therefore, single-modal technologies may be inaccurate in analyzing human emotions. Moreover, most existing technologies' dialogue materials and intervention strategies are mainly based on Western cultural backgrounds, which may not be suitable for users from other cultural backgrounds. For example, they may seem inappropriate when discussing certain topics and fail to meet the needs of users from different cultural backgrounds. Summary of the Invention
[0007] The purpose of this invention is to design and develop a two-way emotional interaction method for the elderly based on artificial intelligence technology. By constructing emotional vectors through multiple modalities, more effective emotional features are extracted, and combined with emotion and context recognition, the accuracy and humanization of emotional interaction for the elderly are improved.
[0008] The technical solution provided by this invention is as follows:
[0009] A two-way emotional interaction method for the elderly based on artificial intelligence technology includes the following steps:
[0010] Step 1: Collect audio and video data from the elderly;
[0011] Step 2: Extract audio feature data and visual feature data from the audio data and video data respectively, and convert the audio data into audio-transcribed text data;
[0012] Step 3: Input the audio feature data, visual feature data, and audio transcribed text data into the sentiment computing model to obtain the context query version, user sentiment vector, and empathy vector;
[0013] Step 4: Input the context query version, user sentiment vector, and empathy vector into the core chat model to obtain the response table;
[0014] Step 5: Pass the reply table into the response candidate sorter, combine the context query version, user sentiment vector and empathy vector to sort the final scores of each answer in the reply table. If the number of replies with a final score greater than 0.7 is less than 10, then arbitrarily select a reply with a score greater than 0.7 as the final reply.
[0015] If there are more than 10 responses with a final score greater than 0.7, then one response will be randomly selected from the top 10% of the candidate responses based on the total score as the final response.
[0016] Preferably, step one further includes preprocessing the audio and video data;
[0017] The preprocessing of the audio data includes:
[0018] Step a: Input the audio data into the VAD model based on energy threshold for actual speech segmentation, and then input the speech segments into the ASR model to obtain the audio data after initial text recognition;
[0019] Step b: After boosting the high-frequency components of the audio data by decibels, input the data into the RNNoise algorithm for noise reduction to obtain preprocessed audio data;
[0020] The preprocessing of the video data includes:
[0021] Step 1: If the user is over 75 years old, resample the original video to 1.2 times slower frame rate before using optical flow to detect facial motion regions. Increase the sampling weight for frame intervals with motion amplitude >15%, with a weighting ratio of [missing information]. Matching the slow movements of the elderly;
[0022] If the user is under 75 years old, optical flow is used to detect facial motion regions. Frame intervals with motion amplitude >15% are given increased sampling weight, with a weighting ratio of [missing information]. ;
[0023] in, The detected motion amplitude;
[0024] Step 2: Divide the video data into... Each part is divided into sections, and samples are taken from each section. A short segment of consecutive frames, which will One containing A series of short segments of consecutive frames are used as the data representation of the entire video;
[0025] Step 3, regarding the above One containing Age-related features are enhanced by short segments of consecutive frames.
[0026] Preferably, extracting audio feature data from the audio data includes:
[0027] Step c: Use Mel frequency cepstral coefficients as the main feature representation for the preprocessed audio data to obtain the time-frequency feature matrix;
[0028] Step d: Input the time-frequency feature matrix into the audio feature extraction model to obtain the first feature matrix;
[0029] The audio feature extraction model includes a 2DResNet-18 model, an average pooling layer, and a temporal attention layer connected in sequence.
[0030] Step e: Multiply the time-frequency feature matrix and the first feature matrix, and then pass the result through an average pooling layer to obtain the audio feature data.
[0031] Preferably, the visual feature data extracted from the video data includes:
[0032] Step 4, the above One containing Short segments of consecutive frames are input into a 3D ResNet-101 neural network model for feature extraction to obtain the second feature matrix.
[0033] Step 5: Input the second feature matrix into the spatial attention module to obtain the spatial attention weights;
[0034] The spatial attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence.
[0035] Step 6: Multiply the second feature matrix with the spatial attention weights to generate weighted spatial features;
[0036] Step 7: Input the weighted spatial features into the channel attention module to obtain the channel attention weights;
[0037] The channel attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence.
[0038] Step 8: Multiply the weighted spatial features by the channel attention weights to obtain the weighted channel features;
[0039] Step 9: Pass the weighted channel features through an average pooling layer to obtain the dimensionality-reduced channel features;
[0040] Step 10: Input the dimensionality-reduced channel features into the temporal attention module to obtain temporal attention weights;
[0041] The temporal attention module includes a one-dimensional convolutional layer, a fully connected layer, and an activation function layer connected in sequence.
[0042] Step 11: Multiply the dimensionality-reduced channel features by the temporal attention weights to obtain the weighted temporal features;
[0043] Step 12: After passing the weighted temporal features through an average pooling layer, visual feature data is obtained.
[0044] Preferably, the conversion of audio data into audio-transcribed text data specifically includes:
[0045] Step 1: Input the audio feature data into the VAD module to separate human voice from noise and obtain effective speech segments;
[0046] Step II: Input the output valid speech segments into the ASR module for speech-to-text transcription.
[0047] Preferably, the context query version of the user query includes:
[0048] Audio-to-text transcription, user intent, current emotional themes, and user personas.
[0049] Preferably, the user emotion vector includes:
[0050] Intent tags, emotion tags, current emotional themes, and user profiles.
[0051] Preferably, the empathic centripetal vector includes:
[0052] Current sentiment, current sentiment themes, and user profiles.
[0053] Preferably, the final score satisfies:
[0054] ;
[0055] In the formula, For the final score, To reply with a rating, The score is based on overall coherence.
[0056] Preferably, the response score satisfies:
[0057] ;
[0058] In the formula, For local consistency score, To retrieve the matching score, The score is based on empathy matching.
[0059] The beneficial effects of this invention are as follows:
[0060] (1) The present invention designs and develops a two-way emotional interaction method for the elderly based on artificial intelligence technology. It uses multimodal emotional features for emotion recognition, which makes up for the problem of insufficient or inaccurate emotional features in a single modality. The multimodal features interact with each other, which helps to extract more comprehensive and effective emotional features.
[0061] (2) The two-way emotional interaction method for the elderly based on artificial intelligence technology designed and developed in this invention introduces empathy analysis and user emotion analysis strategies to better understand user emotions and intentions;
[0062] (3) The present invention designs and develops a two-way emotional interaction method for the elderly based on artificial intelligence technology, and provides a framework and methodology for a two-way emotional interaction model for the elderly based on multimodal and deep learning, which fully considers the unique emotional needs of the elderly. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating the two-way emotional interaction method for the elderly based on artificial intelligence technology described in this invention.
[0064] Figure 2 This is a schematic diagram of the visual feature data extraction process described in this invention.
[0065] Figure 3 This is a schematic diagram of the process of audio transcription text data according to the present invention.
[0066] Figure 4 This is a flowchart illustrating the scenario query understanding sub-model described in this invention.
[0067] Figure 5 This is a flowchart illustrating the user understanding sub-model described in this invention. Detailed Implementation
[0068] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0069] like Figure 1 As shown, the present invention provides a two-way emotional interaction method for the elderly based on artificial intelligence technology, which includes the following steps:
[0070] Step 1: Collect audio and video data from the elderly, specifically including:
[0071] Using high-definition cameras, microphones and other equipment, audio and video data of elderly users are collected in real time to provide raw materials for subsequent processing;
[0072] Furthermore, the audio and video data are preprocessed:
[0073] 1) To better address the issues of slow speech, frequent pauses, and complex background noise among the elderly, audio data preprocessing was performed:
[0074] Step a: Use the energy threshold-based VAD model to segment the audio according to the actual speech segments to avoid mechanical segmentation at silent points. At the same time, use the ASR model for initial audio text recognition to ensure that segmentation does not interrupt complete semantic units.
[0075] Step b: Boost the high-frequency part (>4kHz) of the speech by 3dB to compensate for the frequency band attenuation common in age-related hearing loss, and use the RNNoise algorithm, which is specifically optimized for home environment noise.
[0076] 2) To capture subtle facial expressions (such as frowning and twitching of the mouth) and slow movements in the elderly, video data preprocessing was performed:
[0077] Step 1: If the user is over 75 years old, resample the original video to 1.2 times slower frame rate before using optical flow to detect facial motion regions. Increase the sampling weight for frame intervals with motion amplitude >15%, with a weighting ratio of [missing information]. Matching the slow movements of the elderly;
[0078] If the user is under 75 years old, optical flow is used to detect facial motion regions. Frame intervals with motion amplitude >15% are given increased sampling weight, with a weighting ratio of [missing information]. ;
[0079] in, The detected motion amplitude;
[0080] In this embodiment, the original video has a frame rate of 25 fps, and the resampled video has a frame rate of 20.8 fps.
[0081] In this embodiment, the motion amplitude detected by the model is 20%, so the final weighting is 1.33 times.
[0082] Step 2: Divide the video data into... Each part is divided into sections, and samples are taken from each section. A short segment of consecutive frames, which will One containing Short segments of consecutive frames are used as the data representation of the entire video, making it conform to the input requirements of a three-dimensional convolutional neural network model;
[0083] in, The values of satisfy:
[0084] If the duration of the video data is less than 10 seconds, the video data will be divided into 3 segments to ensure that each segment contains at least 2 seconds of valid content.
[0085] If the video data is 10-30 seconds long, then the video data is divided into 5 segments to simulate the opening, main body, and ending.
[0086] If the duration of the video data is longer than 30 seconds, the video data will be divided into more than 8 segments, each segment lasting 3-5 seconds, in order to better facilitate the processing of the attention mechanism module.
[0087] The values of satisfy:
[0088] The initial value is 5. If the average motion amplitude in consecutive frames is greater than 30%, then according to... Update if multiple If there are 3 consecutive frames, then select the frame with an average weight not less than the weight threshold. A series of consecutive frames;
[0089] The weight threshold is:
[0090] ;
[0091] In the formula, This is the weight threshold;
[0092] Step 3, regarding the above One containing Enhance aging features in short segments of consecutive frames:
[0093] Apply local histogram equalization (CLAHE) to areas such as the corners of the eyes and forehead.
[0094] Step 2: Extract audio feature data and visual feature data from the audio data and video data respectively, and convert the audio data into audio-transcribed text data;
[0095] The audio feature data includes the timing information of the speech, and the visual features include facial images.
[0096] The audio data extraction uses a 2D ResNet-18 neural network model as the basis, and employs an attention mechanism for audio feature extraction, including:
[0097] Step c: Using Mel-frequency cepstral coefficients (MFCCs) as the main feature representation, the preprocessed audio data is used to obtain the time-frequency feature matrix. ;
[0098] Step d: The time-frequency feature matrix Input the audio feature extraction model to obtain the first feature matrix. ;
[0099] The audio feature extraction model includes a 2DResNet-18 model, an average pooling layer, and a temporal attention layer connected in sequence.
[0100] The temporal attention layer comprises a one-dimensional convolutional block, a fully connected layer, and a ReLU activation layer connected in sequence.
[0101] Furthermore, the one-dimensional convolutional block uses three parallel 1D convolutional layers with kernel sizes of 3, 5, and 7, a stride of 1, and 64 output channels.
[0102] Step e: Calculate the time-frequency feature matrix. and the first characteristic matrix After multiplication, the final audio feature data is obtained through an average pooling layer. .
[0103] like Figure 2 As shown, the extraction of visual feature data from the video data employs a 3D ResNet-101 neural network model and an attention mechanism, specifically including:
[0104] Step 4, the above One containing Short segments from consecutive frames are input as independent units into a 3D ResNet-101 neural network model for feature extraction, resulting in a second feature matrix. ;
[0105] Step 5: Transfer the second feature matrix Spatial attention weights are obtained by passing the spatial attention module. This module characterizes the importance of features in the spatial dimension;
[0106] The spatial attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence. Specifically, the input feature matrix is first passed to a one-dimensional convolutional layer, then processed by a fully connected layer, generating spatial relevance features. Finally, the feature matrix is passed through a normalization layer. softmax The function normalizes the attention weights to obtain the spatial attention weight matrix. ;
[0107] Step 6: Transfer the second feature matrix Spatial attention weight matrix Generate a weighted spatial feature matrix by performing matrix multiplication. ;
[0108] Step 7: Calculate the weighted spatial feature matrix. The channel attention module receives the channel attention weight matrix. This module characterizes the importance of features in the spatial dimension;
[0109] The channel attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence. That is, the input features are first fed into a one-dimensional convolutional layer, then processed by a fully connected layer, generating channel-level correlation features. Finally, the input features are processed by a normalization layer. softmax The function normalizes the attention weights to obtain the channel attention weight matrix. ;
[0110] Step 8: Weight the spatial feature matrix With channel attention weight matrix Perform matrix multiplication to obtain the weighted channel feature matrix. ;
[0111] Step 9: Calculate the weighted channel feature matrix. The reduced-dimensional channel feature matrix is obtained after the average pooling layer. This reduces feature dimensionality and smooths feature information;
[0112] Step 10: Calculate the dimensionality-reduced channel feature matrix. The temporal attention weight matrix is obtained by passing it to the temporal attention module. This module characterizes the importance of features over time;
[0113] The temporal attention module comprises a one-dimensional convolutional layer, a fully connected layer, and an activation function layer connected in sequence. Specifically, the input features are first passed to a one-dimensional convolutional layer, then processed by a fully connected layer, generating temporal relevance features. These features are then passed through the activation function layer (…). ReLU The activation function is used to calculate the temporal attention weight matrix. ;
[0114] Step 11: Dimensionally reduced channel feature matrix With the time attention weight matrix Multiplying them yields a weighted time feature matrix. ;
[0115] Step 12: Calculate the weighted time feature matrix. After passing through the average pooling module, the feature dimensionality is reduced and the feature information is smoothed, resulting in the final visual feature data. .
[0116] like Figure 3 As shown, the conversion of audio data into audio-transcribed text data involves performing speech processing and understanding on the audio data, comprehensively separating human voice from background noise, and converting the original audio data into audio-transcribed text data. Specifically, this includes:
[0117] Step I: Transfer the audio feature data The audio is fed into the VAD (Video Activity Detection) module to separate human voice from noise and obtain valid speech segments. Specifically:
[0118] Step a-1: Input audio feature data Cut into Each feature segment, a set of feature segments ;
[0119] Step a-2, Time-based audio segment The input is a VAD model based on a convolutional neural network, which detects speech segments from pre-defined audio feature segments to remove noise and silence.
[0120] ;
[0121] In the formula, It is a moment The probability of voice activity. VAD The function uses a convolutional neural network to calculate... ;
[0122] The VAD model based on a convolutional neural network includes a 1D convolutional layer, a max pooling layer, and a fully connected layer connected in sequence.
[0123] Step a-3, for X After filtering, output the valid audio segments. X s ;
[0124] Specifically, when the probability of speech activity exceeds a probability threshold, it is considered to correspond to a speech signal at that moment, and the corresponding speech segment is a valid speech region; otherwise, it is considered silence or background noise. The purpose is to segment out the valid speech region. x t Speech segments deemed invalid (silence or background noise) are removed from [the list / process]. X Delete from the middle;
[0125] In this embodiment, the probability threshold is set to 0.5, which is derived empirically.
[0126] Step II: Output the valid speech segments X s The text is passed to the ASR (Automatio SpeechRecognition) module for speech-to-text transcription.
[0127] Step a-4: Input valid speech segments X s Frame sequences and text sequences Y Establish a mapping between them and define conditional probabilities:
[0128] ;
[0129] In the formula, The first in the transcribed audio results i_text One character, A It represents all possible frame-to-character alignment paths, indicating how audio frames are mapped to text characters. A ( Y ) represents all possible alignment paths that can generate Y. P ( A | X s ) indicates that the model is aligned to a certain path. A Confidence level, P ( Y | X s ) represents all that can be generated Y The sum of path probabilities is used to calculate the conditional probabilities and the total probability using the forward-backward algorithm;
[0130] Step a-5: Using a greedy decoding method, select the character with the highest probability at each time step, and simultaneously use an external language model (such as N-gram) to hypothesize (i.e., candidate results) all possible text sequences generated during the ASR decoding process. To avoid errors caused by segment context, a re-scoring process is performed to generate a text sequence. For each candidate result Calculate the joint score according to the following formula:
[0131] ;
[0132] In the formula, The score is calculated by the acoustic model CTC. This is the model score calculated by the N-gram model using ternary grammar. λ Let the weight value be , and take it as . The largest for among, .
[0133] In this embodiment, the weight value is set to 0.3;
[0134] Step a-6: Perform the recognition results of the merged speech segments, which involves punctuation restoration, text formatting, and other tasks to obtain the final audio-transcribed text data.
[0135] Step 3: Transfer the audio feature data Visual feature data An emotion computing model employing multimodal emotion recognition methods is used as input for audio-transcribed text data to obtain a context-based query version. Q c User sentiment vector e Q and empathic central vector e R ;
[0136] The sentiment computing model includes a context query understanding sub-model, a user understanding sub-model, and an interpersonal response generation sub-model.
[0137] The audio-transcribed text data is fed into the context query understanding sub-model, the audio feature data is fed into the user understanding sub-model and the interpersonal response generation sub-model, and the visual feature data is fed into the user understanding sub-model and the interpersonal response generation sub-model. The context query understanding sub-model transforms the input into a context query version that characterizes the dialogue context and semantics. Q c The user understanding sub-model transforms the input into vectors that characterize the user's emotions. e Q The interpersonal response generation sub-model transforms the input into a vector characterizing the empathy in the response. e R Specifically, it includes:
[0138] 1) such as Figure 4 As shown, the audio-transcribed text data is fed into the Contextual Query Understanding (CQU) sub-model to generate a contextual query version that characterizes the dialogue context and semantics. Q c This includes pronoun resolution and sentence completion, specifically including:
[0139] The resolution of the reference includes:
[0140] Step b-1: The WordPiece word segmentation algorithm based on the HuggingFace platform is used to convert the input audio transcribed text data into a token sequence. Each token consists of three parts: word embedding, sentence discrimination embedding, and position embedding.
[0141] Step b-2: Calculate the context representation vector for each token using a multi-layer Transformer encoder;
[0142] In this embodiment, the BERT-base model is used, which has a total of 12 transformer encoder layers;
[0143] Step b-3: Pass the context representation vector of each token through a linear layer and then through... Softmax The function is normalized to obtain the predicted tag probability distribution for each token. Here, the predicted tag set is a set of named entity tags, such as person names and place names, which describes the properties of the token. Named entity information for each token is generated based on the tag with the highest probability.
[0144] Steps b-1 to b-3 implement Named Entity Recognition (NER) of the input audio-transcribed text data. The purpose of this operation is to identify key entities (such as names of people, places, and objects) in the user input.
[0145] Step b-4: Extract mention span information from the token's named entity information. (can be) (consider it as a pronoun)
[0146] The extraction process involves passing the word segmentation results through NER to obtain the original NER tag sequence (B / I / O format), then traversing the tag sequence, merging consecutive BI tag information, and generating span information based on the merged result. The mention span information refers to a consecutive sequence of words (i.e., a phrase or word) in the text that may point to a certain entity. This span information can be noun phrases, pronouns, named entities, etc., and they may point to the same entity with other mentions in the same text (i.e., coreference).
[0147] Furthermore, extracting mention span information is not effective for longer texts. For context texts exceeding 512 tokens, segmentation is required, employing two segmentation strategies:
[0148] Strategy 1: Divide the document into non-overlapping segments. Specifically, divide the document into non-overlapping paragraphs of a fixed length (e.g., 512 tokens), and input each paragraph independently into the transformer encoder to generate local hidden vectors. Mention scores are calculated only within paragraphs using hidden vectors;
[0149] Strategy 2: Generate overlapping segments and fuse the representations of the overlapping parts through interpolation. Specifically, the segments are still divided into groups of 512 tokens each. Then, a sliding window with a stride of 256 and a length of 512 is generated to scan the segments. This design allows for overlapping segments. For the overlapping parts, the hidden vector is reconstructed using an interpolation method, calculated according to the following formula:
[0150] ;
[0151] In the formula, It is the first The final hidden vector of the overlapping parts, As the weight, it follows It decreases linearly with the increase. Indicates the first The overlapping segments are segmented k The hidden vector is calculated using the same method as the local hidden vector in Strategy 1. For the first The overlapping segments are segmented The hidden vector, generally and The hidden vectors of the left and right segments corresponding to the overlapping part;
[0152] Step b-5: Calculate mention span information and candidate antecedents using a scoring function. The compatibility between them, the candidate antecedent refers to the named entity that the pronoun may point to, according to the following formula:
[0153] ;
[0154] In the formula, For rating, This indicates that the span information and the first A score indicating the probability that each candidate antecedent points to the same entity. It refers to the score that mentions span information, indicating its score as a reference to a specific valid entity. It is the first The mention score of each candidate antecedent is calculated in the same way as... similar;
[0155] The score for mentioning span information is calculated according to the following formula:
[0156] ;
[0157] In the formula, It is a feature vector that mentions span information. It is obtained by passing the token's context vector through a transformer encoder. FFN() is a feedforward neural network. , These are learnable parameters;
[0158] The probability score for mentioning span information and candidate antecedents pointing to the same entity is calculated using the following formula:
[0159] ;
[0160] In the formula, , These are respectively mentioning span information and the first Feature vectors of candidate antecedents It is a trainable parameter matrix.
[0161] Step b-5 is performed for each segment.
[0162] Step b-6, through softmax The function computes all candidate pairs. , and the probability distribution of its rating, if P ( y_span l | x_span If it is the largest, then use y_span l replace x_span The pronoun referring to a segment, if it mentions a span of information, indicates the probability that it represents a new entity. The largest means x_span It is a new entity information;
[0163] Regarding the current mention x_span Candidate antecedents y_span l Including those in the same paragraph x_span All previous possible representations x_span Referential nouns and special symbols (express x (No common reference, which is a new entity). Specifically, it is calculated according to the following formula:
[0164] ;
[0165] in, It is a learnable empty finger score. For the first One candidate antecedent, For the first One candidate antecedent;
[0166] In particular, for The calculation method is as follows:
[0167] ;
[0168] Steps b-1 to b-6 complete the process of resolving the pronouns in the text. The purpose of this operation is to resolve the pronouns (such as "he" and "this") in the user input into specific referents.
[0169] The sentence completion includes:
[0170] Step c-1: Use a strategy combining rules and models to identify incomplete sentences in the input audio-transcribed text data. For each incomplete sentence... Calculate the probability distribution of candidate words and select the word with the highest probability distribution to fill in the sentence to complete the content;
[0171] Specifically, identifying incomplete sentences includes:
[0172] Rule-based methods refer to the use of linguistic rules to quickly filter obviously incomplete sentences. In this method, rules for determining incomplete sentences are manually specified, including grammatical rules, semantic logic rules, and contextual information rules. The grammatical rules check the completeness of sentence components, and incomplete sentences are eliminated. The semantic logic rules check whether there are semantic contradictions or interruptions in the sentence, and contradictory sentences are eliminated. The contextual information rules check whether the sentence is coherent with the historical dialogue, and incoherent sentences are eliminated.
[0173] The model method calculates the perplexity PPL of a sentence according to the following formula:
[0174] ;
[0175] In the formula, It is the first in the sentence One token, yes All previous tokens, It is a model prediction The conditional probability, N It represents the total number of tokens in the sentence.
[0176] In this embodiment, a PPL threshold is set through experiments. Sentences above the threshold are defined as incomplete sentences. For the selection of the threshold, 1,000 elderly people's dialogue sentences (50% complete and 50% incomplete) are labeled, and the PPL of different sentences is counted. The median of 80, which is the intersection of the PPL intervals of complete sentences and incomplete sentences, is taken as the threshold.
[0177] The range of candidate words is a predefined language model vocabulary. Such vocabularies are usually built based on large-scale corpora (such as Chinese Wikipedia, news, dialogue data, etc.), covering common words, phrases and some professional terms. For dialogue scenarios involving the elderly, special vocabulary in fields such as medical care and daily care can be added.
[0178] For each incomplete sentence The probability distribution of candidate words is calculated using the following formula to complete the content:
[0179] ;
[0180] In the formula, The probability distribution of candidate words. For the incomplete sentence One word, For Transformer The hidden state of the layer It is a learnable parameter matrix.
[0181] Hidden states aggregate contextual information through a multi-head attention mechanism:
[0182] ;
[0183] in, Q , K , V Let H represent the query, key, and value matrices, respectively, which are obtained from the hidden state H of the previous layer through a linear transformation. d k It is the dimension scaling factor, obtained according to the following formula. :
[0184] ;
[0185] In the formula, As the first intermediate parameter, This is the second intermediate parameter;
[0186] Step c-2: Specify the maximum sequence length as 1024. By concatenating documents and adding a "document end" marker, ensure the capture of long-distance dependencies. The generation process continues until the output end marker is reached or the maximum sequence length is reached.
[0187] In this embodiment, the maximum sequence length is 1024.
[0188] Steps c-1 to c-2 completed the work of completing the incomplete sentences in the text.
[0189] Based on steps b-1 to c-2, the final audio transcript obtained after context query understanding was obtained. Ctx Fill it into the context query version Q c .
[0190] 2) For example Figure 5 As shown, the audio transcribed text generated through speech processing and understanding... Ctx Audio feature data Visual feature data The user understanding submodel uses static attribute input from the user to generate sentiment vectors that characterize the user's emotional attributes. e Q It includes intent tags, emotion tags, current emotional themes, and user profiles, specifically:
[0191] The acquisition of the intent tags and emotion tags includes the following steps:
[0192] Step d-1: Transcribe the audio data into text data. Ctx Encoding Chinese text sequences using a MacBERT pre-trained model W Ctx ;
[0193] The MacBERT pre-trained model includes sequentially connected embedding layers, 12 layers of transformer encoders, and Deep Pyramid CNN.
[0194] Step d-2: Use 1D convolution to project features from different modalities (text encoding and audio / video feature data) onto the same dimension, and use the cross-modal fusion model MulT to concatenate speech and text features;
[0195] Step d-3: Apply a multi-task learning strategy architecture for intent recognition and sentiment classification. This strategy optimizes machine learning paradigms for multiple related tasks by sharing the underlying parameters of the model.
[0196] In this task, intent recognition and sentiment classification are performed as a joint task, sharing the front-end feature extraction layer, while the back-end uses an independent task header.
[0197] Specifically, for shared features, the concatenated features are input into a BiLSTM module and a self-attention pooling module to generate shared features h. shared .
[0198] For the task head of intent recognition, the intent types are divided into medical consultation (0), daily companionship (1), life assistance (2), memory review (3) and other purposes (4). The shared features are passed through a fully connected layer and normalized to generate a probability distribution, and the intent label with the highest probability is selected.
[0199] For the task head of emotion recognition, the emotion types are divided into happy (0), sad (1), anxious (2), angry (3), and neutral (4). The shared features are passed through a fully connected layer and normalized to generate a probability distribution. The emotion label with the highest probability is selected.
[0200] Steps d-1 to d-3 combine the input text with audio and video features to complete the intent detection and emotion recognition tasks.
[0201] The acquisition of the current emotional theme includes the following steps:
[0202] Step e-1: Decompose the audio transcribed text data into sentences according to the period and input them into the SBERT model to generate embedding vectors. Here, the embedding vector refers to the feature vector of the sentence. The feature vector length is different for sentences of different lengths. The longest feature vector is used as the standard, and feature vectors that are less than this length are padded with 0.
[0203] Step e-2: Calculate the cosine similarity of different embedding vectors to compare the similarity between embedding vectors. Then, calculate the cosine similarity between all sentences to form a similarity matrix. Based on the similarity matrix, use hierarchical clustering to divide the sentences into several classes. The cosine similarity is calculated according to the following formula:
[0204] ;
[0205] In the formula, u and v are two embedding vectors. and These are the components of the corresponding embedding vector;
[0206] Step e-3: For each type of sentence, calculate the word frequency and inverse document frequency (prevalence in the corpus) of all words in the sentence to form a TF-IDF value. Select the word with the highest TF-IDF in the category as the current sentiment topic, prioritizing the retention of verbs and nouns, while merging synonyms.
[0207] Steps e-1 to e-3 complete the task of topic detection on the input speech-transcribed text.
[0208] The acquisition of the user profile specifically includes the following steps:
[0209] Step f-1: Encode the static attributes of the user-authorized input into a continuous vector using the SBERT model;
[0210] The user input static attributes are user default information stored in the database before the interaction begins. These default attributes are generally static attributes related to the user, characterized in that they do not change or can be changed according to rules, such as the name, gender, and age of an elderly person.
[0211] Step f-2: Input the continuous vector into the LSTM to perform time-dynamic modeling of the user's attribute vector. The final hidden state output is a compact representation of the user's behavior patterns and long-term preferences. This is a fixed-dimensional vector. Long-term dependencies are captured through a gating mechanism. Here, the user's behavior patterns, including preferences and personality, are inferred from the user's static features, which is the so-called user profile.
[0212] In LSTM, the forget gate determines whether to retain or discard historical information, the input gate is used to update newly occurring events, and the output gate is used to generate the current hidden state. The cell state update mechanism combines the historical state and the current input to form an updated long-term memory.
[0213] Step f-3: Set multiple key time points (e.g., medication time), extract key features, including time features (specific time periods), behavioral features (what the user is doing via voice feedback during that time period), and emotional features (the user's characteristics during that time period). Utilize time-aware attention mechanisms to assign higher weights to these key time points. Specifically, this includes the following steps:
[0214] Convert time points into periodic codes using the following formula to capture patterns such as morning / evening and day of the week:
[0215] ;
[0216] In the formula, For timestamps, d For feature dimension, From 0 d -1, for t_code Position encoding of even-indexed positions in the feature vector at time points. for t_code Position encoding of odd-indexed positions in the feature vector at a given time point;
[0217] Calculate the importance score for each time point using the following formula:
[0218] ;
[0219] In the formula, For time points t Importance score , Learnable parameter matrix, It is a point in time. t The behavioral and emotional characteristics are used, with σ being the Sigmoid function. Key time point features are obtained by weighting the features according to the following formula, while suppressing non-critical noise:
[0220] ;
[0221] In the formula, For time points t The weighted average of behavioral and emotional characteristics;
[0222] Steps f-1 to f-3 complete the work of creating a user profile.
[0223] The obtained intent tags, emotion tags, current emotion themes, and user profiles are then populated into the user emotion vector. e Q .
[0224] Improved context query version Q c Including audio transcribed text Ctx User intent, current emotional theme, and user profile. e Q Provide emotional context, Q c Providing semantic context allows both to work together to generate empathetic responses.
[0225] 3) Transcribe audio into text after speech processing and understanding. Ctx Audio feature data Visual feature data The static attributes input by the user are passed into the interpersonal generation sub-model to generate empathy vectors that characterize emotional attributes. e R , e R The existence of this is key to providing emotional and empathetic responses;
[0226] Empathic vector e R In fact, it is a feature representation of a persona manually set by the respondent model, and its composition is the same as... e Q The difference lies in the fact that the former describes the "personality" of the model, while the latter represents the personality of a real person. For example, the setting chosen for this invention is a gentle, kind, and patient "old friend," depicted as a middle-aged or slightly older, approachable character who conveys care, encouragement, and positive emotions in conversations, and is good at listening, comforting, and providing timely encouragement and advice. Gentle humor can be used to adjust the atmosphere when appropriate. The overall style of expression is simple and straightforward, avoiding complex terminology or radical expressions.
[0227] Fill in the empathy vector using components based on user understanding e RSpecifically, the first step is to set a persona template for the model, then encode the persona template into an emotion vector. Simultaneously, the model's emotion vector is dynamically adjusted based on the user's current emotion. Finally, the adjusted emotion vector is converted into the model's current emotion and populated into the system. e R The current sentiment component of the model. This part of the work is done by the sentiment recognition component in User Understanding.
[0228] Next, the preferred themes of the preset model, such as memories, encouragement, or daily care, are analyzed from the user's emotional vector. e Q Retrieve the current sentiment theme and fill it in. e R The current emotional theme section.
[0229] Finally, pre-define some key time points, such as morning = encouragement and bedtime = soothing, assign higher weight to these key time points, and fill in the key features. e R The key time-point features section. This part of the work is done by the user persona component in User Understanding.
[0230] Step 4: Query the context version Q c User sentiment vector e Q and empathic central vector e R Input the core chat model to obtain the response form. ;
[0231] The core chat model specifically includes a paired data retrieval generator sub-model, a neural generator sub-model, and an unpaired data retrieval generator sub-model. Q c , e Q , e R Each response should be passed to one of the three sub-models, which will generate responses based on their respective algorithms and store them in the response table.
[0232] A generator is a component that generates output content based on the features of user input. From the perspective of generation method, it can be divided into retrieval generators based on paired data, neural generators, and unpaired data retrieval generators, specifically including:
[0233] 1) Query the context version of the user query. Q c User sentiment vector e Q empathic central vector e RThe steps for generating a response based on a paired data retrieval generator, as input, include:
[0234] Step h-1: Clean the 3 billion historical dialogues (from internet social data and historical dialogue records), mainly to filter out sensitive information and low-quality content and build an index. The resulting data will be used for searching in subsequent steps.
[0235] Specifically, index creation includes: setting data cleaning rules to remove sensitive information and low-quality content, and using classic models such as IK segmenter, FAISS, and clustering bucketing to create the index;
[0236] Low-quality content includes duplicate and invalid data;
[0237] Step h-2: Based on the current dialogue state (e.g., the context of the user's query), query the version. Q c User sentiment vector e Q Retrieve candidate responses using an index, output the retrieved historical dialogue data, and generate a reply. .
[0238] 2) Query the context version of the user query. Q c User sentiment vector e Q empathic central vector e R The input is used as the neural generator to generate a response;
[0239] Specifically, a GRU-RNN model is used as the neural generator to generate the contextual query version of the user query. Q c User sentiment vector e Q and empathic central vector e R The data is fed into a GRU-RNN neural network for processing. The neural network generator is trained using common social dialogue datasets such as OpenSubtitles and Reddit, i.e., based on user query context records. Q c User empathy vector e Q empathic central vector e R Concatenation to generate intermediate vectors v , the intermediate vector v Input neural generator to obtain natural language response .
[0240] 3) Query the context version of the user query. Q c User sentiment vector e Q empathic central vector e R The unpaired data retrieval generator is fed in as input to generate a response, specifically including:
[0241] Step j-1, from Q c Extract the input entity information and the relationships between entities;
[0242] The entity information includes the entity's attribute information and the attributes described in the text. The entity's attribute information includes visual attributes such as the entity's color, shape, and size.
[0243] By calculating the similarity between entities, the problem of ambiguous entity reference in multimodal data is solved. Then, entities in different modalities are aligned to identify whether they refer to the same real-world object. This can be accomplished by using the reference resolution in steps b-4 to b-6.
[0244] Step j-2: Use the extracted entities as nodes and the relationships as edges to construct the graph structure of the knowledge graph;
[0245] Step j-3: Use single sentences from news articles and speeches as unpaired data, and query versions based on the constructed knowledge graph and the context of the user query. Q c User sentiment vector e Q empathic central vector e R As input to the unpaired data retrieval generator, it searches for relevant unpaired data.
[0246] Step j-4: Retrieve data according to the principle of retaining only frequently co-occurring topics in the dialogue, and finally output the knowledge graph and the retrieved unpaired data, and generate a response. ;
[0247] The high-frequency co-occurrence refers to the top 20% of topics in terms of frequency of occurrence;
[0248] The answers generated from the three sub-models 1), 2), and 3) , , Save reply form .
[0249] Step 5: Pass the response table into the response candidate sorter. Combine the context query version, user sentiment vector, and empathy vector to calculate and sort the scores of each answer in the table. Randomly select one response from among the multiple responses with a final score higher than 0.7 as the final response. If the number of responses with a score higher than 0.7 is more than 10, then randomly select one response from the top 10% of the candidate responses in terms of total score as the final response. Random selection is used to retain a certain degree of randomness and prevent the response from appearing dull and lifeless. Specifically, this includes:
[0250] Step k-1, Regarding the response form For each candidate response, the response candidate sorter uses the DSSN model and the Entity Grid model to calculate four indicators: local consistency, global coherence, retrieval matching, and empathy matching to score and filter candidate responses.
[0251] Local consistency is used to calculate the relevance of the response to the current user input. It is calculated using the DSSN model, specifically:
[0252] Query the context version of the user query. Q c Each candidate response is encoded into fixed-dimensional vectors q and r using a BERT model, and cosine similarity is calculated. ;
[0253] Use the following formula to calculate the word-level relevance of queries and responses, enhancing sensitivity to entity words:
[0254] ;
[0255] In the formula, IDF( w ) for the search term w Frequency of occurrence in a specified corpus f ( w , R ) is a word w In the reply form The frequency of occurrence in the text, where avg_len is the average length of all responses. k 1 = 1.5 b =0.75 is an empirical parameter.
[0256] Local consistency score Calculate using the following formula:
[0257] ;
[0258] Global coherence is used to ensure that responses are consistent with the conversation history (e.g., avoiding a sudden switch to movie topics when a user is discussing music continuously). It is calculated using the Entity Grid model, specifically:
[0259] Query version from the context of the user query. Q c Extract entities and their grammatical roles (different roles of the same word) from candidate responses, construct an entity-role transition matrix, and calculate the coherence score according to the following formula:
[0260] ;
[0261] The LDA model is used to calculate the JS divergence (Theme_Score) of the historical dialogue topic distribution and candidate response topics. The global coherence score is then obtained according to the following formula. :
[0262] ;
[0263] The search calculates the matching score for keywords (BM25 / TF-IDF) and semantics (DSSM), specifically:
[0264] Context version of the user query Q c Each candidate response is encoded into fixed-dimensional vectors q and r using a BERT model, and a matching score is calculated using a deep network, as shown in the formula:
[0265] ;
[0266] In the formula, , , and These are all learnable parameters used in calculating the matching score. For matching scores;
[0267] calculate Q c and R The angle between the TF-IDF vectors is calculated by combining the term frequencies and inverse document frequencies of each word in the two texts. The cosine of the angle between the two is denoted as TFIDF_Score. The final retrieval matching score, Retrieval_Matching, is calculated according to the following formula:
[0268] ;
[0269] Empathy matching compares response empathy vectors. With the expected vector This is used to verify whether the response aligns with the model's persona and the user's emotional needs. Expected vector. The Empathy Score, derived from manually specified response trait codes, is calculated using the following formula:
[0270] ;
[0271] Step k-2: The DSSN model outputs a response score by measuring local consistency, retrieval matching, and empathy matching. The response score is calculated according to the following formula:
[0272] ;
[0273] In the formula, Rate the reply;
[0274] Step k-3: Add the response score and the global coherence score to obtain the final score, i.e.:
[0275] ;
[0276] In the formula, This is the final score.
[0277] This invention designs and develops a two-way emotional interaction method for the elderly based on artificial intelligence technology. It employs multimodal emotional features for emotion recognition, overcoming the shortcomings of insufficient or inaccurate emotional features in a single modality. The interaction between multiple modalities helps extract more comprehensive and effective emotional features, improving accuracy. Empathy analysis and user emotion analysis strategies are introduced to better understand user emotions and intentions. The method is trained using a multimodal two-way emotional interaction dataset for the elderly, fully considering their unique emotional needs. A framework and methodology for a two-way emotional interaction model for the elderly based on multimodal and deep learning are presented. Through deep learning and neural network technologies, this method enables timely communication with elderly people living alone to understand their psychological state and intervene promptly to address imbalances in their mental health, thereby improving their overall mental well-being.
[0278] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A method for two-way emotional interaction among the elderly based on artificial intelligence technology, characterized in that, Includes the following steps: Step 1: Collect audio and video data from the elderly; Step 2: Extract audio feature data and visual feature data from the audio data and video data respectively, and convert the audio data into audio-transcribed text data; Step 3: Input the audio feature data, visual feature data, and audio transcribed text data into the sentiment computing model to obtain the context query version, user sentiment vector, and empathy vector; The sentiment computing model includes a context query understanding sub-model, a user understanding sub-model, and an interpersonal response generation sub-model. The audio-transcribed text data is fed into the context query understanding sub-model, the audio feature data is fed into the user understanding sub-model and the interpersonal response generation sub-model, and the visual feature data is fed into the user understanding sub-model and the interpersonal response generation sub-model. The context query understanding sub-model transforms the input into a context query version Q that characterizes the dialogue context and semantics. c The user understanding sub-model transforms the input into a vector e that characterizes the user's emotions. Q The interpersonal response generation sub-model transforms the input into a vector e that characterizes the empathy in the response. R ; Step 4: Input the context query version, user sentiment vector, and empathy vector into the core chat model to obtain the response table; The core chat model specifically includes a paired data retrieval generator sub-model, a neural generator sub-model, and an unpaired data retrieval generator sub-model. c e Q e R Each response is passed into one of the three sub-models, which then generate responses based on their respective algorithms and store them in the response table. A generator is a component that generates output content based on the features of user input. In terms of generation method, it can be divided into retrieval generators based on paired data, neural generators, and unpaired data retrieval generators. Step 5: Pass the reply table into the response candidate sorter, combine the context query version, user sentiment vector and empathy vector to sort the final scores of each answer in the reply table. If the number of replies with a final score greater than 0.7 is less than 10, then arbitrarily select a reply with a score greater than 0.7 as the final reply. If there are more than 10 responses with a final score greater than 0.7, then one response will be randomly selected from the top 10% of the candidate responses based on the total score as the final response.
2. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 1, characterized in that, Step one also includes preprocessing the audio and video data; The preprocessing of the audio data includes: Step a: Input the audio data into the VAD model based on energy threshold for actual speech segmentation, and then input the speech segments into the ASR model to obtain the audio data after initial text recognition; Step b: After boosting the high-frequency components of the audio data by decibels, input the data into the RNNoise algorithm for noise reduction to obtain preprocessed audio data; The preprocessing of the video data includes: Step 1: If the user is over 75 years old, resample the original video to 1.2 times slower frame rate before using optical flow to detect facial motion regions. Increase the sampling weight for frame intervals with motion amplitude >15%, with a weighting ratio of [missing information]. Matching the slow movements of the elderly; If the user is under 75 years old, optical flow is used to detect facial motion regions. Frame intervals with motion amplitude >15% are given increased sampling weight, with a weighting ratio of [missing information]. Where m is the detected motion amplitude; Step 2: Divide the video data into n parts, and sample m consecutive frames from each part. Use these n short segments containing m consecutive frames as the data representation of the entire video. Step 3: Enhance the aging features of the n short segments containing m consecutive frames.
3. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 2, characterized in that, Extracting audio feature data from the audio data includes: Step c: Use Mel frequency cepstral coefficients as the main feature representation for the preprocessed audio data to obtain the time-frequency feature matrix; Step d: Input the time-frequency feature matrix into the audio feature extraction model to obtain the first feature matrix; The audio feature extraction model includes a 2DResNet-18 model, an average pooling layer, and a temporal attention layer connected in sequence. Step e: Multiply the time-frequency feature matrix and the first feature matrix, and then pass the result through an average pooling layer to obtain the audio feature data.
4. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 3, characterized in that, Visual feature data extracted from video data includes: Step 4: Input the n short segments containing m consecutive frames into the 3D ResNet-101 neural network model for feature extraction to obtain the second feature matrix; Step 5: Input the second feature matrix into the spatial attention module to obtain the spatial attention weights; The spatial attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence. Step 6: Multiply the second feature matrix with the spatial attention weights to generate weighted spatial features; Step 7: Input the weighted spatial features into the channel attention module to obtain the channel attention weights; The channel attention module includes a one-dimensional convolutional layer, a fully connected layer, and a normalization layer connected in sequence. Step 8: Multiply the weighted spatial features by the channel attention weights to obtain the weighted channel features; Step 9: Pass the weighted channel features through an average pooling layer to obtain the dimensionality-reduced channel features; Step 10: Input the dimensionality-reduced channel features into the temporal attention module to obtain temporal attention weights; The temporal attention module includes a one-dimensional convolutional layer, a fully connected layer, and an activation function layer connected in sequence. Step 11: Multiply the dimensionality-reduced channel features by the temporal attention weights to obtain the weighted temporal features; Step 12: After passing the weighted temporal features through an average pooling layer, visual feature data is obtained.
5. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 4, characterized in that, The process of converting audio data into audio-transcribed text data specifically includes: Step 1: Input the audio feature data into the VAD module to separate human voice from noise and obtain effective speech segments; Step II: Input the output valid speech segments into the ASR module for speech-to-text transcription.
6. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 5, characterized in that, The contextual query versions of user queries include: Audio-to-text transcription, user intent, current emotional themes, and user personas.
7. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 6, characterized in that, The user sentiment vector includes: Intent tags, emotion tags, current emotional themes, and user profiles.
8. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 7, characterized in that, The empathic centroid includes: Current sentiment, current sentiment themes, and user profiles.
9. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 8, characterized in that, The final score satisfies: f total =Score+Global_Coherence; In the formula, f total The final score is calculated using the "Score" option, which represents the response score, and "Global_Coherence" which represents the global coherence score.
10. The method for two-way emotional interaction among the elderly based on artificial intelligence technology as described in claim 9, characterized in that, The response score satisfies: Score=0.4·Local_Coherence+0.3·Retrieval_Matching+0.3·Empathy_Score; In the formula, Local_Coherence is the local consistency score, Retrieval_Matching is the retrieval matching score, and Empathy_Score is the empathy matching score.
Citation Information
Patent Citations
Interaction optimization method, system and equipment based on sentiment analysis and storage medium
CN120163166A
Interaction method, system, equipment and medium
CN120196214A