Method for Empathetic Response of Agent Based on Brain-Machine Coupling under the Guidance of Learner Emotions
Through brain-computer coupling technology and multimodal emotion analysis, agents can accurately identify learners' emotional states and generate appropriate empathy feedback, solving the problem that existing agents cannot effectively promote learners' empathy ability in collaborative learning, and achieving the effect of improving learners' emotional cognition and emotional intelligence development.
Patent Information
- Application Number
- CN202510247708.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-04
AI Technical Summary
In collaborative learning, existing agents find it difficult to accurately identify the learner's multimodal emotional state and cannot generate appropriate empathy feedback, resulting in failure to effectively promote the cultivation of learners' empathy ability.
Brain-computer coupling technology is used to collect learners' multimodal emotional data, and emotional analysis is performed through the emotional state discrimination method of multimodal data. A multimodal empathy response generation framework is designed, an empathy response model based on multimodal representation is constructed, and a joint learning training is performed to generate dialogue content that conforms to the emotional state of the learners.
It realizes precise identification of learners' emotional state and dynamically adjusts the dialogue content of the agent, which improves the emotional cognitive ability of learners in the process of collaborative learning and promotes the development of their social and emotional intelligence.
Smart Images

Figure CN119760362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the intersection and integration of artificial intelligence and brain science, and specifically to an intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions. Background Art
[0002] In the current educational environment, learners not only need to master subject knowledge, but also should possess social-emotional skills and empathy skills in order to achieve better learning outcomes in increasingly complex learning tasks. Empathy, as an individual's ability to understand, perceive others' emotions and make appropriate responses in social interactions, is one of the core elements in collaborative learning. Research shows that empathy can effectively promote trust, cooperation and understanding between individuals and others, and thus enhance learners' thinking depth and teamwork ability. Therefore, how to cultivate learners' empathy has become the key to improving learning effects.
[0003] However, the current focus of human-machine collaborative education is mainly on knowledge transfer, and less on how to help learners cultivate social-emotional skills, especially empathy. In the interaction between traditional intelligent agents and learners, although they can perform certain emotion recognition and feedback, it is usually limited to single-modal emotion expression and lacks sufficient emotion understanding ability. More importantly, existing intelligent agents often cannot accurately identify learners' emotional states in different situations, nor can they generate appropriate empathy feedback based on the emotional state, resulting in their failure to effectively promote the cultivation of learners' empathy ability during the collaborative learning process. In addition, with the progress of artificial intelligence and brain-computer interface technologies, intelligent agents should be able to perceive and feedback multi-modal emotion information in the interaction with learners, including visual, auditory and even physiological signals. This provides a new direction for the design of empathy-based intelligent agents, but how to effectively combine these technologies to generate dialogue content that conforms to the learners' emotional states remains a difficult point in the current technology. Summary of the Invention
[0004] In view of the above problems of the prior art, the present invention proposes an intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions, which can accurately identify learners' emotional states and dynamically adjust the dialogue content of the intelligent agent based on the emotional state, effectively improving learners' emotional cognitive ability during the collaborative learning process and promoting the development of their social and emotional intelligence.
[0005] To achieve the above object, the present invention proposes an intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions, and the specific steps are as follows:
[0006] S1. Collect multi-modal emotional data of learners for the empathy response of the intelligent agent. The collected multi-modal emotional data includes learners' brain data for emotional empathy and dialogue data for intelligent agent collaboration;
[0007] S2. Use the emotional state discrimination method based on multimodal data to perform emotion analysis on the multimodal data of learners under brain-computer coupling. The emotion analysis is carried out from three parts: the facial emotion analysis of learners based on the brain-computer coupling learning method, the emotional recognition of learners' dialogue texts based on the second-order interactive attention mechanism, and the discrimination of learners' emotional states by integrating brain-vision-language multimodal data;
[0008] S3. Design a multimodal empathy response generation framework based on the emotion recognition results and dialogue content, construct an empathy response model based on multimodal representations, and conduct joint learning training on the empathy response model to form an intelligent agent empathy response model based on a large language model multi-agent system under the guidance of learners' emotions.
[0009] Preferably, in S1, the specific steps for collecting learners' brain data for emotional empathy are as follows:
[0010] S11. Collect learners' brain data in a laboratory environment, and let learners watch different emotional images in the dataset to obtain EEG signals;
[0011] S12. After the EEG signals are collected, remove the invalid segments and artifacts, and use a Butterworth filter to filter the frequency band of 1 Hz - 75 Hz.
[0012] Preferably, in S1, the specific steps for collecting dialogue data with agent collaboration are as follows: After completing the stage-based classroom learning tasks, learners self-perceive their learning content and learning emotions, report them to the agent, have conversations and exchanges with the agent, and record the context dataset of the conversation as , the dialogue context is recorded as , and the agent model understands the emotions in the dialogue ; among them, represents the th speech, which contains vocabularies, and the context represents a context description composed of words. In the formula, w q is the context background information.
[0013] Preferably, in S2, the specific steps for the facial emotion analysis of learners based on the brain-computer coupling learning method are as follows:
[0014] S211. Use EEGNet and CNNNet to perform preliminary feature extraction on facial expression images and electroencephalogram signals;
[0015] Among them, for facial expression images in the visual field, an improved convolutional neural network CNNNet is used to extract preliminary features, and the preliminary representation in the visual field :
[0016] ;
[0017] In the formula, is the improved convolutional neural network CNNNet, V is the given visual image, is the preliminary representation in the visual field;
[0018] For electroencephalogram signals, the feature extractor of EEGNet is used for preliminary feature extraction, and the preliminary representation in the cognitive field :
[0019] ;
[0020] In the formula, is the compact convolutional neural network designed by EEGNet for the EEG brain-computer interface, R is the given electroencephalogram signal, is the preliminary representation in the cognitive field;
[0021] S212. Construct a cognitive-visual learning framework, project the obtained preliminary representations onto the common channel and the private channel, and encode the common channel and private channel representations;
[0022] Among them, the common channel uses an encoding function with shared parameters , learn to capture the common information between the cognitive and visual fields. After training, the common channel extracts the shared features of the cognitive and visual fields. Given and , then the common representation of the cognitive and visual field modalities is:
[0023] ;
[0024] ;
[0025] In the formula, is an encoding function based on a simple fully connected neural layer, are the shared parameters of the cognitive and visual fields;
[0026] The private representation is output by the private channels of the cognitive and visual fields, capturing the private information related to the cognitive and visual fields. Given and , then the private representation is:
[0027] ;
[0028] ;
[0029] In the formula, is the private encoding function of the private channel in the visual field, The private encoding function for the private channel in the cognitive domain, is implemented through a fully connected neural layer, is the parameter of the private encoding function for the private channel in the visual domain, is the parameter of the private encoding function for the private channel in the cognitive domain;
[0030] S213. After obtaining the common representation and the private representation, perform a concatenation operation on the obtained common representation and private representation, and perform an emotion recognition task;
[0031] Among them, the representations after concatenation of the common representations and private representations in the visual domain and the cognitive domain are respectively defined as:
[0032] ;
[0033] ;
[0034] In the formula, is the representation after concatenation in the visual domain, is the representation after concatenation in the cognitive domain;
[0035] The expressions for the emotion recognition tasks in the visual domain and the cognitive domain are respectively:
[0036] ;
[0037] ;
[0038] In the formula, and are the predicted labels of the image and the EEG signal, and KNN is used as the function model.
[0039] Preferably, in S2, the specific steps for emotion recognition of the learner's dialogue text based on the second-order interactive attention mechanism are as follows:
[0040] S221. Construct explicit information pairs and implicit information pairs for the learner's speech on learning emotions and learning experience reports as inputs, and send them into the information association module;
[0041] S222. Encode the dialogue information, regard the dialogue speech and situation between the learner and the intelligent agent as explicit information, regard the inference knowledge as implicit information, and process the explicit information and the implicit information to obtain the semantic representations of the dialogue speech, the situation and the inference knowledge;
[0042] S223. Use the information association module to capture the important correlation words in the explicit and implicit information, and perform the operation steps of constructing a correlation matrix, first-order interactive attention, second-order interactive attention, and storing the correlation words in the memory to obtain a correlation representation, and its expression is: ;
[0043] Wherein, is the association representation, is the association information, , is the number of associated words in the memory of the memory, is the encoder;
[0044] S224. Input the dialogue speech representation, the situation representation, the association representation, and the inference knowledge representation into the aggregation network to obtain the emotion representation, and use the emotion representation to predict the emotion probability respectively; wherein, the expression of the emotion probability is:
[0045] ;
[0046] ;
[0047] ;
[0048] ;
[0049] Wherein, , is the emotion probability of the dialogue speech representation, is the emotion probability of the situation representation, is the emotion probability of the association representation, is the emotion probability of the inference knowledge representation, is the softmax function, is the number of emotions, and represent aggregation networks with the same structure but different parameters, is the dialogue speech representation, is the situation representation, is the association representation, is the inference knowledge representation;
[0050] S225. Multiply the emotion probabilities as the final emotion probability, and use the log-likelihood loss function to optimize the parameters according to the emotion probability and the true label ; Its expression is: ; , where is the loss function for training and optimizing the model by calculating the logarithmic difference between the predicted emotion probability and the true emotion label.
[0051] Preferably, in S2, the specific steps for discriminating the emotional state of the learner by fusing brain-vision-language multi-modal data are as follows:
[0052] S231. Receive the emotion prediction result obtained from the second-order attention mechanism 1. Visual image-based emotion prediction labels and EEG signal-based emotion prediction labels and concatenate them into a vector ;
[0053] S232. Process the concatenated input through multiple fully connected layers MLP. Each layer is transformed by a weight matrix and an activation function to generate the final emotion prediction result, and the expression is: ; The MLP is a multi-layer perceptron that contains multiple fully connected layers and activation functions.
[0054] Preferably, in S3, the multi-modal empathy response generation framework introduces three models, namely the Perception and Retrieval Model (PRM), the Generation Model (GM), and the Retrieval-Augmented Model (RAM), to process multi-modal data; the multi-modal empathy response generation framework encodes images using ViT-G / 14, Qformer, and a linear layer, decodes images using the DALL·E2 image decoder, and uses GPT-4 for language modeling.
[0055] Preferably, in S3, the specific steps for constructing the empathy response model based on multi-modal representations are as follows:
[0056] S31. Perform multi-modal input perception and feature fusion: Use the conversation content between the learner and the agent as the multi-modal input, convert it into a feature vector that can be solved by a large language model, encode the image through a pre-trained visual encoder, and convert the image into a feature vector through BLIP and MLP and fuse it with the text features; where each text token is embedded as a semantic information vector of the text ;
[0057] S32. Perform multi-modal output and generation: Perform operations such as expanding the vocabulary, text generation, and image generation to obtain a loss function optimization model, and perform retrieval-augmented image generation operations.
[0058] Preferably, in S32, the expressions of the obtained loss functions are:
[0059] ;
[0060] ;
[0061] ;
[0062] In the formula, is the language modeling loss, is the image generation loss, is the image retrieval loss, are the features of the input text and images, represented as generated text tokens and visual tokens, , is the embedded representation of the image and represents visual features, is the text-to-image contrast loss, is the image-to-text contrast loss, is the hidden state in the generation task, is the query feature, is the frozen text encoder, is the description of the image, is the emotion information.
[0063] Preferably, in S3, the empathy response model jointly learns and trains by training the model in an end-to-end manner, using the adapter fine-tuning method for joint fine-tuning, synchronously updating a limited number of parameters in the LLM, and at the same time updating the input linear projection layer and the feature mapping module to obtain the final overall loss function; the final loss function includes the language modeling loss , the image generation loss and the image retrieval loss
[0064] ;
[0065] wherein, is a hyperparameter, for the perception and retrieval model, ; for the generation model, .
[0066] Therefore, the present invention proposes an agent empathy response method based on brain-machine coupling under the guidance of the learner's emotion, and its beneficial effects are as follows:
[0067] (1) The present invention proposes an agent empathy response method based on brain-machine coupling, which enhances the cultivation of the learner's empathy ability by accurately identifying the learner's emotional state and dynamically adjusting the agent's conversation content based on the emotional state.
[0068] (2) The present invention combines electroencephalogram (EEG) signals and multi-modal emotion recognition technology, obtains the learner's emotional feedback in real time through brain-machine coupling, and generates more empathetic conversation content on this basis, improving the learner's emotional cognitive ability in the collaborative learning process and promoting the development of their social and emotional intelligence.
[0069] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1It is the main flowchart of the intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions in the present invention;
[0071] Figure 2 It is the cognitive-visual learning framework diagram of the intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions in the present invention;
[0072] Figure 3 It is the diagram of the learner dialogue text emotion recognition model based on the second-order interactive attention mechanism of the intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions in the present invention;
[0073] Figure 4 It is the multi-modal empathy response generation framework diagram of the intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions in the present invention. Detailed implementation manners
[0074] To make the technical solutions, advantages and objectives of the present invention clearer, the technical solutions of the embodiments of the present invention will be described clearly and completely below. The described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present invention without creative efforts belong to the protection scope of this application.
[0075] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art to which the present invention belongs.
[0076] As Figure 1 shown, the present invention provides an intelligent agent empathy response method based on brain-computer coupling under the guidance of learners' emotions, and the specific steps are as follows:
[0077] S1. Collect multi-modal emotional data of learners for intelligent agent empathy responses. The collected multi-modal emotional data includes learners' brain data for emotional empathy and dialogue data for intelligent agent collaboration;
[0078] The specific steps for collecting learners' brain data for emotional empathy are as follows:
[0079] S11. Collect learners' brain data in a laboratory environment, and let learners watch different emotional images in the dataset to obtain EEG signals;
[0080] After the EEG signals are collected, remove the invalid segments and artifacts, and use a Butterworth filter to filter the frequency band of 1 Hz - 75 Hz.
[0081] Among them, the EEG signals are collected through a NeuroScan 64-channel EEG cap, which includes 62 scalp electrodes and two reference electrodes, placed according to the international 10-20 standard, and the sampling rate is 1 kHz; the images are divided into 7 emotional types: happy, sad, angry, surprised, afraid, disgusted, and neutral. Each facial emotion image is displayed for 0.5 seconds, and there is a 10-second buffer between facial images; to reduce interference, irrelevant devices are turned off during the experiment to ensure a clean environment, and the experiment lasts for about 13 minutes.
[0082] The specific steps for collecting the dialogue data of agent collaboration are as follows: After completing the stage-based classroom learning tasks, the learner self-perceives the learning content and learning emotions and reports them to the agent, generates conversations and exchanges with the agent, and records the context data set of the conversation as The dialogue situation is recorded as The agent model understands the emotion A in the conversation; among them, represents the i-th speech, which contains vocabularies. The situation Q represents a situation description composed of n words. In the formula, w q is the context information of the situation, which helps to understand the context of the conversation and generate appropriate responses accordingly.
[0083] S2. Use the emotional state discrimination method of multi-modal data to perform emotion analysis on the multi-modal data of the learner under brain-computer coupling. The emotion analysis is carried out from three parts: the learner's facial emotion analysis based on the brain-computer coupling learning method, the learner's dialogue text emotion recognition based on the second-order interactive attention mechanism, and the learner's emotional state discrimination by fusing brain-vision-language multi-modal data;
[0084] The specific steps of the learner's facial emotion analysis based on the brain-computer coupling learning method are as follows:
[0085] S211. Extract the preliminary representations of facial expression images and EEG signals: Use EEGNet and CNNNet to perform preliminary feature extraction on facial expression images and EEG signals;
[0086] Among them, for the facial expression images in the visual field, an improved convolutional neural network CNNNet is used to extract preliminary features, and the preliminary representation in the visual field :
[0087] ;
[0088] In the formula, is the improved convolutional neural network CNNNet, V is the given visual image, is the preliminary representation in the visual field;
[0089] For electroencephalogram (EEG) signals, the feature extractor of EEGNet is used for preliminary feature extraction, which is the preliminary representation in the cognitive domain. :
[0090] ;
[0091] In the formula, is a compact convolutional neural network designed by EEGNet for the EEG brain-computer interface, R is the given EEG signal, is the preliminary representation in the cognitive domain;
[0092] Among them, EEG signals and facial expression images are two different data modalities; CNNNet consists of three convolutional modules, each convolutional module includes a convolutional layer, a normalization layer, a non-linear activation layer and a max pooling layer, and the output of the third convolutional module is used as the preliminary representation in the visual domain to reflect the learner's facial emotional response; for the EEG signals in the cognitive domain, EEGNet is used as the feature extractor. EEGNet (GR) is a compact convolutional neural network designed for the EEG brain-computer interface, which includes a standard convolutional layer, a depth convolutional layer and a separable convolutional layer. The output of the third convolutional module is used as the preliminary feature representation in the cognitive domain to capture the changes in the learner's electroencephalogram activity;
[0093] S212. Construct a cognitive-visual learning framework, project the obtained preliminary representations onto the common channel and the private channel, and encode the common channel and private channel representations;
[0094] Among them, the common channel uses an encoding function with shared parameters , to learn to capture the common information between the cognitive and visual domains. After training, the common channel extracts the shared features of the cognitive and visual domains. Given and , then the common representation of the cognitive and visual domain modalities is:
[0095] ;
[0096] ;
[0097] In the formula, is an encoding function based on a simple fully connected neural layer, are the shared parameters of the cognitive and visual domains;
[0098] The private representation is output from the private channels of the cognitive and visual domains, capturing the private information related to the cognitive and visual domains. Given and , then the private representation is:
[0099] ;
[0100] ;
[0101] In the formula, is the private encoding function of the private channel in the visual field, is the private encoding function of the private channel in the cognitive field, implemented through a fully connected neural layer, is the parameter of the private encoding function of the private channel in the visual field, is the parameter of the private encoding function of the private channel in the cognitive field;
[0102] As Figure 2 shown, the cognitive field and the visual field each have independent private channels, and at the same time capture the commonalities between the cognitive field and the visual field through a common channel. The common features of the cognitive-visual field refer to the emotional response patterns (such as emotional activation intensity, emotional category consistency, etc.) that can be shared with EEG signals by observing facial expression images. The private features of the cognitive field refer to those unique to EEG, reflecting the way the brain processes emotions and individual differences (such as differences in brain wave frequencies, differences in EEG fluctuation response intensities, etc.). The private features of the visual field refer to the individual differences and details unique to facial expressions (such as facial morphological features, micro-expression features).
[0103] S213. After obtaining the common representation and the private representation, perform a concatenation operation on the obtained common representation and private representation, and perform an emotion recognition task;
[0104] Among them, the representations after concatenation of the common representations and private representations of the visual field and the cognitive field are respectively defined as:
[0105] ;
[0106] ;
[0107] In the formula, is the representation after concatenation of the visual field, is the representation after concatenation of the cognitive field;
[0108] The expressions for the emotion recognition tasks of the visual field and the cognitive field are respectively:
[0109] ;
[0110] ;
[0111] In the formula, and are the predicted labels of the image and the EEG signal, using KNN as the function model.
[0112] AsFigure 3 As shown, a learner dialogue text emotion recognition model based on a second-order interactive attention mechanism is proposed. In a real classroom environment, the second-order interactive attention mechanism is used to analyze the explicit and implicit connotations of dialogue texts to improve the accuracy of learners' emotions. After completing a stage of learning tasks, learners self-perceive their learning content and learning emotions and report them to the intelligent agent through text or emoji pictures for emotional communication and feedback. This process regards the dialogue content as explicit information, while the implicit information includes the learners' emotional states and situational inferences. On this basis, the dialogue content is encoded from two perspectives: explicit information and implicit information.
[0113] The specific steps for learner dialogue text emotion recognition based on the second-order interactive attention mechanism are as follows:
[0114] S221. Construct explicit information pairs and implicit information pairs for the learner's speeches on learning emotions and learning feelings reports as inputs and send them to the information association module to understand the learner's current speeches on learning emotions and learning feelings reports ;
[0115] The constructed explicit information pairs:
[0116] ;
[0117] Implicit information pairs:
[0118] ;
[0119] In the formula, , represents memory, which is used to store correlation words and is initialized to be empty; when , both the dialogue and the memory are empty;
[0120] Among them, the explicit information pairs are used to capture the important associations between the current speech and the situation, dialogue history, and memory. The implicit information pairs are used to capture the implicit important associations between the current speech and the situation, dialogue history, and memory.
[0121] S222. Encode the dialogue information, regard the dialogue speeches and situations between the learner and the intelligent agent as explicit information, regard the reasoning knowledge as implicit information, and process the explicit information and implicit information to obtain the semantic representations of the dialogue speeches, situation representations, and reasoning knowledge;
[0122] Among them, the process of processing the explicit information is as follows: Add special start markers before the dialogue speech and the situation respectively to obtain the word sequence and the situation description ; For each round of speech and the situation description Encode the speech Use a speech-level encoder Generate a sentence representation:
[0123] ;
[0124] Wherein is a word embedding is a positional embedding is a status embedding, and the status embedding is used to distinguish the speaker (learner or agent) and the listener (agent or learner);
[0125] For the entire conversation Use a conversation-level encoder To encode:
[0126] ;
[0127] Wherein , , where and respectively represent the length of the i-th speech and the total number of conversation speeches, represents the hidden layer size, represents a concatenation operation;
[0128] For the situational word sequence , use the encoder To learn the situational representation:
[0129] ;
[0130] Wherein and respectively represent the word embedding and positional embedding of the situation, , is the number of words included in the situation;
[0131] The processing process of implicit information is: generate inference knowledge through the COMET model and , add before the inference knowledge tag, and encode the inference knowledge; the inference knowledge
[0132] ;
[0133] Wherein represents the semantic representation of the inference knowledge That is, for the i-th round of speech and the context reasoning knowledge; type ∈
[0134] uEffect, uReact, uIntent, uNeed, uWant, and ⊕ represents the concatenation operation;
[0135] S223. The information association module is used to capture important correlation words in explicit and implicit information, and perform operations such as constructing a correlation matrix, first-order interactive attention, second-order interactive attention, and storing correlation words in memory to obtain a correlation representation, so as to understand the conversation between the learner and the intelligent agent more coherently and comprehensively;
[0136] Among them, the correlation words exist in both types of sentences (speech or context). Therefore, the correlation words are selected bidirectionally in the two types of sentences, and the specific steps for constructing the correlation matrix are as follows:
[0137] Select the correlation words in the context based on the speech (i.e., ), select the correlation words in the speech based on the context (i.e., ). In order to capture rich features, a multi-head correlation matrix is constructed: , representing the correlation score from the context word to the speech word; represents the correlation score between the speech word and the context word. The expression of the multi-head correlation matrix is:
[0138] ;
[0139] ;
[0140] In the formula, , is the Sigmoid function, is the number of multi-heads, represents the -th head index in the multi-head attention mechanism. , represents the weight matrix and word vector used in the multi-head attention mechanism;
[0141] The operation steps for first-order interactive attention are as follows:
[0142] Identify the key words in the context and select the context words with a higher correlation degree with the speech words as the key words:
[0143] ;
[0144] ;
[0145] Among them, is the average function, is the screening function, represents the number of selected keywords, ;
[0146] The operation steps for second-order interactive attention are as follows:
[0147] Select the utterance word with the highest correlation with the key situation word as the important correlation word:
[0148] ;
[0149] Among them, and are respectively the representation and score of the utterance word related to the situation word. is the number of correlation words;
[0150] The operation steps for storing associative memory are as follows:
[0151] For explicit information, select the keyword between the utterance and the situation , the correlation word between the utterance and the dialogue history and the correlation word between the utterance and the memory :
[0152] ;
[0153] Implicit memory:
[0154] ;
[0155] When iterating to the last sentence, combine the explicit information memory and the implicit information memory as the final associative information :
[0156] ;
[0157] Among them, , ;
[0158] Based on the memory of the correlation words, learn the associative information through the encoder. The expression of the associative information is:
[0159]
[0160] In the formula, is the associative representation, is the associative information, , is the number of correlation words in the memory of the memory device, is the encoder;
[0161] S224. Input the dialogue speech representation, situation representation, association representation, and inference knowledge representation into the aggregation network to obtain the emotion representation, and use the emotion representation to predict the emotion probability respectively; where the expression of the emotion probability is:
[0162] ;
[0163] ;
[0164] ;
[0165] ;
[0166] In the formula, , is the emotion probability of the dialogue speech representation, is the emotion probability of the situation representation, is the emotion probability of the association representation, is the emotion probability of the inference knowledge representation, is the softmax function, is the number of emotions, and represent aggregation networks with the same structure but different parameters, is the dialogue speech representation, is the situation representation, is the association representation, is the inference knowledge representation;
[0167] S225. Multiply the emotion probabilities as the final emotion probability, and use the log-likelihood loss function to optimize the parameters according to the emotion probability and the true label ; its expression is: ; , where is the loss function for training and optimizing the model by calculating the logarithmic difference between the predicted emotion probability and the true emotion label.
[0168] The specific steps for discriminating the emotional state of the learner by fusing brain-vision-language multi-modal data are as follows:
[0169] S231. Receive the emotion prediction result obtained from the second-order attention mechanism, the emotion prediction label based on the visual image, and the emotion prediction label based on the EEG signal, and concatenate them into a vector ;
[0170] S232. The concatenated input Processed by multiple fully connected layers MLP, each layer is transformed by a weight matrix and an activation function to generate the final sentiment prediction result, and the expression is: ; where MLP is a multi-layer perceptron, which contains multiple fully connected layers and activation functions
[0171] S3. Design a multi-modal empathy response generation framework based on the emotion recognition result and the conversation content, construct an empathy response model based on multi-modal representation, and conduct joint learning and training of the empathy response model to form an agent empathy response model of the large language model multi-agent system under the guidance of the learner's emotion.
[0172] As Figure 4 shown, the multi-modal empathy response generation framework has the ability to perceive the learner's emotion and generate text responses and emoticons based on the learner's emotion recognition result of EEG and text data and the conversation content (including text input and emoticon input). Through different image generation strategies, three models including a perception and retrieval model (PRM), a generation model (GM), and a retrieval enhanced model (RAM) included in the multi-modal empathy response generation framework are derived based on this framework; the multi-modal empathy response generation framework encodes images using ViT-G / 14, Qformer, and linear layers, uses GPT-4 for language modeling, and the image decoder uses DALL·E 2.
[0173] The specific steps for constructing the empathy response model based on multi-modal representation are as follows:
[0174] S31. Perform multi-modal input perception and feature fusion: Use the conversation content between the learner and the agent as multi-modal input, convert it into a feature vector that can be solved by the large language model, encode the image through a pre-trained visual encoder, and convert the image into a feature vector through BLIP and MLP and fuse it with the text feature; where each text token is embedded as a semantic information vector of the text ;
[0175] S32. Perform multi-modal output and generation: Perform operations of expanding the vocabulary, text generation, and image generation to obtain a loss function to optimize the model, and perform retrieval enhanced image generation operations.
[0176] Among them, the specific steps for expanding the vocabulary are:
[0177] Add a set of visual tokens , divide the visual tokens into two groups, where the first tokens are used for image retrieval, and the last tokens are used for image generation, and use the image information as part of the generated content, and its expression is:
[0178] ;
[0179] ;
[0180] ;
[0181] In the formula, for the perception and retrieval model and the retrieval enhancement model, for the generation model and the retrieval enhancement model;
[0182] The specific steps for text generation are as follows:
[0183] After receiving the multimodal input, generate a joint sequence of text tokens and visual tokens , and represent the generated tokens as , where , then the loss function is defined as:
[0184] ;
[0185] Among them, and are the features of the input text and image, is the embedded representation of the image and represents the visual features, used for alignment with the text embedding, and the loss function optimizes the model by maximizing the generation probability.
[0186] In image retrieval, the agent empathy response model aligns the hidden state corresponding to to the retrieval space through contrastive learning and uses cosine similarity to measure the similarity of the projection vectors for the image retrieval loss , combined with the contrastive loss of text-to-image ( ) and image-to-text ( ), to optimize the mapping between the image and the text. The image retrieval loss is:
[0187] ;
[0188] Among them, is the loss used to optimize the retrieval projection layer;
[0189] The specific steps for image generation are as follows:
[0190] In the image generation stage, the generated image is generated by the DALL·E 2 decoder, and the loss between the generated image and the actual image is minimized.
[0191] The loss function is expressed as:
[0192] ;
[0193] Among them, is the hidden state in the generation task, is the query feature, is the frozen text encoder, is the description of the image, is the emotion information;
[0194] The steps for performing retrieval-enhanced image generation operations are as follows:
[0195] Retrieve an image as a latent representation to enhance the generation process. During the image generation process, still utilize as a condition. This method continues to generate on the retrieved image, which can expand the diversity of the image while maintaining the image quality.
[0196] The specific steps for the joint learning and training of the empathy response model are as follows:
[0197] Train the model in an end-to-end manner, use the adapter fine-tuning method for joint fine-tuning, synchronously update a limited number of parameters in the LLM, and at the same time update the input linear projection layer and the feature mapping module to obtain the final overall loss function; the final loss function includes the language modeling loss , the image generation loss and the image retrieval loss , then the overall loss function is expressed as:
[0198] ;
[0199] In the formula, is a hyperparameter. For the perception and retrieval model, ; for the generation model, .
[0200] Through this multi-task joint training, the intelligent agent empathy response model can generate reasonable and visually element-rich responses in empathy-rich conversations, improving the diversity and emotional understanding ability of the intelligent agent's empathy response and conversation.
[0201] Therefore, the intelligent agent empathy response method based on brain-computer coupling under the guidance of the learner's emotion provided by the present invention accurately identifies the learner's emotional state, dynamically adjusts the intelligent agent's conversation content based on the emotional state, thereby enhancing the cultivation of the learner's empathy ability. Combining electroencephalogram (EEG) and multi-modal emotion recognition technology, it obtains the learner's emotional feedback in real time through the brain-computer coupling method, and generates more empathetic conversation content on this basis, improving the learner's emotional cognitive ability during the collaborative learning process and promoting the development of their social and emotional intelligence.
[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions of the present invention or make equivalent replacements, and these modifications or equivalent replacements do not enable the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. An intelligent agent empathy response method based on brain-computer coupling under the guidance of learner emotions, characterized in that: Here are the steps: S1. Collect multimodal emotional data of learners for agent-oriented empathic response, where the collected multimodal emotional data includes learners' brain data for emotional empathy and dialogue data for agent collaboration; S2. Use the emotion state discrimination method of multimodal data to perform emotion analysis on the multimodal data of learners collected under brain-computer coupling. The emotion analysis is performed from three parts: facial emotion analysis of learners based on brain-computer coupling learning method, emotion recognition of learners' dialogue text based on second-order interactive attention mechanism, and emotion state discrimination of learners integrating brain-vision-language multimodal data; S3. Design a multimodal empathy response generation framework based on emotion recognition results and conversation content, build an empathy response model based on multimodal representation, and conduct joint learning and training of the empathy response model to form an intelligent agent empathy response model based on a large language model multi-agent system under the guidance of learner emotions; In S3, the specific steps for constructing an empathy response model based on multimodal representation are as follows: S31. Perform multimodal input perception and feature fusion: The conversation content between the learner and the agent is used as multimodal input and converted into a feature vector that can be solved by a large language model. The image is encoded through a pre-trained visual encoder and converted into a feature vector through BLIP and MLP. Fusion with text features; each text tag is embedded as a semantic information vector of the text , where is a real vector space, is the feature vector dimension; S32, perform multimodal output and generation: perform operations of expanding the vocabulary, text generation, and image generation to obtain a loss function optimization model, and perform retrieval enhanced image generation operations.
2. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S1, the specific steps for collecting learner brain data for emotional empathy are: S11. Collect learners’ brain data in a laboratory environment and let them watch different emotional images in the dataset to obtain EEG signals. S12. After EEG signal acquisition, invalid segments and artifacts were removed, and the frequency band of 1 Hz-75 Hz was filtered using a Butterworth filter.
3. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S1, the specific steps of conversation data collection for agent collaboration are as follows: after completing the staged classroom learning tasks, learners conduct self-perception of learning content and learning emotions, report them to the agent, have conversations and exchanges with the agent, and record the conversation context data set as , the dialogue situation is recorded as , the intelligent agent model understands the emotions in the conversation ;in, Indicates Speeches, including words, situations Indicated by The situation description consists of words, where w q Provides contextual background information.
4. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S2, the specific steps of learner facial emotion analysis based on the brain-computer coupling learning method are as follows: S211, use EEGNet and CNNNet to perform preliminary feature extraction on facial expression images and EEG signals; Among them, for facial expression images in the visual field, an improved convolutional neural network CNNNet is used to extract preliminary features and preliminary representation of the visual field : ; In the formula, For the improved convolutional neural network CNNNet, For a given visual image, For the initial representation of the visual field; For EEG signals, the feature extractor of EEGNet is used to perform preliminary feature extraction and preliminary representation of cognitive domains : ; In the formula, EEGNet is a compact convolutional neural network designed for EEG brain-computer interface. R is a given EEG signal. It is a preliminary representation of the cognitive domain; S212, construct a cognitive-visual learning framework, project the obtained preliminary representation to the public channel and the private channel, and encode the public channel and the private channel representation; Among them, the public channel adopts the encoding function of shared parameters , learning to capture the common information between cognitive and visual fields. After training, the common channel extracts the shared features of cognitive and visual fields. Given and , then the common representation of cognitive and visual domain modalities is: ; ; In the formula, is an encoding function based on a simple fully connected neural layer, for shared parameters in cognitive and visual domains; The private representation is output by the private channels in the cognitive and visual domains, capturing the private information related to the cognitive and visual domains. and , then the private representation is: ; ; In the formula, is the private encoding function of the private channel in the visual field, is the private encoding function of the private channel in the cognitive domain, This is achieved through fully connected neural layers. is the parameter of the private encoding function of the private channel in the visual field, are the parameters of the private encoding function of the private channel of the cognitive domain; S213, after obtaining the public representation and the private representation, performing a splicing operation on the obtained public representation and the private representation, and performing an emotion recognition task; Among them, the concatenated representations of the public and private representations in the visual and cognitive domains are defined as: ; ; In the formula, It is the spliced representation of the visual field. It is a representation after splicing of cognitive domains; The expressions of emotion recognition tasks in the visual domain and cognitive domain are: ; ; In the formula, and is the predicted label of the image and EEG signal, using KNN as the function Model, is the decoding function of the emotion recognition task in the visual field, It is the decoding function of emotion recognition task in cognitive domain.
5. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S2, the specific steps of emotion recognition of learner dialogue text based on the second-order interactive attention mechanism are as follows: S221, constructing explicit information pairs and implicit information pairs based on the learners' learning emotions and learning experience reports as inputs, and sending them to the information association module; S222, encode the dialogue information, regard the dialogue speech and situation between the learner and the agent as explicit information, regard the reasoning knowledge as implicit information, and process the explicit information and implicit information to obtain the semantic representation of the dialogue speech representation, the situation representation and the reasoning knowledge; S223, using the information association module to capture important associated words in explicit and implicit information, constructing an associated matrix, first-order interactive attention, second-order interactive attention, and storing associated words in memory, to obtain an associated representation, the expression of which is: ; In the formula, For related information, , is the number of associated words in the memory, For encoder; S224, input the dialogue speech representation, situation representation, association representation and reasoning knowledge representation into the aggregation network to obtain the emotion representation, and use the emotion representation to predict the emotion probability respectively; wherein the expression of the emotion probability is: ; ; ; ; In the formula, , Indicates the length The real vector space of , is the amount of emotion, represents the emotional probability for the dialogue utterances, represents the sentiment probability for the context, represents the sentiment probability for the association, Represent sentiment probability for inference knowledge, is the softmax function, The aggregation network is used to process the relevant inputs of speech representation and context representation. and represents an aggregate network with the same structure but different parameters, Used to process associated representations of related inputs, Used to process input related to reasoning knowledge representation, Speaking for the dialogue, To express the situation, For association representation, for reasoning knowledge representation; S225, multiply the emotion probability as the final emotion probability, and use the log-likelihood loss function to calculate the emotion probability and the true label Perform parameter optimization; its expression is: ; ,in, The loss function for training and optimizing the model is calculated by calculating the logarithmic difference between the predicted emotion probability and the true emotion label.
6. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S2, the specific steps of identifying the learner's emotional state by integrating brain-vision-language multimodal data are as follows: S231. Receive sentiment prediction results from the second-order attention mechanism , sentiment prediction tags based on visual images and emotion prediction labels based on EEG signals , and concatenate them into a vector ; S232, the spliced input Through multiple fully connected layers MLP processing, each layer is transformed by weight matrix and activation function to generate the final sentiment prediction result, which is expressed as: ; The MLP is a multi-layer perceptron, comprising multiple fully connected layers and activation functions.
7. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S3, the multimodal empathy response generation framework introduces three models: perception and retrieval model PRM, generation model GM, and retrieval enhancement model RAM to process multimodal data; the multimodal empathy response generation framework uses ViT-G / 14, Qformer, and linear layers to encode images, uses DALL·E2 image decoder for image decoding, and uses GPT-4 for language modeling.
8. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: S31, the feature expressions of the input text and image in the multimodal input perception and feature fusion process are: ; ; In the formula, is a contextual embedding sequence, which is provided to the language model for conditional content generation. Indicates The embedding vector of modality tags, subscript is the modal indicator, indicating the embedded modal type, and the text is , the image is ; After receiving multimodal input, the generated textual and visual tags are expressed as a joint sequence: ; In the formula, , is generated A joint sequence of textual and visual tags, It is a mixed vocabulary space containing both textual and visual tokens; In S32, the expressions of the loss functions obtained are: ; ; ; In the formula, for language modeling loss, Generate loss for the image, is the image retrieval loss, is the embedding representation of the image and represents the visual features, is the contrast loss from text to image, is the image-to-text contrast loss, is the hidden state in the generation task, is the query feature, is a frozen text encoder, is the description of the image, It’s emotional information.
9. The method for empathic response of an intelligent agent based on brain-computer coupling under the guidance of learner emotions according to claim 1, characterized in that: In S3, the empathy response model joint learning training adopts an end-to-end model training method, uses an adapter fine-tuning method for joint fine-tuning, synchronously updates a limited number of parameters in the LLM, and simultaneously updates the input linear projection layer and feature mapping module to obtain the final overall loss function; the final loss function includes the language modeling loss , image generation loss and image retrieval loss , then the overall loss function is expressed as: ; In the formula, is a hyperparameter. For perception and retrieval models, ; For the generative model, .
Citation Information
Patent Citations
Multi-round dialogue semantic comprehension subsystem based on multi-modal emotion identification system
CN108877801A
Emotion recognition method based on brain-computer modal common space
CN113974628A