Emotion monitoring and adjusting system based on real-time dialogue semantic analysis
By combining the MobileViT and CSN networks into a real-time dialogue semantic analysis module, along with the Bandit algorithm and hardware devices, the limitations of existing technologies in monitoring and regulating the emotions of both parties in a dialogue are overcome. This enables real-time monitoring and effective regulation of the emotions of both parties, improving the regulation effect and user experience.
Patent Information
- Application Number
- CN202511006138.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-10-31
AI Technical Summary
Existing real-time dialogue semantic analysis emotion monitoring and regulation technologies cannot simultaneously take into account the emotional states of both parties in the conversation. The regulation methods are too simplistic, the accuracy and real-time performance of semantic analysis are insufficient, and the hardware and software do not work closely together, resulting in poor emotion regulation effects.
A real-time dialogue semantic analysis module combining MobileViT and CSN networks is used to capture emotional associations through an attention mechanism and generate personalized adjustment strategies using the Bandit algorithm. The hardware and software work together in conjunction with bone conduction devices and LED feedback devices.
It enables the monitoring and regulation of emotions of both parties in a dialogue, improving the pertinence and real-time nature of regulation strategies. The synergy between hardware and software enhances the effectiveness of emotion regulation and user experience.
Smart Images

Figure CN120877786A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of emotion monitoring and regulation technology, specifically to an emotion monitoring and regulation system based on real-time dialogue semantic analysis. Background Technology
[0002] Emotions, as a core element of human psychological activity, play a crucial role in daily communication. They directly influence the direction, atmosphere, and final outcome of conversations. Positive emotions promote smooth and in-depth communication, while negative emotions can lead to communication barriers, escalation of conflicts, and even a series of other problems. In modern society, emotion monitoring and regulation technologies have extremely wide applications, including intelligent customer service for timely detection of customer emotions and adjustment of service strategies; education for teachers to understand students' learning emotional states; mental health for professionals to provide real-time emotion assessment data; and remote work collaboration for efficient communication among team members.
[0003] With the continuous development of technology, emotion monitoring and regulation technology based on real-time dialogue semantic analysis has emerged. Existing technical solutions usually use algorithms such as speech recognition and semantic analysis to analyze the speech content in the dialogue, extract emotion-related features such as speech rate, tone, and keywords, and combine them with machine learning models to classify and predict emotional states. At the same time, some systems will also provide corresponding regulation suggestions or intervention measures based on the analysis results.
[0004] However, existing real-time dialogue semantic analysis-based emotion monitoring and regulation technologies have several problems. First, in terms of emotion monitoring, most can only monitor a single object, failing to simultaneously consider the emotional states of both parties in the dialogue, making it difficult to comprehensively grasp the emotional interactions during the conversation. Second, the regulation methods are relatively simple, often relying on simple prompts or suggestions, lacking diverse and targeted regulation tools. Third, there are still shortcomings in the accuracy and real-time performance of semantic analysis. Due to the complexity and diversity of dialogue scenarios, existing algorithms have difficulty handling ambiguous semantics and polysemous expressions, resulting in inaccurate analysis results. At the same time, the lack of real-time processing capabilities leads to a lag in the implementation of regulation measures. In addition, the collaboration between hardware and software is not close enough, the performance of hardware devices is not fully utilized, and the software regulation strategies cannot be well adapted to the responsiveness of hardware devices, thus affecting the overall effectiveness of emotion regulation. Summary of the Invention
[0005] The purpose of this invention is to solve the above problems by designing an emotion monitoring and regulation system based on real-time dialogue semantic analysis.
[0006] The first aspect of this invention provides an emotion monitoring and regulation system based on real-time dialogue semantic analysis, the system comprising:
[0007] The information acquisition module is used to collect real-time voice information from both parties in the conversation;
[0008] The real-time dialogue semantic analysis module is used to perform semantic analysis on real-time voice information using the MobileViT network, extract emotional features, and combine the CSN network to dynamically capture emotional associations in the dialogue through an attention mechanism.
[0009] The two-way emotion regulation module is used to determine the user's current emotional state based on the output of the real-time dialogue semantic analysis module, map the emotional state to the behavioral pattern, and use the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue.
[0010] The hardware and software collaborative interaction module is used to control the bone conduction device and LED feedback device based on the emotion regulation strategy generated by the two-way emotion regulation module, so as to realize the collaborative work of hardware and software.
[0011] Optionally, in a first implementation of the first aspect of the present invention, the real-time dialogue semantic analysis module includes:
[0012] The conversion submodule is used to convert real-time speech information into Mel spectrograms, and to perform noise reduction on the Mel spectrograms using spectral subtraction to obtain the noise-reduced Mel spectrograms.
[0013] The segmentation submodule is used to segment the Mel spectrogram according to the time window to obtain an image block sequence, and input the image block sequence into the MobileViT network;
[0014] The sliding computation submodule is used by the MobileViT network to extract local features from each image block using convolutional layers. By sliding the convolutional kernel on the image block, it obtains local texture and detail information within the image block. A convolutional kernel generator is added to the convolutional layer to dynamically generate convolutional kernel parameters based on the input image block.
[0015] The capture submodule is used to input the features extracted by the convolutional layer into the Transformer encoder. The Transformer encoder models the global dependency between different image patches through a self-attention mechanism, captures the global features of speech information in the time dimension, and obtains a feature vector containing speech semantics and preliminary emotional information. An emotion perception module is added to the Transformer encoder to dynamically adjust the attention weight by analyzing the emotional dimension in the features.
[0016] The filtering submodule is used to filter and enhance feature vectors through the classification head part of the MobileViT network, highlighting the emotion-related features to obtain emotion features. The classification head includes a scene perception module and a dynamic weight adjustment module. The scene perception module determines the emotional tone of the current scene by analyzing the context information of the dialogue, and the dynamic weight adjustment module adjusts the weight parameters of feature filtering according to the output of the scene perception module, enhancing the feature dimensions related to the current scene's emotion.
[0017] Optionally, in a second implementation of the first aspect of the present invention, the real-time dialogue semantic analysis module further includes:
[0018] The input submodule is used to input the sentiment features extracted by the MobileViT network into the CSN network, where the CSN network includes a slow path and a fast path.
[0019] The parallel processing submodule is used to obtain the static and dynamic features of sentiment over time through parallel processing of slow and fast paths.
[0020] The fusion submodule is used to fuse the static and dynamic features output by the CSN network to obtain a fused sentiment feature sequence.
[0021] The weighted summation submodule is used to calculate the attention weights between features at different time points in the fused emotional feature sequence through an attention mechanism. Based on the attention weights, the fused emotional feature sequence is weighted and summed to dynamically capture the mutual influence and correlation of emotions during the dialogue process, and obtain emotional features containing emotional correlation information.
[0022] Optionally, in a third implementation of the first aspect of the present invention, the bidirectional emotion regulation module includes:
[0023] The matching submodule is used to match the output of the real-time dialogue semantic analysis module with the sentiment word vectors in the preset sentiment lexicon, calculate the similarity, and determine the current sentiment state based on the similarity.
[0024] The modeling submodule is used to model emotional states and behavioral patterns as a graph structure, dynamically calculate attention weights through the GAT network, and capture the changes in associations under different situations.
[0025] The generation submodule is used to initialize the policy pool of the Bandit algorithm for both parties in the dialogue, set the initial exploration probability and utilization probability for each policy in the policy pool, and generate personalized emotion regulation strategies for both parties in the dialogue using the Bandit algorithm.
[0026] Optionally, in a fourth implementation of the first aspect of the present invention, the step of modeling emotional states and behavioral patterns as a graph structure, dynamically calculating attention weights through a GAT network, and capturing changes in associations under different situations includes:
[0027] Emotional states and behavioral patterns are modeled as graph structures, and node embedding representations are computed in each attention head of each layer of the GAT network.
[0028] For each emotion node in the graph structure, the GAT network dynamically updates the attention weights to capture the changes in the relationship between emotion and behavior in different contexts.
[0029] After processing through a multi-layer GAT network, the final feature representations of emotion nodes and behavior nodes are fused to obtain a contextualized emotion-behavior association representation, which represents the strength and pattern of association between different emotional states and behavior patterns in the current dialogue context.
[0030] Optionally, in a fifth implementation of the first aspect of the present invention, the emotion nodes in the graph structure are connected by directed edges, the behavior nodes are connected by directed edges, and the emotion nodes and behavior nodes are connected by bidirectional edges.
[0031] Optionally, in a sixth implementation of the first aspect of the present invention, the step of computing the node embedding representation in each attention head of each layer of the GAT network includes:
[0032] The feature vector of each node in the graph structure is linearly transformed using the weight matrix to obtain the transformed feature representation.
[0033] Calculate the attention coefficient of the current node to all its neighboring nodes, and evaluate the importance of the neighboring nodes to the current node through the attention mechanism;
[0034] The transformed features of all neighboring nodes are weighted and summed to obtain the new feature representation of the current node. The outputs of multiple attention heads are concatenated to obtain the node embedding representation.
[0035] Optionally, in the seventh implementation of the first aspect of the present invention, the use of the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue includes:
[0036] The Bandit algorithm obtains the current emotional state and behavioral patterns of both parties in the conversation and selects candidate strategies from the strategy pool.
[0037] The candidate strategies are selected based on the initial exploration and utilization probabilities. After the adjustment strategy is executed, real-time feedback data from users is collected, and the immediate reward value of the strategy is calculated based on the feedback data.
[0038] The reward value is fed back to the Bandit algorithm to update the historical reward record of the corresponding strategy, and the average reward of all strategies is recalculated to adjust the exploration probability and utilization probability of each strategy in the strategy pool.
[0039] Optionally, in an eighth implementation of the first aspect of the present invention, the bone conduction device is controlled to play voice prompts or adjust music, and the LED feedback device is controlled to display corresponding colors and brightness according to the emotional state.
[0040] A second aspect of the present invention provides a method for implementing an emotion monitoring and regulation system based on real-time dialogue semantic analysis, the method comprising the following steps:
[0041] Collect real-time voice information from both parties in the conversation;
[0042] The MobileViT network is used to perform semantic analysis on real-time voice information and extract emotional features. Combined with the CSN network, the emotional associations in the dialogue are dynamically captured through the attention mechanism.
[0043] Based on the output of the real-time dialogue semantic analysis module, the user's current emotional state is determined, the emotional state is mapped to the behavioral pattern, and the Bandit algorithm is used to generate personalized emotion regulation strategies for both parties in the dialogue.
[0044] Based on the emotion regulation strategy generated by the two-way emotion regulation module, the bone conduction device and LED feedback device are controlled to achieve the coordinated work of hardware and software.
[0045] The technical solution provided by this invention collects real-time voice information from both parties in a dialogue; it uses a MobileViT network to perform semantic analysis on the real-time voice information, extracts emotional features, and combines it with a CSN network to dynamically capture emotional associations in the dialogue through an attention mechanism; based on the output of the real-time dialogue semantic analysis module, it determines the user's current emotional state, maps the emotional state to a behavioral pattern, and uses the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue; according to the emotion regulation strategies generated by the two-way emotion regulation module, it controls the bone conduction device and the LED feedback device to achieve collaborative work between hardware and software; this invention analyzes dialogue semantics and extracts emotional features, while simultaneously achieving two-way emotion regulation for both parties in the dialogue, overcoming the limitations of single-object monitoring and regulation in existing technologies. The bone conduction method does not affect normal dialogue listening, and the LED feedback intuitively displays the emotional state. The two work together to improve the effect of emotion regulation and user experience, and can generate appropriate regulation strategies according to different user situations, improving the targeting of the regulation strategies; through unique algorithms and hardware-software collaboration, it achieves real-time monitoring and effective regulation of the emotions of both parties in the dialogue. Attached Figure Description
[0046] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.
[0047] Figure 1 A schematic diagram of the structure of an emotion monitoring and regulation system based on real-time dialogue semantic analysis provided in an embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of the structure of the real-time dialogue semantic analysis module provided in an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the structure of the bidirectional emotion regulation module provided in an embodiment of the present invention. Detailed Implementation
[0050] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 The present invention provides a schematic diagram of the structure of an emotion monitoring and regulation system based on real-time dialogue semantic analysis. The system includes an information acquisition module, a real-time dialogue semantic analysis module, a two-way emotion regulation module, and a hardware and software collaborative interaction module.
[0051] The information acquisition module is used to collect real-time voice information from both parties in the dialogue. Specifically, when collecting real-time voice information from both parties, the information acquisition module uses a circular array composed of eight high-sensitivity microphones. This array uses beamforming technology to accurately locate the positions of both parties in the dialogue and effectively filter out environmental noise from other directions, so that the voice signal can be clearly captured even in noisy public places. At the same time, the array has an adaptive gain control function, which can automatically adjust the pickup sensitivity according to the volume of both parties in the dialogue, so as to avoid signal saturation caused by one party's volume being too loud or information loss caused by the other party's volume being too soft, and ensure that the acquired voice information maintains a good signal-to-noise ratio within a dynamic range of 60dB-90dB. In addition, the voice information is recorded in real time at a sampling rate of 16kHz and quantization precision, and continuously stored in 500ms time slices, providing coherent and high-quality voice data input for the subsequent real-time dialogue semantic analysis module.
[0052] The real-time dialogue semantic analysis module is used to perform semantic analysis on real-time voice information using the MobileViT network, extract emotional features, and combine it with the CSN network to dynamically capture emotional associations in the dialogue through an attention mechanism.
[0053] The two-way emotion regulation module is used to determine the user's current emotional state based on the output of the real-time dialogue semantic analysis module, map the emotional state to the behavioral pattern, and use the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue.
[0054] The hardware and software collaborative interaction module is used to control the bone conduction device and LED feedback device based on the emotion regulation strategy generated by the two-way emotion regulation module, so as to realize the collaborative work of hardware and software.
[0055] In this embodiment, the bone conduction device is controlled to play voice prompts or adjust music, and the LED feedback device is controlled to display corresponding colors and brightness according to the emotional state. The software logic in the hardware and software collaborative interaction module receives the emotion regulation strategy generated by the two-way emotion regulation module in real time, and then sends precise control commands to the bone conduction device and the LED feedback device according to the strategy content. When the strategy involves playing soothing music for an emotionally agitated client and prompting the service provider to slow down their speech, the software logic first retrieves a preset soothing music audio file. Through the bone conduction device's driver, the audio signal is converted into a bone conduction vibration signal, allowing the client to perceive the music through their skull. Simultaneously, the music volume is controlled at 30-40 decibels to avoid interfering with the client's ability to hear the service provider's speech. For the service provider, the software logic generates a voice prompt to slow down their speech, which is transmitted to them via the bone conduction device. At the same time, the software controls the LED bracelet worn by the client to display a pulsed green light, indicating that their emotions are in a state of adjustment. The software controls the LED bracelet worn by the service provider to display a stable yellow light, reminding them to adjust their speech rate. Throughout this process, the software logic monitors the hardware's operating status at 100ms intervals to ensure accurate execution of the adjustment strategy and achieve seamless coordination between software commands and hardware actions.
[0056] Please see Figure 2 The real-time dialogue semantic analysis module includes a transformation submodule, a segmentation submodule, a sliding calculation submodule, a capture submodule, a filtering submodule, an input submodule, a parallel processing submodule, a fusion submodule, and a weighted summation submodule.
[0057] The conversion submodule is used to convert real-time voice information into Mel spectrograms, and to perform noise reduction on the Mel spectrograms using spectral subtraction to obtain the noise-reduced Mel spectrograms.
[0058] The segmentation submodule is used to segment the Mel spectrogram according to the time window to obtain an image block sequence, and input the image block sequence into the MobileViT network;
[0059] The sliding computation submodule is used by the MobileViT network to extract local features from each image block using convolutional layers. By sliding the convolutional kernel on the image block, it obtains local texture and detail information within the image block. A convolutional kernel generator is added to the convolutional layer to dynamically generate convolutional kernel parameters based on the input image block.
[0060] The capture submodule is used to input the features extracted by the convolutional layer into the Transformer encoder. The Transformer encoder models the global dependency between different image patches through a self-attention mechanism, captures the global features of speech information in the time dimension, and obtains a feature vector containing speech semantics and preliminary emotional information. An emotion perception module is added to the Transformer encoder to dynamically adjust the attention weight by analyzing the emotional dimension in the features.
[0061] The filtering submodule is used to filter and enhance feature vectors through the classification head part of the MobileViT network, highlighting the emotion-related features to obtain emotion features. The classification head includes a scene perception module and a dynamic weight adjustment module. The scene perception module determines the emotional tone of the current scene by analyzing the context information of the dialogue. The dynamic weight adjustment module adjusts the weight parameters of feature filtering according to the output of the scene perception module, enhancing the feature dimensions related to the emotion of the current scene.
[0062] The input submodule is used to input the sentiment features extracted by the MobileViT network into the CSN network, where the CSN network includes a slow path and a fast path;
[0063] The parallel processing submodule is used to obtain the static and dynamic features of sentiment features over time through parallel processing of slow and fast paths.
[0064] The fusion submodule is used to fuse the static and dynamic features output by the CSN network to obtain a fused sentiment feature sequence.
[0065] The weighted summation submodule is used to calculate the attention weights between features at different time points in the fused emotional feature sequence through an attention mechanism. Based on the attention weights, the fused emotional feature sequence is weighted and summed to dynamically capture the mutual influence and correlation of emotions during the dialogue process, and obtain emotional features containing emotional correlation information.
[0066] In this embodiment, when processing the Mel spectrogram, a fixed time window parameter is first determined. This parameter is set according to the size requirements of the input image block in the MobileViT network, and is usually between 20ms and 50ms. Then, starting from the beginning of the Mel spectrogram according to the set time window, segments are extracted one by one. Each segment forms an image block until the entire Mel spectrogram is segmented, resulting in a continuous sequence of image blocks. During the segmentation process, to ensure the continuity and integrity of the image block sequence, a certain overlap area is set between adjacent time windows, with an overlap ratio of generally 20%-50%. Finally, the segmented image block sequence is input into the MobileViT network in chronological order.
[0067] The convolutional layers of the MobileViT network use convolutional kernels of a preset size, which are determined based on the resolution of the image patch and the requirements for local feature extraction. Starting from the top left corner of the image patch, the kernel slides across the patch with a set stride. At each position, it performs a convolution operation with the corresponding region of the image patch. By calculating the sum of the products of the pixel values in that region and the weights of the convolutional kernel, the local feature values of that region are obtained. Through this sliding calculation method, the entire image patch is traversed to obtain local texture and detail information at different locations within the image patch.
[0068] After receiving the local features extracted by the convolutional layer, they are organized into a feature sequence in chronological order and input into the Transformer encoder. The Transformer encoder contains multiple self-attention layers. Each self-attention layer calculates the correlation between the features of each image patch in the feature sequence. That is, it obtains the self-attention weight by calculating the correlation of the Query, Key, and Value matrices. The higher the weight, the stronger the dependency between the corresponding image patches. Based on these weights, the local features of all image patches are weighted and fused to model the global dependency between different image patches, effectively capturing the global features of speech information in the time dimension, and finally outputting a feature vector containing speech semantics and preliminary emotional information.
[0069] The classification head in the MobileViT network calls the sentiment feature filtering model learned during pre-training. This model is trained on a large amount of sentiment-annotated data and can identify key dimensions related to sentiment expression in the feature vector. The classification head analyzes the input feature vector dimension by dimension, strengthening the values of those sentiment-related feature dimensions and weakening or even suppressing feature dimensions that are not related to sentiment. For example, it highlights features corresponding to pitch changes and speech rate fluctuations in speech and filters out features that are left by background noise. After such filtering and strengthening, sentiment features focused on sentiment expression are obtained.
[0070] After receiving the sentiment features output from the MobileViT network, the system first performs format conversion and standardization to ensure that the dimensions and numerical range of the sentiment features meet the input requirements of the CSN network. Subsequently, according to the input interface protocol of the CSN network, the processed sentiment features are simultaneously transmitted to the slow path and the fast path of the CSN network. The input of the slow path uses a lower sampling frequency to extract the sentiment features, while the fast path uses a higher sampling frequency to extract them, in order to adapt to the different requirements of the two paths for feature processing speed.
[0071] The slow path processes the input sentiment features at a lower frame rate, typically 1 / 8 to 1 / 4 of the fast path's frame rate. By aggregating and analyzing features over long windows, it focuses on capturing relatively stable components of the sentiment features, such as static features like the emotional tone maintained by the speaker over a period of time. The fast path, on the other hand, processes sentiment features at a higher frame rate. Through detailed analysis over short time windows, it focuses on capturing rapidly changing parts of the sentiment features, such as dynamic features like sudden emotional fluctuations triggered by a particular sentence. The two paths work independently and in parallel, outputting the corresponding static and dynamic features respectively.
[0072] The system receives static features from the slow path and dynamic features from the fast path of the CSN network. First, it performs dimensional alignment on the two types of features to ensure that they are consistent in feature dimensions. Then, it uses a weighted fusion method to combine the static and dynamic features. The weights are set according to the importance of the static and dynamic features in the current dialogue scenario. For example, in dialogues with relatively stable emotions, the weights of static features are higher, while in dialogues with large emotional fluctuations, the weights of dynamic features are higher. Finally, it obtains a sentiment feature sequence that integrates static and dynamic information.
[0073] First, the fused sentiment feature sequence is aligned along the timeline to identify the time point corresponding to each feature. Next, an attention mechanism is used to calculate the attention weights between features at different time points. Specifically, each time point's feature is used as a query vector, and features from other time points are used as key and value vectors. The weight is obtained by calculating the similarity between the query vector and the key vector; the higher the similarity, the greater the weight, indicating a greater influence of the feature at that time point on the feature at the current query time point. Then, the features at all time points are weighted and summed based on these attention weights, enabling the fused sentiment feature sequence to highlight the emotional relationships between different time points. This dynamically captures the mutual influence of emotions during the dialogue, resulting in sentiment features containing emotional association information.
[0074] In this embodiment, a dynamic convolutional kernel generation mechanism is introduced into the convolutional layers of the MobileViT network. Traditional convolutional layers use fixed-size convolutional kernels, which cannot adapt to the feature variations of different speech segments. The improved system dynamically generates convolutional kernel parameters based on the local features of the input image patch. Specifically, a convolutional kernel generator is added to the convolutional layer. This generator generates a matching convolutional kernel size and shape by analyzing the spectral characteristics of the image patch, such as energy distribution and frequency components. For example, when the image patch contains many high-frequency details, such as rapid emotional changes, the generator generates a smaller convolutional kernel, such as 3×3, to capture the details. When the image patch mainly contains low-frequency information, such as a stable speech segment, the generator generates a larger convolutional kernel, such as 7×7, to improve processing efficiency. In this way, the convolutional layer can adaptively extract different types of local features.
[0075] This paper improves the Transformer encoder structure by introducing a hierarchical self-attention mechanism. Traditional Transformer encoders use the same attention calculation method for all features, leading to the overloading of crucial information. The improved encoder is divided into multiple layers: the base layer handles global dependencies and captures long-distance semantic associations; the enhancement layer refines emotion-related features by introducing an emotion-aware weight matrix to strengthen the expression of emotional features. Specifically, an emotion-aware module is added to the encoder. This module dynamically adjusts the attention weights by analyzing emotion-related dimensions such as pitch and speech rate in the feature vectors, making the model focus more on emotion-related information. For example, when a speech segment with high emotional intensity is detected, the emotion-aware module increases the attention weight at the corresponding position to ensure that key emotional features are not ignored.
[0076] The classification head of MobileViT has been redesigned, introducing a scene-adaptive mechanism. Traditional classification heads use fixed feature selection rules, which cannot adapt to the needs of different dialogue scenarios. The improved classification head includes a scene-aware module and a dynamic weight adjustment module. The scene-aware module determines the emotional tone of the current scene by analyzing contextual information such as the topic and the relationship between the speakers. For example, business negotiation scenarios tend to convey a tense and determined emotion, while casual conversations with friends focus more on a relaxed and pleasant emotion. The dynamic weight adjustment module adjusts the weight parameters of feature selection based on the scene-aware results, strengthening feature dimensions related to the current scene's emotion while suppressing irrelevant features. For example, in a business negotiation scenario, the classification head increases the weight of features related to "determination" and "confidence," while in a casual conversation with friends, it increases the weight of features related to "humor" and "friendliness."
[0077] In this embodiment, emotional states such as anger, joy, and anxiety, and behavioral patterns such as raising volume, frequent nodding, and remaining silent are modeled as two types of nodes in the graph. Emotional nodes are connected by directed edges, representing the possibility of emotion transfer, such as anger potentially turning into frustration. Behavioral nodes are connected by directed edges, representing the continuity of behavior, such as raising volume potentially being accompanied by increased gestures. Emotional nodes and behavioral nodes are connected by bidirectional edges, representing the behaviors that emotions may trigger and the emotions that behaviors may reflect. Each edge is assigned an initial weight, representing the strength of the association. The initial weights are set based on psychological theories and historical data statistics.
[0078] Feature vectors are constructed for each emotion node and behavior node. The feature vectors of emotion nodes include dimensions such as emotion intensity, duration, and arousal level; the feature vectors of behavior nodes include dimensions such as behavior frequency, duration, and amplitude. The emotion and behavior data extracted from real-time dialogue are standardized and mapped to a unified feature space to ensure that the feature vectors of different types of nodes are comparable.
[0079] Construct a GAT network, with each layer containing multiple attention heads; each attention head contains a learnable weight matrix and an attention mechanism; the weight matrix is used to project the input node features into a new feature space, and the attention mechanism is used to calculate the attention weights between nodes; initialize all learnable parameters, using a random initialization method to ensure that the parameters have appropriate initial values;
[0080] The constructed graph structure and node features are input into the GAT network. For each attention head in each layer, the feature vector of each node is linearly transformed through the weight matrix to obtain the transformed feature representation. The attention coefficients of the current node to all its neighboring nodes are calculated to evaluate the importance of the neighboring nodes to the current node through the attention mechanism. The transformed features of all neighboring nodes are weighted and summed to obtain the new feature representation of the current node. The outputs of multiple attention heads are concatenated or averaged to enhance the expressive power of the model. Through the processing of the multi-layer GAT network, the feature representation of the node gradually incorporates the information of its neighboring nodes, capturing the correlation in the graph structure.
[0081] When processing real-time dialogue data, the GAT network dynamically calculates attention weights based on the current emotional and behavioral node characteristics. For each emotional node, the network re-evaluates the importance of the behavioral nodes connected to it, and vice versa. For example, when the network detects that a user is angry, it may increase the attention weight of behavioral nodes such as volume and intensified body movements, while decreasing the weight of behavioral nodes such as smiling and nodding. This dynamic adjustment enables the model to capture changes in the relationship between emotions and behaviors in different contexts.
[0082] After processing by a multi-layer GAT network, the feature representation of each node contains the association information in the global graph structure; the final feature representations of emotion nodes and behavior nodes are fused, such as by concatenation or weighted summation, to obtain a contextualized emotion-behavior association representation; this representation reflects the association strength and pattern between different emotional states and behavior patterns in the current dialogue context.
[0083] A training set is constructed using historical dialogue data, with each sample containing an emotion sequence, a behavior sequence, and a corresponding contextual label during the dialogue. Loss functions such as cross-entropy loss and mean squared error loss are defined to evaluate the difference between the model's predicted correlations and the true labels. The gradient descent optimization algorithm is used to iteratively update the parameters of the GAT network and minimize the loss function. During training, a validation set is used for model selection and hyperparameter tuning to ensure that the model has good generalization ability.
[0084] Please see Figure 3 The two-way emotion regulation module includes a matching submodule, a modeling submodule, and a generation submodule.
[0085] The matching submodule is used to match the output of the real-time dialogue semantic analysis module with the emotional word vectors in the preset emotional word library, calculate the similarity, and determine the current emotional state based on the similarity.
[0086] The modeling submodule is used to model emotional states and behavioral patterns as a graph structure, dynamically calculate attention weights through the GAT network, and capture the changes in associations under different situations.
[0087] The generation submodule is used to initialize the policy pool of the Bandit algorithm for both parties in the dialogue, set the initial exploration probability and utilization probability for each policy in the policy pool, and use the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue.
[0088] In this embodiment, independent strategy pools are established for both Party A and Party B in the dialogue. These pools contain various preset emotion regulation strategies, such as playing music of different styles, providing text prompts with different content, and providing voice feedback at different frequencies. These strategies are designed based on common emotion regulation scenarios and methods to ensure their diversity and relevance. For each strategy in both parties' strategy pools, an initial exploration probability and utilization probability are set, with the sum of the two being 100%. Initially, the exploration probability is set to a higher value, and the utilization probability is set to a lower value to ensure that the algorithm can fully try different strategies and accumulate sufficient feedback data in the early stages. For example, the strategy of playing soothing music in Party A's strategy pool has an initial exploration probability of 70% and a utilization probability of 30%. The strategies in Party B's strategy pool are also set according to this rule.
[0089] The Bandit algorithm obtains the current emotional state, intensity, and corresponding behavioral patterns of both parties in the conversation. The algorithm receives information in real time and clarifies the emotional state, emotional intensity level, and corresponding typical behavioral patterns of Party A and Party B. For example, the emotional state is negative for Party A and neutral for Party B. The emotional intensity level is moderate for Party A and slight for Party B. The behavioral pattern is faster speech for Party A and silence for Party B.
[0090] The Bandit algorithm filters candidate strategies based on the current information of both parties. For Party A, the algorithm selects strategies from its strategy pool that match the combination of negative-moderate intensity-increased speech rate, such as playing soothing music or sending reassuring text prompts. For Party B, the algorithm selects strategies from its strategy pool that match the combination of neutral-weak intensity-silence, such as sending guiding text prompts or reducing the frequency of voice feedback, thus forming a subset of candidate strategies for each party.
[0091] The Bandit algorithm selects strategies based on exploration and utilization probabilities. For Party A's subset of candidate strategies, the selection is made according to the initial exploration and utilization probabilities of each strategy: with a 70% probability, one strategy is randomly selected from the candidate strategies, such as randomly selecting to send a soothing text prompt; with a 30% probability, the strategy with the highest historical average reward in the subset is selected, such as if playing soothing music previously yielded the highest reward in similar scenarios; Party B's strategy selection process is the same as Party A's, based on the probability rules of its subset of candidate strategies.
[0092] The algorithm executes the selected strategy and updates the probabilities. After executing the selected emotion regulation strategy for both parties, the algorithm collects real-time feedback data from both parties, such as changes in emotion intensity and whether behavioral patterns have improved. Based on the feedback, the algorithm calculates the reward value of the strategy. If the emotion intensity decreases, the reward value is positive; otherwise, it is negative. Subsequently, the algorithm updates the utilization probability and exploration probability of the strategy based on the reward value: if the reward value is positive, the utilization probability is increased and the exploration probability is decreased; if the reward value is negative, the utilization probability is decreased and the exploration probability is increased, providing an updated probability basis for the next strategy selection.
[0093] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. A sentiment monitoring and regulation system based on real-time dialogue semantic analysis, characterized in that, include: The information acquisition module is used to collect real-time voice information from both parties in the conversation; The real-time dialogue semantic analysis module is used to perform semantic analysis on real-time voice information using the MobileViT network, extract emotional features, and combine the CSN network to dynamically capture emotional associations in the dialogue through an attention mechanism. The two-way emotion regulation module is used to determine the user's current emotional state based on the output of the real-time dialogue semantic analysis module, map the emotional state to the behavioral pattern, and use the Bandit algorithm to generate personalized emotion regulation strategies for both parties in the dialogue. The hardware and software collaborative interaction module is used to control the bone conduction device and LED feedback device based on the emotion regulation strategy generated by the two-way emotion regulation module, so as to realize the collaborative work of hardware and software.
2. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 1, characterized in that, The real-time dialogue semantic analysis module includes: The conversion submodule is used to convert real-time speech information into Mel spectrograms, and to perform noise reduction on the Mel spectrograms using spectral subtraction to obtain the noise-reduced Mel spectrograms. The segmentation submodule is used to segment the Mel spectrogram according to the time window to obtain an image block sequence, and input the image block sequence into the MobileViT network; The sliding computation submodule is used by the MobileViT network to extract local features from each image block using convolutional layers. By sliding the convolutional kernel on the image block, it obtains local texture and detail information within the image block. A convolutional kernel generator is added to the convolutional layer to dynamically generate convolutional kernel parameters based on the input image block. The capture submodule is used to input the features extracted by the convolutional layer into the Transformer encoder. The Transformer encoder models the global dependency between different image patches through a self-attention mechanism, captures the global features of speech information in the time dimension, and obtains a feature vector containing speech semantics and preliminary emotional information. An emotion perception module is added to the Transformer encoder to dynamically adjust the attention weight by analyzing the emotional dimension in the features. The filtering submodule is used to filter and enhance feature vectors through the classification head part of the MobileViT network, highlighting the emotion-related features to obtain emotion features. The classification head includes a scene perception module and a dynamic weight adjustment module. The scene perception module determines the emotional tone of the current scene by analyzing the context information of the dialogue, and the dynamic weight adjustment module adjusts the weight parameters of feature filtering according to the output of the scene perception module, enhancing the feature dimensions related to the current scene's emotion.
3. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 1, characterized in that, The real-time dialogue semantic analysis module also includes: The input submodule is used to input the sentiment features extracted by the MobileViT network into the CSN network, where the CSN network includes a slow path and a fast path. The parallel processing submodule is used to obtain the static and dynamic features of sentiment over time through parallel processing of slow and fast paths. The fusion submodule is used to fuse the static and dynamic features output by the CSN network to obtain a fused sentiment feature sequence. The weighted summation submodule is used to calculate the attention weights between features at different time points in the fused emotional feature sequence through an attention mechanism. Based on the attention weights, the fused emotional feature sequence is weighted and summed to dynamically capture the mutual influence and correlation of emotions during the dialogue process, and obtain emotional features containing emotional correlation information.
4. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 1, characterized in that, The bidirectional emotion regulation module includes: The matching submodule is used to match the output of the real-time dialogue semantic analysis module with the sentiment word vectors in the preset sentiment lexicon, calculate the similarity, and determine the current sentiment state based on the similarity. The modeling submodule is used to model emotional states and behavioral patterns as a graph structure, dynamically calculate attention weights through the GAT network, and capture the changes in associations under different situations. The generation submodule is used to initialize the policy pool of the Bandit algorithm for both parties in the dialogue, set the initial exploration probability and utilization probability for each policy in the policy pool, and generate personalized emotion regulation strategies for both parties in the dialogue using the Bandit algorithm.
5. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 4, characterized in that, The process of modeling emotional states and behavioral patterns as a graph structure, dynamically calculating attention weights through a GAT network, and capturing changes in associations across different contexts includes: Emotional states and behavioral patterns are modeled as graph structures, and node embedding representations are computed in each attention head of each layer of the GAT network. For each emotion node in the graph structure, the GAT network dynamically updates the attention weights to capture the changes in the relationship between emotion and behavior in different contexts. After processing through a multi-layer GAT network, the final feature representations of emotion nodes and behavior nodes are fused to obtain a contextualized emotion-behavior association representation, which represents the strength and pattern of association between different emotional states and behavior patterns in the current dialogue context.
6. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 5, characterized in that, In the graph structure, emotion nodes are connected by directed edges, behavior nodes are connected by directed edges, and emotion nodes and behavior nodes are connected by bidirectional edges.
7. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 5, characterized in that, The calculation of node embedding representations in each attention head of each layer of the GAT network includes: The feature vector of each node in the graph structure is linearly transformed using the weight matrix to obtain the transformed feature representation. Calculate the attention coefficient of the current node to all its neighboring nodes, and evaluate the importance of the neighboring nodes to the current node through the attention mechanism; The transformed features of all neighboring nodes are weighted and summed to obtain the new feature representation of the current node. The outputs of multiple attention heads are concatenated to obtain the node embedding representation.
8. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 4, characterized in that, The Bandit algorithm is used to generate personalized emotion regulation strategies for both parties in the conversation, including: The Bandit algorithm obtains the current emotional state and behavioral patterns of both parties in the conversation and selects candidate strategies from the strategy pool. The candidate strategies are selected based on the initial exploration and utilization probabilities. After the adjustment strategy is executed, real-time feedback data from users is collected, and the immediate reward value of the strategy is calculated based on the feedback data. The reward value is fed back to the Bandit algorithm to update the historical reward record of the corresponding strategy, and the average reward of all strategies is recalculated to adjust the exploration probability and utilization probability of each strategy in the strategy pool.
9. The emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 1, characterized in that, Control the bone conduction device to play voice prompts or adjust music, and control the LED feedback device to display corresponding colors and brightness according to emotional state.
10. A method for implementing an emotion monitoring and regulation system based on real-time dialogue semantic analysis as described in claim 1, characterized in that, The method includes the following steps: Collect real-time voice information from both parties in the conversation; The MobileViT network is used to perform semantic analysis on real-time voice information and extract emotional features. Combined with the CSN network, the emotional associations in the dialogue are dynamically captured through the attention mechanism. Based on the output of the real-time dialogue semantic analysis module, the user's current emotional state is determined, the emotional state is mapped to the behavioral pattern, and the Bandit algorithm is used to generate personalized emotion regulation strategies for both parties in the dialogue. Based on the emotion regulation strategy generated by the two-way emotion regulation module, the bone conduction device and LED feedback device are controlled to achieve the coordinated work of hardware and software.
Citation Information
Cited By
An instant messaging interaction optimization method, system, client and electronic device based on bidirectional emotional compensation
CN122601626A