Dialogue interaction system based on multi-modal emotion perception and knowledge graph dynamic enhancement
Through a dialogue interaction system with multimodal emotion perception and dynamic enhancement of knowledge graphs, the problems of non-text information capture and dynamic association of knowledge graphs in artificial intelligence dialogue systems are solved, and efficient response and fluency of intelligent dialogue interaction are achieved.
Patent Information
- Application Number
- CN202511274967.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing artificial intelligence dialogue systems have difficulty capturing non-text information, resulting in stiff interactions and a lack of empathy. Traditional knowledge graphs are unable to dynamically associate real-time information with user emotions, and lack in-depth knowledge mining and interactive resonance.
Through a conversational interaction system that dynamically enhances multimodal emotion perception and knowledge graphs, we use user edge terminals to acquire multimodal data, employ a lightweight Transformer fusion network for feature fusion, and use the cloud service platform's dynamic knowledge graph enhancement model to mine emotional relationships and entity knowledge in real time, building an intelligent reply prototype.
It achieves efficient and intelligent response of the dialogue system, improves the targeted response and interaction fluency, and enhances user satisfaction and natural interaction.
Smart Images

Figure CN120745853A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence. More specifically, the embodiments of the present application relate to a dialogue interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs. Background Art
[0002] At present, with the development of artificial intelligence technology, artificial intelligence dialogue systems are becoming more and more popular.
[0003] In related technologies, chatbots (such as customer service robots, Siri, Alexa, etc.) rely primarily on text semantic analysis, making it difficult to capture non-text information that is crucial in communication (such as the emotions implied in voice intonation, the psychological state represented by facial micro-expressions, the pressure implied in the voice, etc.). This leads to awkward interactions, a lack of empathy, and a tendency to misunderstand the user's true intentions and emotional state, especially in complex, emotionally charged conversation scenarios (such as psychological counseling, complaint handling, and companionship care). Furthermore, in related technologies, traditional knowledge graphs are often static or pre-set references in conversations, making it difficult to dynamically associate real-time information in the conversation context with potential knowledge needs. They lack the ability to dynamically mine relevant knowledge and make analogies based on the current emotional state, resulting in answers that are relevant but may not be in-depth or meet user needs, leading to a lack of resonance during user interaction.
[0004] Therefore, there is an urgent need to design a more efficient dialogue interaction solution to solve at least one of the above technical problems. Summary of the Invention
[0005] In this context, the embodiments of the present application hope to provide a dialogue interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs, which can realize intelligent dialogue interaction and improve the targeted response and interaction fluency.
[0006] In a first aspect of the embodiments of the present application, a conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs is provided. The conversational interaction system includes at least a user edge terminal and a cloud service platform. The system includes:
[0007] Obtain multimodal data generated by users during the interaction process through user edge terminals;
[0008] Through the user edge terminal, a lightweight Transformer fusion network is used to convert user feature vectors under different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment. The fused feature vector and historical conversations within the preset rounds are compressed in real time and uploaded to the cloud.
[0009] The received compressed data is input into the dynamic knowledge graph enhancement model DKGE through the cloud service platform. During the dialogue interaction, the fused feature vector and the entity knowledge and emotional relationships implicit in the historical dialogue are mined in real time to update the dynamic knowledge graph. Based on the updated dynamic knowledge graph, knowledge enhancement processing is performed to construct the initial response prototype for the current round of interaction with the user and send it to the user's edge terminal.
[0010] Through the cloud service platform, the fused feature vector and the updated dynamic knowledge graph are input into the dialogue strategy model to determine the response strategy and knowledge call direction for the current round of dialogue, and then sent to the user edge terminal;
[0011] Through the user edge terminal, real-time interactive reply information is generated and output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction.
[0012] In a second aspect of the embodiments of the present application, a conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs is provided, comprising:
[0013] The user edge terminal is used to obtain multimodal data generated by users during the interaction process. It uses a lightweight Transformer fusion network to convert user feature vectors under different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment. The fused feature vector and historical conversations within a preset round are compressed in real time and uploaded to the cloud.
[0014] The cloud service platform is configured to input the received compressed data into the dynamic knowledge graph enhancement model DKGE, mine the fused feature vector and the entity knowledge and sentiment relationships implicit in the historical conversations in real time during the conversation interaction to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial response prototype for the current round of interaction with the user, and send it to the user's edge terminal;
[0015] The cloud service platform is further configured to input the fused feature vector and the updated dynamic knowledge graph into the dialogue strategy model, determine the response strategy and knowledge call direction for the current dialogue round, and send them to the user edge terminal;
[0016] The user edge terminal is further configured to generate real-time interactive response information output to the user based on the initial response prototype, the response strategy, and the knowledge calling direction.
[0017] In a third aspect of the implementation of the present application, a terminal device is provided, comprising: at least one processor, a memory, and an input-output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute any one of the first aspects of the dialogue interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs.
[0018] In a fourth aspect of the embodiments of the present application, a computer-readable storage medium is provided, which includes instructions that, when executed on a computer, enable the computer to execute the dialogue interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs as described in any one of the first aspects.
[0019] In a fifth aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the dialogue interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs as described in any one of the first aspects.
[0020] According to an embodiment of the present application, a conversational interaction system based on multimodal emotion perception and dynamic enhancement of a knowledge graph is applied to a cloud-edge collaborative conversational interaction system, which includes at least a user edge terminal and a cloud service platform. In this embodiment, the user edge terminal obtains multimodal data generated by the user during the interaction process; a lightweight Transformer fusion network is used to convert user feature vectors in different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment; the fused feature vector and historical conversations within a preset round are compressed in real time and uploaded to the cloud. Furthermore, the cloud service platform inputs the received compressed data into a dynamic knowledge graph enhancement model (DKGE). During the conversational interaction, the fused feature vector and the entity knowledge and emotional relationships implicit in the historical conversations are mined in real time to update the dynamic knowledge graph. Based on the updated dynamic knowledge graph, knowledge enhancement processing is performed to construct an initial response prototype for the current round of interaction with the user and send it to the user edge terminal. The fused feature vector and the updated dynamic knowledge graph are input into a conversational strategy model to determine the response strategy and knowledge call direction for the current round of conversation, which are then sent to the user edge terminal. Finally, through the user edge terminal, real-time interactive reply information is generated and output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction.
[0021] In this application, multimodal data acquisition breaks through the limitations of single text, comprehensively capturing multimodal interaction information such as user voice and expressions. This allows for a multi-dimensional understanding of user intent and lays a data foundation for subsequent processing. A lightweight Transformer fusion network is then used to fuse multimodal features using cross-modal attention and a dynamic weighting mechanism, addressing data heterogeneity. This reduces edge computing load while improving feature semantic richness, providing high-quality fused features for subsequent steps. By jointly compressing historical conversations and fused features, data transmission volume is reduced, network latency is lowered, and system response speed is improved, while retaining key information to support cloud processing. Next, a dynamic knowledge graph enhancement model mines entity knowledge and emotional relationships in conversations in real time, dynamically updating the knowledge graph to form a knowledge system tailored to user needs and providing precise knowledge support for reply prototype construction. Based on the updated knowledge graph and fused features, natural language generation technology is used to construct initial reply prototypes containing core semantics, improving the logic and knowledge content of replies. A dialogue strategy model combines user status and knowledge graph information to determine response strategies and knowledge retrieval directions, enabling intelligent dialogue process control and improving response relevance and interaction fluency. Finally, the edge terminal integrates the initial prototype, response strategy, etc., and optimizes the generation of natural language responses that are both knowledge-rich and emotionally adaptable, thereby improving the naturalness of interaction and user satisfaction, and enabling the system to respond efficiently and intelligently to user needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A flowchart of a conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs provided in one embodiment of the present application;
[0023] Figure 2 A schematic diagram of the structure of a conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs provided in one embodiment of the present application;
[0024] Figure 3 The structural diagram of a medium in an embodiment of the present application is schematically shown. DETAILED DESCRIPTION
[0025] Reference below Figure 1 , Figure 1 A flowchart of a conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs provided in one embodiment of the present application.
[0026] Figure 1 The process of the conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs provided in one embodiment of the present application includes:
[0027] Step S101: acquiring multimodal data generated by the user during the interaction process through the user edge terminal;
[0028] Step S102: Using a lightweight Transformer fusion network at the user edge terminal, user feature vectors under different modalities are converted into fused feature vectors through cross-modal attention fusion and dynamic weight adjustment.
[0029] Step S103: compress the fused feature vector and historical conversations within a preset round in real time and upload them to the cloud;
[0030] In step S104, the received compressed data is input into the dynamic knowledge graph enhancement model DKGE through the cloud service platform. During the dialogue interaction, the fused feature vector and the entity knowledge and emotional relationships implicit in the historical dialogue are mined in real time to update the dynamic knowledge graph. Based on the updated dynamic knowledge graph, knowledge enhancement processing is performed to construct an initial response prototype for the current round of interaction with the user, and the prototype is sent to the user's edge terminal.
[0031] Step S105: Input the fused feature vector and the updated dynamic knowledge graph into the dialogue strategy model through the cloud service platform, determine the response strategy and knowledge call direction for the current dialogue round, and send them to the user edge terminal;
[0032] Step S106 : generating real-time interactive reply information output to the user through the user edge terminal according to the initial reply prototype, the response strategy, and the knowledge call direction.
[0033] In the embodiment of the present application, the conversation interaction system is applied to cloud-edge collaboration. Further optionally, the conversation interaction system includes at least a user edge terminal and a cloud service platform.
[0034] Specifically, the user edge terminal is the front-end device that directly interfaces with the user in the conversational interaction system, fulfilling the core functions of multimodal data collection and initial processing. At the hardware level, it integrates microphones, cameras, and various sensors to capture multimodal data generated by users during interactions, including voice, text, expressions, and gestures, in real time, enabling comprehensive perception of user interaction information. Algorithmically, it utilizes a lightweight Transformer fusion network. Through a cross-modal attention mechanism and dynamic weight adjustment strategy, it fuses feature vectors from different modalities into a unified semantic representation, effectively addressing the heterogeneity of multimodal data. Furthermore, the terminal is responsible for real-time compression of historical conversations and fused features within a preset round, reducing data volume by removing redundant information and preparing for subsequent transmission to the cloud. Furthermore, the edge terminal receives information such as initial response prototypes and response strategies from the cloud, and combines this with a local language generation model and emotion rendering module to generate real-time interactive responses for the user, providing feedback through output devices such as speakers and screens.
[0035] The cloud service platform is the core brain of the conversational interaction system, primarily responsible for knowledge processing, policy decision-making, and response framework construction. During the data reception phase, it receives compressed data from the user's edge terminal, including fused feature vectors and historical conversation information, providing data support for subsequent knowledge graph updates. The core module, the Dynamic Knowledge Graph Enhancement (DKGE) model, analyzes this data in real time. Using techniques such as graph neural networks, it mines implicit entity knowledge and sentiment relationships, dynamically updating the nodes and edges of the knowledge graph, enabling the knowledge graph to evolve over the course of the conversation and form a knowledge system tailored to user needs. Graph neural networks, such as GAT, use an attention mechanism to calculate weights between nodes in the graph and are suitable for processing node interactions and feature transfer within the knowledge graph. Based on the updated knowledge graph, the cloud utilizes natural language generation technology to construct an initial response prototype. This prototype contains the core semantics and response framework that matches the user's current interaction. Simultaneously, the conversation policy model integrates feature vectors and dynamic knowledge graph information, using methods such as reinforcement learning to determine the response strategy (e.g., inquiry, answer, guidance) and knowledge retrieval direction (e.g., domain knowledge, user preferences, etc.) for the current turn, enabling intelligent control of the conversational flow. Finally, the cloud sends the initial response prototype and policy decision results to the edge terminal, providing key input for the terminal to generate the final response.
[0036] Entity knowledge, which is structured semantic units extracted from conversations, such as specific concepts or objects like weather and mood, is the basic node of the knowledge graph. Emotional relationships are associations between entities with emotional polarity, such as the "cause" relationship in "rain → cause → bad mood" with the emotional intensity of "low."
[0037] In practical applications, cloud-edge collaboration offloads some data processing tasks to edge terminals, preventing large amounts of raw data from being directly uploaded to the cloud. For example, in a dialogue system, the edge terminal first performs feature fusion and compression on multimodal data, transmitting only key information to the cloud, which can significantly reduce network traffic. This edge preprocessing mode and cloud-based deep processing mode can reduce data transmission latency and ensure real-time interaction. For example, user voice commands can be quickly converted into feature vectors at the edge, without waiting for the cloud to fully parse the original audio, improving dialogue response speed.
[0038] Specifically, the cloud boasts powerful computing and storage resources, making it suitable for tasks such as large-scale knowledge graph updates and complex policy decisions. Edge terminals, on the other hand, excel at local real-time data collection and lightweight computation (such as multimodal feature fusion). For example, in a conversational system, the edge utilizes lightweight Transformers to process multimodal data, preventing cloud resources from being consumed by massive amounts of raw data and allowing the cloud to focus on deep mining of dynamic knowledge graphs and building response frameworks. This division of labor maximizes the advantages of cloud-edge resources and reduces overall system energy consumption and computing costs.
[0039] Edge terminals can perform some data processing locally, maintaining basic interactive functions (such as generating simple responses based on locally cached knowledge) even during network outages, preventing complete system failure. Furthermore, sensitive data (such as user expressions and voice features) can be desensitized or feature extracted at the edge, reducing the risk of uploading raw data to the cloud. For example, transmitting only compressed feature vectors to the cloud, rather than complete audio and video, combined with local edge encryption technology, can effectively enhance user privacy protection.
[0040] By integrating data from multiple edge devices, the cloud builds and continuously optimizes a global knowledge graph. Edge devices can also provide personalized knowledge supplements based on local user interaction history. For example, in a conversational system, edge devices record users' emotional preferences in specific scenarios, upload them to the cloud, and dynamically update the knowledge graph, enabling the cloud to provide response strategies more tailored to the user's habits. This achieves a two-way synergy between personalized edge perception and global knowledge enhancement in the cloud, enhancing the personalization and intelligence of interactions.
[0041] In the implementation of this application, multimodal data acquisition overcomes the limitations of a single text modality, comprehensively capturing user interaction information and understanding user intent from multiple dimensions, including language expression, emotional state, and behavioral characteristics, laying an accurate data foundation for subsequent analysis and processing. Furthermore, a lightweight Transformer fusion network effectively addresses the heterogeneity of multimodal data, reducing the computational load on edge terminals while improving the integrity and semantic richness of feature representation, ensuring real-time processing and providing high-quality fused features for subsequent steps. Real-time compression and upload of historical conversations reduces data transmission volume, reduces network bandwidth usage and latency, and improves system response speed, while preserving key information and balancing data integrity and transmission efficiency, enabling the cloud to operate based on valid data. Next, the Dynamic Knowledge Graph Enhancement (DKGE) model allows the knowledge graph to evolve in real time with conversations, accumulating relevant entity knowledge and emotional associations to form a knowledge system tailored to user needs. This provides precise and dynamic knowledge support for reply prototype construction, enhancing the system's knowledge representation and reasoning capabilities. The initial response prototype is constructed based on the updated knowledge graph and fused features, leveraging natural language generation technology to generate a response framework with rich knowledge support and semantic accuracy. This provides a foundation for subsequent optimization and improves the logic and knowledge content of the response. The dialogue strategy model determines the response strategy and knowledge retrieval direction based on the user's real-time status and knowledge graph information, enabling intelligent interaction process control, accurately acquiring the required knowledge, improving the relevance and effectiveness of responses, and enhancing the fluency and user experience of the conversation. Finally, real-time interactive response information generation integrates the results of various components to generate responses that are rich in knowledge, meet policy requirements, and reflect the user's emotional state. This is then fed back to the user through natural language expression, enhancing the naturalness of the interaction and satisfaction, and making the entire system efficient and intelligent in responding to user needs.
[0042] In step S101, the user edge terminal acquires multimodal data generated during user interaction. Specifically, the user edge terminal implements multimodal data acquisition through hardware devices and sensor networks. Microphones capture voice signals and convert them into audio waveform data. Cameras collect visual information such as facial expressions and body movements. Sensors (such as accelerometers and gyroscopes) record user gestures. Text input devices (keyboards and touchscreens) capture text commands. After analog-to-digital conversion of this raw data, the edge terminal performs noise reduction and normalization through a preprocessing module to generate a structured multimodal data sequence. For example, voice data is framed to extract Mel-Frequency Cepstral Coefficients (MFCCs), and image data is processed through a convolutional neural network to extract visual features. Ultimately, data from different modalities is stored as feature vectors, providing a foundation for subsequent cross-modal fusion.
[0043] For example, when a user interacts with an intelligent customer service terminal, the terminal's microphone records the user's voice in real time (such as "Help me check tomorrow's weather"), and the camera captures the user's facial expressions when speaking (such as a slight frown may indicate concern). At the same time, if the user holds a device with an accelerometer, the sensor will record their gestures (such as pointing out the window). These data are synchronously transmitted to the pre-processing module of the edge terminal: the voice signal is converted into a spectrogram and acoustic features are extracted, the facial image is used to identify the expression type through key point detection, and the gesture data is normalized into a motion trajectory vector. Ultimately, a multimodal dataset containing speech text features, expression features, and action features is formed for subsequent analysis of user intentions and emotional states.
[0044] Through the above-mentioned step S101, it is possible to break through the limitation of single text input and obtain user interaction information from dimensions such as voice intonation, facial expressions, and body movements. For example, when a user says "it's okay", combined with a frowning expression, it can be judged that his true emotion is negative, thus avoiding semantic understanding deviation. The edge terminal processes raw data locally to reduce the amount of data uploaded and the delay. For example, after the voice data is feature extracted at the edge, only the feature vector needs to be transmitted instead of the complete audio, thereby improving the system response speed. It is suitable for interaction needs in complex environments, such as combining lip reading visual features to assist voice recognition in noisy scenes, or replacing text input with gestures in barrier-free interaction, thereby expanding the scope of application of the system.
[0045] As an optional embodiment, it is assumed that the multimodal data includes at least: voice data, video data, text data, and biological signal data collected by wearable devices. It is assumed that the lightweight Transformer fusion network includes at least: feature extraction layer, spatiotemporal alignment layer, fusion layer, and adaptive weight allocation layer. Based on the above assumptions, in step S102, through the user edge terminal, a lightweight Transformer fusion network is used to convert user feature vectors under different modalities into fused feature vectors through cross-modal attention fusion and dynamic weight adjustment, including:
[0046] In the feature extraction layer, for different modal data, a voice emotion feature extraction module, a speech intonation feature extraction module, a micro-expression feature recognition module, a text emotion semantic analysis module, and a physiological change feature recognition module are respectively constructed; and through the constructed different modal processing modules, user feature vectors under different modalities are extracted from the multimodal data; wherein the user feature vectors include at least: the user's voice emotion features, speech intonation features, micro-expression features, text emotion semantic features, and physiological change features;
[0047] Through the spatiotemporal alignment layer and the fully connected layer, the user feature vector is projected into a unified dimensional space. The emotional semantic nodes in the text emotional semantic features are used as the time axis reference, and the user feature vectors under other modalities are positioned to each emotional semantic node to construct a spatiotemporally consistent emotional expression view.
[0048] Through the fusion layer, combined with the emotional expression view, the emotional semantic nodes in the text emotional semantic features are used as the time axis benchmark. Different user feature vectors are used to perform cross-modal retrieval on other user feature vectors according to the time sequence of the time axis benchmark to obtain corresponding matching feature vectors. The fused feature vectors under different branches are then obtained based on the cross-modal retrieval results.
[0049] Through the adaptive weight allocation layer, the signal quality evaluation and semantic consistency detection are performed on the fused feature vectors under different branches. Based on the evaluation and detection results, the weight parameters corresponding to different modalities are dynamically adjusted to obtain the final output fused feature vector, so as to eliminate data conflicts in the fused feature vector and improve the accuracy of the output results.
[0050] In the embodiments of the present application, the lightweight Transformer fusion network adopts a lightweight variant based on the Transformer architecture, such as MobileViT, TinyBERT, etc. By reducing the number of attention heads and the hidden layer dimension to balance the computational efficiency and cross-modal feature interaction capabilities, a lightweight version of a cross-modal attention network such as CLIP can also be introduced to align text, speech, vision and other modal features through a comparative learning mechanism, and optimize the parameter scale for edge computing scenarios. It is also possible to draw on the idea of Mixture of Experts (MoE) to construct a dynamic weight fusion model, use a lightweight neural network to calculate the reliability score of each modality, and realize the intelligent weighting of conflicting features.
[0051] In the aforementioned network, the voice emotion feature extraction module converts speech into a spectrogram through short-time Fourier transform, extracts frequency domain features using Mel-frequency cepstral coefficients (MFCC), and then uses lightweight CNNs such as MobileNet to capture acoustic patterns corresponding to emotions such as anger and sadness. For example, when a user says "I'm fine," the module can extract anxiety features from the spectrum of the trembling voice. The speech intonation feature extraction module uses a bidirectional LSTM to analyze prosodic features such as fundamental frequency and speech rate. For example, it extracts the temporal features of rising intonation in the interrogative tone of "Really?" The micro-expression feature recognition module uses a lightweight convolutional network (such as ShuffleNet) to detect facial landmarks and locate micro-expressions such as drooping eyelids and twitching mouth corners, capturing, for example, the downward movement of the eye corners when a user speaks. The text sentiment semantic analysis module uses a lightweight version of BERT (such as ALBERT) to extract semantic vectors for emotional words such as "sad" and "disappointed" and combines syntactic analysis to construct semantic nodes. The physiological change feature recognition module uses wavelet transforms to extract stress response features, such as the physiological signal of a sudden increase in heart rate, from heart rate and electrodermal signals collected by wearable devices. These modules process multimodal data in parallel, providing multidimensional feature vectors for subsequent fusion. This allows for comprehensive feature capture from speech prosody to physiological responses, avoiding the semantic understanding bias of a single modality.
[0052] The spatiotemporal alignment layer projects features from various modalities into a unified dimensional space through a fully connected layer. Using the emotional semantic nodes within the text's emotional semantic features (e.g., the text fragment "I'm sad") as the timeline reference, it uses the Dynamic Time Warping (DTW) algorithm to align features such as speech pauses and micro-expressions to corresponding semantic nodes. For example, when a user utters the text "I'm sad," the sobbing features extracted by the speech module and the eyelid drooping captured by the vision module are synchronously mapped to the timestamp of the text, forming a spatiotemporal matrix that encompasses the text's semantics, speech pauses, and facial expressions. This text-anchored alignment resolves the issue of asynchronous sampling rates among multimodal signals, constructing a spatiotemporally consistent view of emotional expression. This allows the system to correlate emotional signals from different modalities temporally. For example, when analyzing the text "I'm fine," the system can combine a smiling expression with the physiological signal of an elevated heart rate during the same time period to determine the conflict between the user's true emotion and the text's semantics.
[0053] The fusion layer uses the timeline of text semantic nodes as a reference and implements feature interaction through a cross-modal attention mechanism. For each text semantic node (e.g., "I'm disappointed"), the speech feature vector is used as a query vector to retrieve the corresponding micro-expression matching vector (e.g., a forced smile) from the visual features. Simultaneously, the text semantic vector also retrieves the corresponding stress response vector (e.g., sudden changes in skin conduction signals) from the physiological features. This cross-modal search uses cosine similarity or dot product operations to identify semantically related feature pairs across modalities. Branched fusion features are then generated through concatenation or weighted summation. For example, when processing the text node "I'm disappointed," the fusion layer cross-modally correlates the falling intonation feature from speech, the smiling expression feature from visuals, and the stable heart rate feature from physiological features. This identifies conflicting signals between the speech and visual modalities and provides a basis for subsequent weighting adjustments. This mechanism enables features from different modalities to complement each other on the timeline, enhancing the representation of complex emotions (e.g., masked emotions) and avoiding misjudgments caused by single-modality analysis.
[0054] The adaptive weight assignment layer handles feature conflicts through a dual mechanism. First, it uses metrics such as the signal-to-noise ratio (SNR) to assess the signal quality of each modality. For example, a speech feature with a low SNR in a noisy environment will be assigned a lower weight. Second, a semantic consistency detection model (such as a lightweight BERT classifier) calculates the probability of matching features from different modalities with the semantics of the text. For example, when a user smiles and says "I'm disappointed," the model combines the user's historical "social smile" behavior in the knowledge graph to calculate the probability of semantic conflict between the visual modality feature "smile" and the text "disappointment," thereby reducing the weight of the visual modality. Weight adjustment is performed through dynamic optimization using gradient descent. The final output fused feature vector weakens conflicting modalities (such as the visual characteristics of a social smile) and strengthens consistent modalities (such as the disappointment characteristic of speech intonation). Taking "Smile and Say Disappointment" as an example, this layer will increase the weight of voice features to 0.6, reduce the weight of visual features to 0.3, and maintain the weight of physiological features at 0.1 through historical data reasoning and contextual probability calculation, thereby eliminating modal conflicts, making the fused features more accurately reflect the user's true emotions, and improving the accuracy and robustness of emotion recognition.
[0055] Further optionally, in the above steps, extracting user feature vectors in different modalities from the multimodal data by constructing different modal processing modules includes:
[0056] Through the voice emotion feature extraction module, the user's voice emotion features are extracted from the voice data based on the Mel spectrum and the CNN attention mechanism; through the voice intonation feature extraction module, the user's voice intonation features are extracted from the voice data based on the Transformer and BiLSTM hybrid model; through the micro-expression feature recognition module, the user's micro-expression features are extracted from the video data based on the VisionTransformer and the spatiotemporal convolutional network; through the text emotion semantic analysis module, the user's text emotion semantic features are extracted from the text data based on BERT and the attention mechanism; through the physiological change feature recognition module, the user's physiological change features are extracted from the biological signal data using the LSTM model.
[0057] The voice emotion feature extraction module first converts the speech data into a Mel spectrogram through a Mel filter bank to simulate the human ear's perception characteristics of sounds at different frequencies. Then, it uses the convolutional layers of a CNN (such as ResNet) to capture the local acoustic patterns in the Mel spectrogram, and combines an attention mechanism to focus on key frequency bands (for example, sad emotions often correspond to enhanced low-frequency energy). The specific process is as follows: speech framing → short-time Fourier transform → Mel spectrogram calculation → CNN feature extraction → attention weight assignment → emotion feature vector output. When the user says "I'm fine", after the module converts the speech into a Mel spectrogram, the CNN detects energy fluctuations in the low-frequency region (80 - 300 Hz). The attention mechanism assigns a high weight to this region, and combines with a pre-trained model to identify the anxiety features corresponding to voice tremors, and outputs a feature vector containing the probability of "nervous" emotion being 0.7. It accurately extracts the emotion-related acoustic features from the speech signal, overcomes the ambiguity between semantics and emotions (for example, the text "fine" may correspond to multiple emotions), improves the accuracy of emotion recognition, and provides emotional evidence in the acoustic dimension for multimodal fusion.
[0058] The speech intonation feature extraction module uses a hybrid model of Transformer and BiLSTM to process speech prosody features. BiLSTM captures the long-term dependencies of temporal features such as fundamental frequency, speech rate, and pauses (such as the intonation rising pattern at the end of an interrogative sentence). The self-attention mechanism of Transformer models the global associations at each time step (such as the intonation fluctuation rhythm of the whole sentence). The module first extracts the prosody parameters of the speech (fundamental frequency curve, energy entropy, etc.), and then encodes them into an intonation feature vector through the hybrid model. For example, when the user asks "Will it rain today?", the module extracts the frequency rising trend of the last syllable in the fundamental frequency curve (rising from 200 Hz to 250 Hz). BiLSTM records the temporal changes of this rising pattern, and Transformer pays attention to the association between the intonation fluctuation of the whole sentence and the position of the interrogative word "吗", and outputs a vector containing the "question" intonation feature, representing the intensity and rhythm of the rising intonation. Thus, it effectively captures the prosody emotion clues in the speech, distinguishes the semantic differences of different intonations for the same text (such as the intonation difference between a declarative sentence and an interrogative sentence), provides support at the prosody level for dialogue intention recognition, and enhances the temporal expression ability of multimodal features.
[0059] The micro-expression feature recognition module processes video frame sequences based on the Vision Transformer (ViT) and a spatiotemporal convolutional network. ViT segments facial images into blocks and uses a self-attention mechanism to capture global facial movements (such as the overall characteristic of the mouth corners drooping). A spatiotemporal convolutional network analyzes muscle movement trajectories in consecutive frames (such as the temporal sequence of eyelid closure). The module first detects facial landmarks using an MTCNN and then inputs the keypoint coordinate sequence into the model to extract the spatiotemporal features of micro-expressions. For example, if a user displays a "drooping eyelid" micro-expression while speaking, the module locates eye landmarks (such as the corners of the eyes and the eyelid contour) in the video frames. ViT captures the global shape changes of the eye region, and the spatiotemporal convolutional network analyzes the eyelid descent trajectory over three consecutive frames (downward eyelid coordinates in frames 10-12 at 25 frames per second) to output a feature vector containing a probability of 0.8 of a "sad" micro-expression. This module extracts transient micro-expression features from visual modalities to identify implicit emotions that are difficult to express in text and speech (such as socially disguised smiles). This complements the spatiotemporal dynamics of facial movements, improving the ability of multimodal sentiment analysis to capture subtle expressions.
[0060] Furthermore, algorithms such as Dlib or MediaPipe detect 68 key facial points (such as the corners of the eyes and the tip of the nose). The coordinate change sequence is input into an LSTM or GRU to extract the temporal motion features of these key points. For example, a micro-expression of sadness may be accompanied by the continuous movement of the corners of the eyes. Furthermore, texture analysis is performed on local areas (such as the eye area and forehead), using 3D convolution to capture the dynamic changes in skin wrinkles. For example, a micro-expression of anger may manifest as the rapid contraction and relaxation of forehead wrinkles.
[0061] The text sentiment semantics analysis module utilizes the BERT pre-trained model and an attention mechanism to analyze the sentiment of text. BERT uses multiple layers of Transformers to encode the contextual semantics of the text (e.g., the varying emotional intensity of the word "disappointment" in different contexts). The attention mechanism focuses on sentiment keywords (e.g., "sad" and "awful"), and combines this with syntactic analysis to construct semantic nodes. The module first tokenizes the text and feeds it into BERT, which then uses a pooling layer to extract sentiment semantic feature vectors. For example, for the text "Today's event was canceled. I'm so disappointed," BERT's encoding process focuses on the contextual association between "cancelled" and "disappointment." The attention mechanism assigns a high weight to the word "disappointment," resulting in a feature vector with a score of 0.9 for the sentiment dimension of "disappointment," including the semantic representation of "event canceled." This extracts deep semantics and sentiment polarity from the text, constructing semantic nodes and sentiment labels for the text. This provides a semantic baseline for the text dimension for multimodal fusion and ensures that features from each modality are aligned with the text's semantic timeline.
[0062] The physiological change feature recognition module uses an LSTM model to process the temporal features of biosignal data (such as heart rate and electrodermal signals). The LSTM memory unit captures long-term trends in physiological signals (such as a sustained increase in heart rate due to stress), filtering out noise through a gating mechanism while retaining key physiological responses. The module first performs denoising preprocessing on the biosignal before inputting it into the LSTM to extract a temporal feature vector. For example, if a user experiences a sudden increase in heart rate (from 70 bpm to 90 bpm) during a conversation, the LSTM model identifies the temporal correlation between this increase and the conversation context of "work review." Combining historical data, it identifies this as a stress response and outputs a vector containing physiological stress features, representing the amplitude and duration of the heart rate change. This extracts objective emotional stress features from physiological signals, compensating for the concealment of subjective expressions such as speech and facial expressions (such as physiological reactions when feigning calmness), providing objective physiological evidence for multimodal emotion analysis, and enhancing the reliability of emotion recognition.
[0063] As an optional embodiment, it is assumed that the dynamic knowledge graph includes at least an emotional expression view. Based on the above assumption, in step S103, the received compressed data is input into the dynamic knowledge graph enhancement model through the cloud service platform. During the dialogue interaction, the fused feature vector and the entity knowledge and emotional relationships implicit in the historical dialogue are mined in real time to update the dynamic knowledge graph, including:
[0064] The compressed data is decompressed through the decompression module to restore the fused feature vector and the context summary of the historical conversation; the user emotional features of the current round of conversation, the user feature vectors under different modalities, and the confidence of different modalities are extracted from the fused feature vector through the extraction module; the missing information in the context summary is supplemented through the conversation history reconstruction module in combination with the user's emotional expression view stored in the cloud to expand the context summary into complete historical conversation context information; the entity knowledge nodes matching the current round of conversation and the semantic association relationship between each entity knowledge node are constructed through the knowledge extraction module based on the user feature vectors under different modalities and the historical conversation context information; the newly added entity knowledge nodes are fused into the user's emotional expression view stored in the cloud based on the message passing mechanism of the graph neural network GAT in combination with the user's emotional features and the confidence of different modalities to obtain the conversation subgraph of the current round of conversation.
[0065] The conversation subgraph for the current conversation turn includes at least: all entity knowledge nodes corresponding to the current conversation turn and the semantic associations between all entity knowledge nodes. The conversation subgraph is a local substructure of the knowledge graph generated by the current conversation turn, containing all entity nodes and associated edges in that turn. It serves as the incremental update unit of the dynamic knowledge graph.
[0066] It's worth noting that the Dynamic Knowledge Graph Enhancement (DKGE) model is a core component of the cloud service platform. Essentially, it's a dynamic knowledge modeling system based on graph neural networks. Based on the emotional expression view, this model continuously mines entity knowledge and emotional relationships within conversations by analyzing fused feature vectors and historical conversation data uploaded by edge terminals in real time, enabling the dynamic evolution of the knowledge graph. Its core function is to transform the multimodal emotional signals and semantic information generated during user interactions into a structured knowledge network. This allows the knowledge graph to continuously accumulate personalized entity associations and emotional dependencies as the conversation progresses, providing precise knowledge support for subsequent response generation and strategic decision-making, addressing the problem that traditional static knowledge graphs are unable to adapt to real-time conversation scenarios.
[0067] The decompression module uses a semantic-based compression-restoration algorithm to reverse process the compressed data uploaded by the edge terminal. By parsing the keyword index and feature mask retained during compression, the dimensions of the fused feature vector and the contextual summary of the historical conversation are restored to their original format. For example, a 128-dimensional fused vector compressed through feature selection is restored to a 512-dimensional feature vector containing speech, vision, and physiological modalities, and the keyword-compressed conversation summary is expanded into a text sequence including a timestamp.
[0068] Compressed data is lightweight data generated by semantically compressing multimodal features and historical conversations at the edge terminal. It contains the key dimensions of the fused feature vector and the conversation summary, reducing network transmission load. Furthermore, a dynamic compression method based on reinforcement learning dynamically adjusts the compression strategy through interactive learning between the agent and the data environment. The agent can determine the compression algorithm and compression ratio in real time based on the characteristics of the multimodal data (such as data importance and transmission bandwidth). For example, when low network bandwidth is detected, the agent increases the compression ratio for voice data while using reinforcement learning to learn how to preserve the emotional characteristics of the voice as much as possible at a high compression ratio. For text data containing important semantic information, the compression ratio is appropriately reduced. Through trial and error and reward feedback, the agent can find the optimal compression strategy to minimize network transmission load while ensuring data validity.
[0069] This compression method, based on a variational autoencoder (VAE), utilizes the VAE's encoding-decoding structure to compress and restore data. The encoder maps the original multimodal data to a low-dimensional latent space. By constraining the distribution of the latent space, the compressed data retains the key information of the original data while also exhibiting a certain degree of generalization. The decoder reconstructs the original data from the low-dimensional representation of the latent space. Taking historical conversation data as an example, VAE can compress lengthy conversation text into a short vector containing core semantics and emotional tendencies, which can then be restored when needed. This compression method significantly reduces the data volume while preserving the semantic integrity of the data and is also robust to noisy data.
[0070] The context summary is a compressed version of the key information in the conversation history, retaining entity keywords and semantic fragments, but requiring historical reconstruction to complete the complete logic.
[0071] For example, when receiving the compressed data "[user features: compressed vector {speech 0.7, visual 0.5}, conversation summary: 'bad weather → mood']" uploaded by the edge terminal, the decompression module will restore the speech feature dimensions to 20-dimensional features of Mel-Frequency Cepstral Coefficients (MFCCs) and the visual features to 68-dimensional coordinates of facial key points based on the predefined compression dictionary, and expand the conversation summary to the complete text "The user said 'The weather is bad today and my mood is not good'".
[0072] In this way, lossless or nearly lossless restoration of compressed data can be achieved, ensuring that the fusion features and historical conversation information obtained in the cloud have complete semantic details, providing an accurate data basis for subsequent knowledge extraction, and avoiding the loss of features caused by compression that affects the accuracy of knowledge graph updates.
[0073] The extraction module uses a multi-dimensional feature parsing algorithm to separate user emotional features and modality-specific features from the fused feature vector. A pre-trained sentiment classifier (such as a BERT-based sentiment analysis model) is used to extract the current round's emotional polarity (e.g., "sad," "joyful") and intensity from the fused vector. A modal mask matrix is used to separate feature vectors of different modalities, such as speech intonation and micro-expressions. Confidence is calculated based on metrics such as the variance and entropy of each modal feature (e.g., the confidence level of speech features decreases in noisy environments).
[0074] For example, for the "speech features [flat tone 0.3, slow speech speed 0.8], visual features [drooping mouth corners 0.6, drooping eyelids 0.9]" in the fused feature vector, the extraction module will use the emotion classifier to determine that the overall emotion is "depressed" and calculate the speech modal confidence of 0.6 (due to the influence of environmental noise) and the visual modal confidence of 0.9 (clear image).
[0075] In this way, the accurate deconstruction of multimodal features can be achieved, the core emotional features and modal details can be separated, and the emotional dimension annotation information can be provided for the knowledge graph. At the same time, the confidence assessment provides a weight basis for subsequent modal fusion, thereby improving the reliability of knowledge extraction.
[0076] The conversation history reconstruction module uses a temporal association algorithm to match the current context summary with the user's emotional expression view stored in the cloud. By identifying entity keywords in the summary (such as "weather" and "mood"), it retrieves relevant emotional expression view fragments from the historical conversation and supplements the missing information in the summary, such as timestamps and the semantics of the preceding conversation. For example, if the current summary is "user mentioned the weather → feeling depressed," the module will search the historical view for records where "user discussed the weather → feeling happy three rounds ago" to complete the context logic.
[0077] For example, if the context summary of the current compressed data is "The user said 'I don't want to go out'", the reconstruction module will retrieve the previous two rounds of dialogue "The user mentioned 'raining' → low voice tone → increased heart rate" from the emotional expression view stored in the cloud, and complete the missing causal relationship between the "raining" entity and "don't want to go out", forming a complete historical context "raining → don't want to go out → low mood".
[0078] In this way, the problem of missing contextual information caused by compressed data is solved. By reusing the knowledge of historical sentiment expression views, a complete dialogue logic chain is constructed, so that knowledge extraction can be based on the full context, avoiding misjudgment of entity relationships due to information fragmentation.
[0079] The knowledge extraction module uses a cross-modal semantic association algorithm to identify entities and relationships from multimodal features and historical conversations. A named entity recognition (NER) model extracts entities such as "weather" and "mood" from text. Visual entities are supplemented with object recognition from visual features (e.g., images of umbrellas captured by cameras). Co-occurrence analysis and temporal associations are used to construct relationships between entities (e.g., "rain → impact → mood"), and emotional signals from multimodal features (e.g., low voice intonation) are used as sentiment weights between entities.
[0080] For example, in the user conversation "It's raining today, and I'm in a bad mood", the knowledge extraction module extracts the entities "rain" and "mood" from the text, identifies the entity "umbrella" corresponding to the umbrella image from the visual features, and constructs the semantic relationship "rain → causes → bad mood" through temporal association. The low tone (0.8 intensity) in the speech feature is used as the emotional weight of the relationship, forming an entity association with emotional annotations.
[0081] In this way, the transformation from multimodal data to structured knowledge is achieved, and the semantic information and emotional signals in user interactions are abstracted into nodes and edges in the knowledge graph, so that the knowledge graph can not only express the semantic association of entities, but also store dynamic features such as emotional intensity, thereby improving the semantic richness and emotional expression ability of the knowledge graph.
[0082] The fusion module integrates newly added entity knowledge nodes into the existing sentiment expression view based on the message passing mechanism of the Graph Attention Network (GAT). The message passing mechanism is the process of feature propagation between nodes in the GAT, allowing nodes to aggregate information from neighboring nodes based on attention weights, achieving semantic fusion within the knowledge graph. Entity knowledge nodes represent specific concepts or objects in the knowledge graph and contain semantic and sentimental attributes (e.g., the "rain" node has the attribute "wet" associated with the sentiment "affecting mood"). The sentiment expression view is a timeline-based multimodal sentiment association graph that stores the spatiotemporal alignment between sentiment signals in conversations and semantic nodes. The attention mechanism calculates the semantic similarity between newly added nodes and existing nodes (e.g., the correlation between "rain" and the historical node "weather"). The message passing mechanism then propagates features from existing nodes (e.g., the sentiment in historical weather conversations) to the new node. Edge weights are then adjusted based on user sentiment and modality confidence. For example, when adding a new "bad mood" node, GAT will calculate its association strength with the "raining" node, and combine the "frown" feature with high confidence in the visual modality to enhance the weight of the "raining → bad mood" edge.
[0083] For example, when the entity "umbrella" and the relationship "raining → use → umbrella" are newly added, the fusion module calculates the semantic similarity (0.7) between "umbrella" and the historical node "raining" through GAT, and assigns the historical emotional features stored in the "raining" node (such as the user's low mood twice in the past three rainy conversations) to the "umbrella" node through message transmission. At the same time, based on the recognition confidence of the umbrella in the visual modality of 0.9, the weight of the relationship edge is set to 0.7×0.9=0.63, forming an enhanced association with emotional transmission.
[0084] Therefore, the dynamic attention mechanism of GAT is used to realize the incremental update of the knowledge graph, so that the new knowledge can generate semantic interaction with historical knowledge. At the same time, through the weighting of emotional features and modal confidence, it is ensured that the emotional relationship in the knowledge graph can accurately reflect the emotional intensity and modal reliability in the user's real-time interaction, thereby improving the dynamic evolution capability of the knowledge graph.
[0085] Further optionally, in the above steps, the conversation history reconstruction module is used to supplement the missing information in the context summary in combination with the user's emotional expression view stored in the cloud, so as to expand the context summary into complete historical conversation context information, including:
[0086] Using the context summary as an index, the corresponding emotional semantic node is retrieved from the user's emotional expression view; the retrieved emotional semantic node is used to query the corresponding historical user feature vectors under other modes; based on the queried historical user feature vectors, the corresponding historical conversation information is predicted according to the emotional development trend, and the predicted historical conversation information is merged and constructed into the historical conversation context information according to the order of the emotional semantic nodes in the timeline.
[0087] Further optionally, in the above steps, the knowledge extraction module constructs entity knowledge nodes that match the current round of conversation and semantic associations between entity knowledge nodes based on user feature vectors in different modalities and historical conversation context information, including:
[0088] Key named entities that match the current round of conversation and the conversation topic are extracted from user feature vectors under different modalities, and the semantic association relationships between key named entities are extracted by combining historical conversation context information; based on the extraction results, entity knowledge nodes corresponding to key named entities and semantic association relationships between each entity knowledge node are established.
[0089] For example, the knowledge extraction module uses cross-modal collaborative perception and semantic deep mining technology to break the limitations of traditional single-modal analysis. Its core principle is to build a joint representation model based on a multimodal Transformer. This model uses a multi-head attention mechanism to simultaneously process the acoustic features of speech, the semantic vectors of text, the visual features of video, and the fluctuation data of physiological signals. For example, when a user excitedly says in a video call: "I just got tickets to someone's concert!", the system not only recognizes named entities such as "someone", "concert", and "tickets" from the text, but also detects the user's excited characteristics of rising tone and accelerated speaking speed through the voice module, uses the camera to capture the user's facial expressions of raised mouth corners and shining eyes, and combines it with the heart rate acceleration data monitored by wearable devices to verify and enhance the accuracy of entity extraction from multiple dimensions.
[0090] During the relationship extraction phase, the module innovatively introduces a temporal graph convolutional network (TGCN), modeling historical conversation data as a temporal knowledge graph. By analyzing the co-occurrence patterns and semantic evolution patterns of entities in the temporal dimension, it explores the deep semantic associations between entities. For example, if a user has previously expressed their love for someone's music many times, the system can infer a stable relationship of "user-love-someone" and further derive a potential semantic connection of "user-expectation-someone's concert." In addition, the module also uses an adversarial training mechanism, using a generative adversarial network (GAN) to enhance the model's robustness to noisy data and ambiguous expressions, enabling the model to accurately identify referential expressions such as "a certain Dong's performance" and "that awesome live."
[0091] This multimodal collaborative knowledge extraction approach has improved the accuracy and recall of entity and relationship extraction in practical applications. Compared to traditional methods, this module improves the accuracy of entity recognition and relationship extraction in complex conversational scenarios. It can accurately capture implicit information and potential connections in user expressions, providing a rich and accurate knowledge base for the construction of dynamic knowledge graphs.
[0092] Furthermore, through the fusion module, combined with user emotional characteristics and the confidence of different modalities, based on the message passing mechanism of the graph neural network, the newly added entity knowledge nodes are integrated into the user's emotional expression view stored in the cloud, including:
[0093] Based on user emotional characteristics and the confidence of different modalities, corresponding emotional labels are added to the semantic association relationships between newly added entity knowledge nodes; among them, emotional labels are used to indicate the type of emotional state corresponding to the connected entity knowledge nodes; according to the propagation trend of the emotional trend expansion subgraph, an emotional reasoning chain including newly added entity knowledge nodes and semantic association relationships is formed; a dynamic pruning mechanism is used to delete emotional reasoning chains with an overall confidence lower than a set threshold from the emotional expression view; the node status of historical conversations in the emotional expression view is inherited across conversation rounds and transferred to the newly added entity knowledge nodes and semantic association relationships.
[0094] For example, the fusion module, based on a dynamic graph neural network and sentiment inference mechanism, enables real-time updates and sentiment evolution of the knowledge graph. This works by leveraging the message passing mechanism of the Graph Attention Network (GAT), combining the user's sentiment feature vectors and the confidence scores of each modal data, to sentimentally annotate and dynamically fuse newly added entity knowledge nodes and semantic associations. For example, when a user says with a hint of regret, "It's a pity the concert is on a weekday; I might not be able to go," the system first adds the sentiment label "regret" to relationships such as "concert-time-weekday" and "user-attitude-regret" based on the depressed tone of the voice (confidence level 0.85), the negative vocabulary in the text, and the frowning gesture in the micro-expressions (confidence level 0.8).
[0095] Next, the module uses the Spreading Activation Algorithm (SAA) to construct and expand sentiment reasoning chains within the knowledge graph, based on sentiment tags and semantic associations. Starting with "concert - time - weekday," and combining historical conversations with users' complaints about busy work schedules, it infers the complete logical chain of "weekday - leads to - busy work - resulting in the user being unable to attend the concert." To ensure the quality of the knowledge graph, the module also incorporates a dynamic pruning mechanism driven by reinforcement learning. This mechanism calculates the overall confidence level of the reasoning chain based on the product of the confidence scores of each node and edge, automatically removing weakly associated paths below a threshold (e.g., 0.6) to prevent the accumulation of invalid information.
[0096] The module also incorporates a cross-conversation memory transfer mechanism. Using attention-weighted methods, it transfers emotional states and knowledge related to the current topic from past conversations to newly added nodes. For example, if a user previously expressed regret over a time conflict while discussing another activity, the system will transfer this emotional tendency and response strategy to the knowledge graph of the current conversation, enhancing the emotional coherence and reasoning capabilities of the knowledge graph.
[0097] In practical applications, this fusion module enables dynamic knowledge graphs to rapidly adapt to changes in conversational scenarios, updating knowledge structures and emotional expressions in real time. Dialogue systems employing this module improve the accuracy of emotional understanding and enhance performance in complex semantic reasoning tasks. This effectively achieves dynamic knowledge accumulation and accurate emotional expression, enhancing the naturalness and intelligence of conversational interactions.
[0098] As an optional example, assume that the dynamic knowledge graph also includes: expert knowledge bases and domain knowledge graphs associated with each entity knowledge node in the emotion expression view. Based on this, knowledge enhancement processing is performed on the updated dynamic knowledge graph to construct an initial response prototype for the user's current interaction, including:
[0099] Extract the user's current interaction intention information in the current conversation round from the conversation subgraph of the current conversation round; extract the user's historical interaction intention information in previous historical conversation rounds from the conversation subgraph corresponding to the historical conversation in the emotion expression view; perform intent parsing on the current interaction intention information and the historical interaction intention information to obtain a semantic framework for representing the user's interaction intention; use each entity knowledge node in the conversation subgraph of the current conversation round as an index, retrieve corresponding entity knowledge information from the expert knowledge base and / or various domain knowledge graphs, and generate a knowledge filling unit based on the retrieved entity knowledge information; the knowledge filling unit includes at least: a conversation content template constructed based on the entity knowledge information and a slot to be filled; the slot to be filled is supplemented in the user's edge terminal to improve cloud-edge transmission efficiency; adopt an emotion-knowledge mapping mechanism, use the emotion label between each entity knowledge node in the conversation subgraph of the current conversation round as an index, and query the emotion association matrix corresponding to the current conversation round for the emotion change type that matches the emotion development trend as the reply emotion guide for the current conversation round; construct the initial reply prototype based on the semantic framework, knowledge filling unit, and reply emotion guide determined in the above steps.
[0100] First, in the above steps, a deep understanding of user intent is achieved through cross-turn conversation intent correlation analysis and dynamic semantic modeling. This is achieved by using a Transformer-based multi-turn conversation encoder to jointly model the entity nodes in the current conversation subgraph and the intent sequences of historical conversation subgraphs, and using an attention mechanism to capture clues about intent evolution.
[0101] The conversation subgraph for the current conversation turn is essentially a local instantiation of a dynamic knowledge graph in a specific conversational scenario. Focusing on the current user input, it extracts key entities and semantic relationships from multimodal features (text, voice, emoticons, etc.) to form a miniature graph structure focused on the current interaction topic. The construction logic is similar to a "close-up shot," cropping out portions of the global knowledge graph that are strongly relevant to the current conversation. Through a triple structure consisting of entity knowledge nodes (such as named entities), semantic association edges (such as "belongs to" and "associated"), and sentiment labels (such as "positive" and "questionable"), the user's immediate intent is converted into a machine-understandable graph data model. This subgraph does not exist independently, but rather forms a "graph-to-graph" relationship with historical conversation subgraphs and the global sentiment expression view. Through a cross-turn node state inheritance mechanism, semantic coherence of the conversation context is achieved.
[0102] For example, in a medical consultation scenario, the user says in the current round: "I feel uncomfortable in my stomach after taking aspirin" (the current intent is to consult about drug side effects). The system extracts "I asked about the contraindications of ibuprofen last week" (the historical intent is drug selection) from the historical conversation subgraph. Through comparative analysis, it is found that the user has a persistent interest in the topic of "side effects of non-steroidal anti-inflammatory drugs." Technically, the dynamic time warping (DTW) algorithm is used to align the semantic spaces of intents in different rounds, so that slots such as "allergy history" and "dosage" in historical intents can be automatically associated with the current intent, forming a cross-round intent chain. Compared with traditional single-round intent recognition, this method can improve the accuracy of intent understanding in multi-round conversations. It is particularly suitable for scenarios that require long-term tracking (such as chronic disease management) to avoid response bias caused by ignoring historical information.
[0103] The intent parsing module uses frame semantics and dynamic slot generation technology to convert free-text intents into structured semantic frames. This works by identifying core predicates in the intent (e.g., "query" and "side effect") using a pretrained frame semantics model (such as BERT-Framenet) and generating corresponding argument slots (e.g., "drug name" and "symptoms") based on the domain knowledge graph. For example, in the education domain, if a user asks, "Recommend physics introductory books for junior high school students," the system parses the "resource recommendation" frame, which includes slots such as "education level = junior high school students," "subject = physics," and "resource type = introductory books." It also extracts the "previously recommended math books" frame from past conversations and automatically adds the implicit slot "user preference = illustrated text." This process incorporates a dynamic slot pruning mechanism, which calculates the information gain of slots relative to the intent and removes non-critical slots such as "publication date," thereby reducing the frame dimensionality. The technical effect is a standardized semantic representation that provides clear, structured guidance for subsequent knowledge retrieval, improving knowledge matching efficiency in intelligent customer service scenarios.
[0104] Furthermore, a joint graph-text search is performed between the dynamic knowledge graph and the expert knowledge base. This involves using the embedding vectors of entity knowledge nodes (such as the semantic representations generated by the TransE model) as query keys, performing an approximate nearest neighbor search within the expert knowledge base, and inferring and expanding upon the relationship paths within the domain knowledge graph. For example, in a legal consultation, if the user's current conversation subgraph contains the entity nodes "Contract Breach" and "Compensation Amount," the system not only retrieves the text of Article 577 of the Civil Code but also, through the relationship chain "Contract Breach → Compensation Calculation → Burden of Proof" in the domain graph, generates a template containing the "Evidence Type" and "Calculation Method" slots to be filled. The knowledge filling unit adopts a hierarchical structure of "core content + variable slots," such as: "According to Article 577 of the Civil Code, the breaching party is required to bear the responsibility of continuing to perform, taking remedial measures, or compensating for losses (fixed content). The calculation of the compensation amount requires consideration of the __loss type__ and __evidence materials__ (unfilled slots)." This setting allows the cloud to only transmit the template framework (about 1KB) instead of the complete text (about 10KB). The edge terminal fills the slots based on local user data, reducing the cloud-edge transmission volume while ensuring the accuracy of knowledge and scenario adaptability.
[0105] Through graph propagation and sentiment trend prediction within the emotion expression view, knowledge and emotion are jointly modeled. This approach uses emotion labels (such as "anxiety" and "doubt") in the conversation subgraph as query signals for the graph neural network, performing diffusion activation within the sentiment correlation matrix and calculating the transition probability of each emotion type. For example, a user in a financial investment consultation stated, "The stock market has recently plummeted, and I've lost 20% of my principal" (emotional label: "anxiety"). Using the GAT's message passing mechanism, the system discovered that the emotion "anxiety" in historical conversations often accompanies the knowledge need for "risk tolerance assessment." Furthermore, the sentiment weight of the "stop-loss strategy" entity node in the current conversation subgraph is 0.8, thus determining the response's sentiment orientation as "comfort + rational advice." The system innovatively introduces an emotion entropy reduction optimization objective to prioritize response styles that reduce user emotional uncertainty. For example, when the user's emotion entropy exceeds a threshold, the system automatically increases the proportion of soothing statements. In a psychological counseling scenario experiment, this mechanism improved the emotional relevance of responses and increased user engagement in subsequent interactions.
[0106] The reply prototype construction module adopts a generation framework of multi-source information fusion to structurally integrate the semantic framework, knowledge filling unit and emotional orientation. The principle is to align the slots of the semantic framework with the to-be-filled items of the knowledge filling unit through template instantiation and content reordering algorithm (such as the "drug name" slot matches the "aspirin" entry in the knowledge base), and then adjust the expression of the content according to the emotional orientation (such as converting "side effects include stomach pain" to "some users may experience mild stomach discomfort, it is recommended to take with meals"). Taking smart health management as an example, the user's current intention is "hypertension medication consultation", and the historical intention includes "exercise advice". The prototype structure constructed by the system is:
[0107] Emotionally oriented: "Your concern about blood pressure control is very important, and we will provide you with detailed answers" (soothing tone)
[0108] Knowledge filling layer: "The mechanism of action of commonly used antihypertensive drugs such as __drug name__ is __action mechanism__. During medication, you need to pay attention to monitor __monitoring indicators__" (slots are retrieved from the knowledge base)
[0109] Historical association layer: "Based on your previous consultation on aerobic exercise, it is recommended that you avoid strenuous exercise for __ time__ after taking the medicine" (cross-round intention association)
[0110] This prototype generation process incorporates an adversarial rewriting mechanism, using a discriminator to ensure that the generated content meets both knowledge accuracy (such as the correctness of the drug's mechanism of action) and emotional fluency (such as consistency in tone). The generated response prototypes here combine both intellectual depth and emotional warmth. In medical Q&A scenarios, compared to traditional template generation methods, this results in higher user satisfaction with responses and improved information retention.
[0111] As an optional embodiment, after the cloud constructs the initial response prototype through the dynamic knowledge graph enhancement model (DKGE), it needs to be distributed differently based on the capability differences of the edge terminals.
[0112] Specifically, through Quantify the terminal computing power, memory and emotional rendering capabilities, for example, high-performance terminals (such as smartphones, ) can accept complete prototypes (including 3D animation); low-computing devices (such as smart watches, ) only receives text summaries. The GFLOPS is a metric that represents the computing power of the edge node's central processing unit (CPU), measured in giga floating-point operations per second (GFLOPS). For example, the CPU computing power of a smartphone is typically 10-100 GFLOPS, while a home gateway may reach over 200 GFLOPS. This measures the node's ability to handle complex tasks, such as 3D emotional animation rendering, which requires high , while text summary generation requires less computing power.
[0113] Memory capacity (GB). The size of a node's random access memory directly impacts data caching and multitasking capabilities. For example, smartwatches typically have 1-2 GB of memory, while high-performance edge servers can have over 16 GB. This determines the number of tasks a node can handle simultaneously. For example, loading large knowledge subgraphs requires ample memory support.
[0114] Emotional rendering ability (0-10 points) measures the output capability of a node in emotional interaction through quantitative indicators, including: the emotional expressiveness of speech synthesis (such as the range of voice intonation); the emotional adaptability of visual rendering (such as image tone adjustment and animation smoothness); and the synchronization of multimodal fusion (such as the matching degree between speech and expression animation).
[0115] Then, using the matching formula Compute the compatibility of semantic units with terminals to ensure that high sentiment weight units (such as =0.9)) are given priority to those with strong emotional expression ability ( =0.8) devices, avoid low Emotional distortion caused by the terminal.
[0116] in, and Represents the importance ratio of data type matching and emotional adaptation, and The sum is 1. In high emotional demand scenarios (such as psychological counseling), it can be dynamically improved. To 0.6, strengthen the matching priority of emotional rendering capabilities. The data type matching degree. is the data type of the current semantic unit (such as text, link, structured data), For edge nodes The set of supported data types. is the sentiment weight of the semantic unit, indicating the emotional intensity or orientation of the unit. Here, the sentiment dimension can correspond to a specific emotion type (e.g., anxiety, comfort) or intensity (e.g., weak to strong), output by the upstream sentiment analysis module. Renders a preference for the node's sentiment. Indicates the node's tendency in expressing emotions, which is determined by the device's capabilities. Calculate the difference between emotional needs and node capabilities. The smaller the difference, the higher the matching degree. For example, =0.9 strong emotion unit and =0.8, the node difference is 0.1, and the corresponding matching item is 1 - 0.1 = 0.9.
[0117] After constructing the initial response prototype for the user's current round of interaction, the process of sending it to the user's edge terminal faces challenges such as low latency, resource adaptation, and emotional fidelity.
[0118] Here, a layered distribution architecture, centered on semantics and edge node adaptation, aims to address device heterogeneity. The principle is to break down the initial response prototype into independent units based on semantic logic and distribute them differently based on data sensitivity and device capabilities. Through dynamic semantic segmentation, the prototype is divided into modules such as "Problem Attribution" and "Solution," annotated with data types and sentiment weights. For example, in a medical consultation scenario, "Drug Efficacy Description" and "Psychological Support Link" are separated into different units: the former is labeled as text with a sentiment weight of 0.7, while the latter is labeled as a link with a sentiment weight of 0.9. Furthermore, based on a security classification strategy, medical records containing user privacy are only distributed to trusted edge nodes, while general recommendations are distributed to all devices. During the dynamic matching phase at the edge node, each device maintains a capability profile containing information such as computing power and memory. The cloud then allocates tasks accordingly: high-performance devices receive the complete prototype for 3D emotional animation rendering, while low-power devices only receive the text summary and sentiment symbols. If device resources are insufficient, neighboring nodes can collaborate to complete the rendering. For example, when a user uses a smartwatch to ask a health question, the cloud distributes complex text and image responses to the home gateway, which then renders and pushes concise text to the watch, while the full content is displayed on the phone. This architecture improves resource utilization on edge devices and effectively avoids rendering failures or resource waste caused by varying terminal performance.
[0119] Furthermore, incremental streaming based on emotion weighting focuses on ensuring the timeliness and accuracy of emotional responses. This approach prioritizes transmitted data based on emotion tags, establishing an emotional priority queue. Emotions (P0 level), such as high anger and sadness, which may trigger crisis intervention, are given the highest transmission priority, interrupting lower-priority tasks. Neutral or low-intensity emotions (P2 level) are used for routine Q&A. For example, if a user expresses suicidal thoughts, P0-level data containing information about psychological hotlines immediately takes over the regular information transmission channel, ensuring delivery within 200ms. Streaming progressive rendering is also employed, with the terminal gradually displaying responses in the order they are received, presenting the core conclusion first, followed by loading supporting data and emotional suggestions. Emotional micro-tags are embedded in text units to guide the terminal in adjusting speech synthesis parameters, such as a gentler speech rate and falling intonation for sadness. For example, when a user complains about a product issue, the terminal first displays a soothing conclusion, "We deeply understand your concerns," followed by a solution, while the voice is delivered in a caring tone. This transmission method significantly shortens response times in crisis scenarios, and progressive rendering reduces the perceived wait time, making data transmission more tailored to user emotional needs.
[0120] Furthermore, context-aware delivery optimization can be further enhanced by dynamically adjusting delivery strategies through real-time perception of user status and network environment. In principle, this approach leverages biometric signals (such as heart rate and galvanic skin response) collected by wearable devices and terminal location information to determine the user's mood and current situation. If an abnormally elevated heart rate is detected, the system identifies anxiety as a sign of anxiety. Knowledge units are compressed and emotionally comforting content is added. If the user is in a sensitive environment, such as a hospital, medical terminology is replaced with euphemisms. Network adaptability dynamically adjusts the multimodal delivery strategy based on bandwidth: A complete prototype with both text and images is provided when bandwidth is high, while only a text summary and emotional symbols are sent when bandwidth is low. Commonly used knowledge is pre-fetched from the local cache. For example, when a user asks for directions while hiking, if the network is poor, the phone will only receive textual directions and pre-cache information about nearby attractions. Upon returning home, when the network recovers, images and descriptions of the attractions are automatically loaded. This optimization reduces bandwidth requirements and ensures that responses meet user emotional needs while being transmitted reliably in complex network scenarios and with fluctuating user status, avoiding emotional distortion caused by network fluctuations.
[0121] To address the risk of data leakage during knowledge transmission, security and privacy enhancements employ differential privacy injection and lightweight edge decryption technologies. This approach obfuscates the original data by adding semantic noise to knowledge units. For example, a precise statement like "85% response rate from targeted drugs" can be rewritten as "Most patients report significant relief." This approach preserves core semantics while protecting data details. Highly sensitive data is encrypted using the national secret algorithm SM4, with the key managed by the terminal's trusted execution environment (TEE). This ensures that data is only decrypted in a secure environment and immediately destroyed after use. For example, diagnostic reports from medical consultations are encrypted before transmission. Upon arrival at a home health terminal, they are decrypted and displayed by the terminal's TEE, preventing data theft or tampering throughout the entire process. These embodiments meet stringent compliance requirements such as GDPR and HIPAA, ensuring the security of knowledge transmission while protecting user privacy and enhancing user trust in the system.
[0122] As an optional embodiment, in the above steps, the fused feature vector and the updated dynamic knowledge graph are input into the dialogue strategy model through the cloud service platform to determine the response strategy and knowledge retrieval direction of the user's current dialogue turn, including:
[0123] The fused feature vector and the conversation subgraph of the current round of dialogue are input into the dialogue strategy model; through the dialogue strategy model, the corresponding emotional state features are respectively extracted from the fused feature vector according to the three dimensions of emotional type, emotional intensity, and emotional mixed components; through the cross-modal cross-validation mechanism, the dominant emotional type of the user in the current round of dialogue is judged based on the emotional state features under the emotional type and emotional mixed components; based on the preset response strategy library of the collaborative mapping of the dominant emotional type and the emotional intensity, the response strategy matching the current round of dialogue is selected; the knowledge type to which each entity knowledge node in the conversation subgraph of the current round of dialogue belongs is identified; based on the proportion of knowledge types occupied by the entity knowledge nodes, the knowledge base type matching the current round of dialogue is determined as the knowledge call direction of the current round of dialogue.
[0124] For example, dialogue strategy models achieve precise interaction strategy decisions through multi-dimensional sentiment analysis and knowledge graph reasoning. These models can be a hybrid of experts (MoE) and an attention mechanism, dynamically routing different types of dialogue requests to specialized expert modules through a gating mechanism. For example, in medical consultations, "symptom analysis" experts handle diagnostic questions, while "psychological support" experts handle emotional comfort. The gating network determines which experts to activate based on input features (such as sentiment intensity and keywords). By incorporating sentiment features into the gating function, for example, activating "empathy experts" is prioritized when anxiety is high. A certain oncology consultation system improves knowledge matching accuracy by assigning knowledge types (such as medications and surgeries) to corresponding knowledge base modules based on the conversation subgraph. Efficient parameter fine-tuning significantly reduces cross-domain adaptation costs by updating only the expert module parameters while preserving the shared layer.
[0125] The principle behind this approach is to construct a three-dimensional emotion deconstruction framework, extracting emotion type (such as anger, joy), emotion intensity (numerical quantification from weak to strong), and mixed emotion components (such as a complex emotion of 70% anxiety and 30% expectation) from the fused feature vector. For example, a user excitedly expresses during an intelligent financial consultation, "My stock price plummeted and I lost 20%. What should I do?" The model analyzes high-frequency fluctuations in voice intonation, negative vocabulary in the text, and frowning micro-expressions to identify the emotion type as "anxiety," with an intensity of 8 (out of 10), and mixed components including "concern for financial security (60%)" and "seeking solutions (40%)." A cross-modal cross-validation mechanism compares the consistency of multimodal signals to confirm that "anxiety" is the dominant emotion. It then uses a pre-set response strategy of "comfort first + professional advice" from the pre-set strategy library, prioritizing a reassuring message like "We understand your concerns. We'll analyze them for you immediately."
[0126] Knowledge retrieval direction determines the path for extracting specific types of knowledge from the knowledge base during a conversational interaction, based on the current conversation content, user needs, and contextual information. Its core approach is to precisely identify the required knowledge by analyzing entities, semantics, and sentiment in the conversation, enabling efficient and accurate knowledge retrieval and application. Incorporating emotion-driven response strategies (for example, prioritizing solution-related knowledge when a user is angry) can help prevent a disconnect between knowledge retrieval and emotional needs.
[0127] To determine the direction of knowledge retrieval, the model categorizes entity knowledge nodes in the conversation subgraph into different types (e.g., financial concepts, operational procedures, and risk cases). For example, in a financial management conversation, nodes such as "stock crash" and "capital loss" all refer to the "investment risk" knowledge type. By calculating the percentage of nodes of this type reaching a certain threshold, the "investment risk response" knowledge base is determined to be retrieved. This decision-making mechanism, which combines both emotion and knowledge, improves the accuracy of conversation strategies compared to traditional single-rule matching. This is particularly effective in handling complex emotions (such as anger implied by anxiety) and invoking specialized domain knowledge, effectively avoiding strategic misjudgments and knowledge mismatches.
[0128] As an optional embodiment, in step S106, the user edge terminal generates real-time interactive response information output to the user based on the initial response prototype, the response strategy, and the knowledge call direction, including:
[0129] Based on the knowledge call direction, a locally stored knowledge subgraph is loaded; wherein, whether the knowledge call direction of the current round of dialogue is consistent with the knowledge call direction of the historical dialogue is monitored according to a preset strategy; if inconsistent, the matching knowledge base is pre-called, and the corresponding knowledge subgraph is obtained by segmenting from the matching knowledge base; according to the knowledge filling unit and the reply emotion orientation in the initial reply prototype, the matching entity knowledge information is loaded from the knowledge subgraph and filled into the to-be-filled slot of the knowledge filling unit; based on the response strategy, the filled knowledge filling unit is converted into a complete sentence, and the rationality of the complete sentence is checked based on the semantic framework in the initial reply prototype to obtain the first real-time interactive reply information; through a lightweight rendering engine, according to the reply emotion orientation in the initial reply prototype, the language style and / or visual style in the first real-time interactive reply information is dynamically adjusted to obtain the second real-time interactive reply information that is finally output, so as to avoid negative targets of the emotional change type contained in the second real-time interactive reply information.
[0130] Specifically, user edge terminals utilize dynamic knowledge loading and strategic content generation technologies to transform cloud-based instructions into natural interactive responses. The core principle is to establish a dynamic monitoring mechanism for knowledge invocation. When it detects that the knowledge invocation direction of the current conversation (e.g., "investment risk") is inconsistent with the historical conversation (e.g., "product returns"), the corresponding knowledge subgraph is pre-segmented from the locally stored knowledge base to avoid delays caused by real-time requests. For example, when a user switches from consulting fund returns to stock risks, the terminal immediately retrieves the "Stock Risk Control" knowledge subgraph, which contains entity knowledge such as "Stop-Loss Strategy" and "Market Volatility Analysis."
[0131] Based on the initial response prototype filling process, content matching is performed based on the response sentiment and knowledge subgraph. For example, to address user anxiety, reassuring professional content such as "diversified investments reduce risk" is selected from the knowledge subgraph and populated into the prototype's "Solution" slot. The response strategy drives sentence conversion, transforming the filled knowledge unit from "diversified investments, risk control" into a complete statement such as "We recommend that you diversify your investment portfolio to effectively reduce the risk associated with a single stock." The logical rationality of the sentence is then checked based on the semantic framework.
[0132] The lightweight rendering engine uses an emotion mapping algorithm to adjust the language and visual style. For example, when detecting user anxiety, text replies are shortened to short sentences and interjected with exclamations ("Don't worry! We have a solution!"). In visual interfaces, the background is also changed to a softer hue. This mechanism uses positive emotional guidance to avoid negative expressions such as "heavy losses" and "irreversible," improving the emotional relevance of replies. In intelligent customer service scenarios, this increases user satisfaction with replies, effectively mitigates negative emotional conflicts during interactions, and achieves the dual goals of emotional resonance and knowledge transfer.
[0133] In this application, multimodal data acquisition breaks through the limitations of single text, comprehensively capturing multimodal interaction information such as user voice and expressions. This allows for a multi-dimensional understanding of user intent and lays a data foundation for subsequent processing. A lightweight Transformer fusion network is then used to fuse multimodal features using cross-modal attention and a dynamic weighting mechanism, addressing data heterogeneity. This reduces edge computing load while improving feature semantic richness, providing high-quality fused features for subsequent steps. By jointly compressing historical conversations and fused features, data transmission volume is reduced, network latency is lowered, and system response speed is improved, while retaining key information to support cloud processing. Next, a dynamic knowledge graph enhancement model mines entity knowledge and emotional relationships in conversations in real time, dynamically updating the knowledge graph to form a knowledge system tailored to user needs and providing precise knowledge support for reply prototype construction. Based on the updated knowledge graph and fused features, natural language generation technology is used to construct initial reply prototypes containing core semantics, improving the logic and knowledge content of replies. A dialogue strategy model combines user status and knowledge graph information to determine response strategies and knowledge retrieval directions, enabling intelligent dialogue process control and improving response relevance and interaction fluency. Finally, the edge terminal integrates the initial prototype, response strategy, etc., and optimizes the generation of natural language responses that are both knowledge-rich and emotionally adaptable, thereby improving the naturalness of interaction and user satisfaction, and enabling the system to respond efficiently and intelligently to user needs.
[0134] After introducing the system of the exemplary embodiment of the present application, next, reference is made to Figure 2A conversational interaction system based on multimodal emotion perception and dynamic enhancement of knowledge graphs according to an exemplary embodiment of the present application is described. The system includes:
[0135] The user edge terminal is used to obtain multimodal data generated by users during the interaction process. It uses a lightweight Transformer fusion network to convert user feature vectors under different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment. The fused feature vector and historical conversations within a preset round are compressed in real time and uploaded to the cloud.
[0136] The cloud service platform is configured to input the received compressed data into the dynamic knowledge graph enhancement model DKGE, mine the fused feature vector and the entity knowledge and sentiment relationships implicit in the historical conversations in real time during the conversation interaction to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial response prototype for the current round of interaction with the user, and send it to the user's edge terminal;
[0137] The cloud service platform is further configured to input the fused feature vector and the updated dynamic knowledge graph into the dialogue strategy model, determine the response strategy and knowledge call direction for the current dialogue round, and send them to the user edge terminal;
[0138] The user edge terminal is further configured to generate real-time interactive response information output to the user based on the initial response prototype, the response strategy, and the knowledge calling direction.
[0139] The above system can implement each step described in the above embodiment, and the specific implementation method of each step will not be repeated here.
[0140] After introducing the system of the exemplary embodiment of the present application, the following describes a terminal device of the exemplary embodiment of the present application, which can be implemented as a user edge terminal. The user edge terminal is used to obtain multimodal data generated by the user during the interaction process; a lightweight Transformer fusion network is used to convert user feature vectors under different modalities into fused feature vectors through cross-modal attention fusion and dynamic weight adjustment; the fused feature vectors and historical conversations within a preset round are compressed in real time and uploaded to the cloud. Thus, the received compressed data is input into the dynamic knowledge graph enhancement model DKGE through the cloud service platform, and the entity knowledge and emotional relationships implicit in the fused feature vectors and the historical conversations are mined in real time during the dialogue interaction to update the dynamic knowledge graph, and knowledge enhancement processing is performed based on the updated dynamic knowledge graph to construct an initial response prototype for the current round of interaction with the user and send it to the user edge terminal; the fused feature vectors and the updated dynamic knowledge graph are input into the dialogue strategy model to determine the response strategy and knowledge call direction of the current round of dialogue, and send it to the user edge terminal. Furthermore, the terminal device is further configured to generate real-time interactive response information output to the user based on the initial response prototype, the response strategy, and the knowledge call direction. The above system can implement each step described in the above embodiment, and the specific implementation method of each step will not be repeated here.
[0141] After introducing the system and terminal device of the exemplary embodiment of the present application, Figure 3 For a description of the computer-readable storage medium of the exemplary embodiment of the present application, please refer to Figure 3 The computer-readable storage medium shown is an optical disc 30, which stores a computer program (i.e., a program product). When the computer program is executed by a processor, it implements each step described in the above embodiment. The specific implementation method of each step is not repeated here.
[0142] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which will not be described in detail here. The above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-mentioned embodiments, ordinary technicians in this field should understand that any technician familiar with this technical field can still modify the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in this application, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement, characterized by: The system includes at least a user edge terminal and a cloud service platform, and the system includes: Obtain multimodal data generated by users during the interaction process through user edge terminals; Through the user edge terminal, a lightweight Transformer fusion network is used to convert user feature vectors under different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment. The fused feature vector and historical conversations within the preset rounds are compressed in real time and uploaded to the cloud. The received compressed data is input into the dynamic knowledge graph enhancement model DKGE through the cloud service platform. During the dialogue interaction, the fused feature vector and the entity knowledge and emotional relationships implicit in the historical dialogue are mined in real time to update the dynamic knowledge graph. Based on the updated dynamic knowledge graph, knowledge enhancement processing is performed to construct the initial response prototype for the current round of interaction with the user and send it to the user's edge terminal. Through the cloud service platform, the fused feature vector and the updated dynamic knowledge graph are input into the dialogue strategy model to determine the response strategy and knowledge call direction for the current round of dialogue, and then sent to the user edge terminal; Through the user edge terminal, real-time interactive reply information is generated and output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction.
2. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 1 is characterized in that: The multimodal data includes at least: voice data, video data, text data, and biosignal data collected by wearable devices; The lightweight Transformer fusion network includes at least: a feature extraction layer, a spatiotemporal alignment layer, a fusion layer, and an adaptive weight distribution layer; The method uses a lightweight Transformer fusion network through the user edge terminal to convert user feature vectors under different modalities into fused feature vectors through cross-modal attention fusion and dynamic weight adjustment, including: In the feature extraction layer, for different modal data, a voice emotion feature extraction module, a speech intonation feature extraction module, a micro-expression feature recognition module, a text emotion semantic analysis module, and a physiological change feature recognition module are respectively constructed; and through the constructed different modal processing modules, user feature vectors under different modalities are extracted from the multimodal data; wherein the user feature vectors include at least: the user's voice emotion features, speech intonation features, micro-expression features, text emotion semantic features, and physiological change features; Through the spatiotemporal alignment layer and the fully connected layer, the user feature vector is projected into a unified dimensional space. The emotional semantic nodes in the text emotional semantic features are used as the time axis reference, and the user feature vectors under other modalities are positioned to each emotional semantic node to construct a spatiotemporally consistent emotional expression view. Through the fusion layer, combined with the emotional expression view, the emotional semantic nodes in the text emotional semantic features are used as the time axis benchmark. Different user feature vectors are used to perform cross-modal retrieval on other user feature vectors according to the time sequence of the time axis benchmark to obtain corresponding matching feature vectors. The fused feature vectors under different branches are then obtained based on the cross-modal retrieval results. Through the adaptive weight allocation layer, the signal quality evaluation and semantic consistency detection are performed on the fused feature vectors under different branches. Based on the evaluation and detection results, the weight parameters corresponding to different modalities are dynamically adjusted to obtain the final output fused feature vector, so as to eliminate data conflicts in the fused feature vector and improve the accuracy of the output results.
3. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 2 is characterized in that: The constructing of different modality processing modules to extract user feature vectors in different modalities from the multimodal data includes: The voice emotion feature extraction module extracts the user's voice emotion features from the speech data based on the Mel spectrum and CNN attention mechanism; The voice and intonation feature extraction module extracts the user's voice and intonation features from the voice data based on the Transformer and BiLSTM hybrid model; Through the micro-expression feature recognition module, the user's micro-expression features are extracted from the video data based on the Vision Transformer and the spatiotemporal convolutional network; Through the text sentiment semantic analysis module, the user's text sentiment semantic features are extracted from the text data based on BERT and attention mechanism; The physiological change feature recognition module uses an LSTM model to extract the user's physiological change features from the biosignal data.
4. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 1 is characterized in that: The dynamic knowledge graph includes at least: sentiment expression view; The received compressed data is input into the dynamic knowledge graph enhancement model DKGE through the cloud service platform, and the entity knowledge and emotional relationships implicit in the fused feature vector and the historical conversations are mined in real time during the conversation interaction to update the dynamic knowledge graph, including: Decompressing the compressed data by a decompression module to restore the fused feature vector and the context summary of the historical conversation; Extracting the user emotion features of the current round of conversation, the user feature vectors under different modalities, and the confidence of different modalities from the fused feature vector through an extraction module; By using a conversation history reconstruction module and combining the user's emotional expression view stored in the cloud, the missing information in the context summary is supplemented to expand the context summary into complete historical conversation context information; The knowledge extraction module builds entity knowledge nodes that match the current round of conversation, as well as semantic associations between entity knowledge nodes, based on user feature vectors in different modalities and historical conversation context information. Through the fusion module, combined with user emotional features and the confidence of different modalities, based on the message passing mechanism of the graph neural network GAT, the newly added entity knowledge nodes are integrated into the user's emotional expression view stored in the cloud to obtain the conversation subgraph of the current round of conversation; The conversation subgraph of the current round of conversation includes at least: all entity knowledge nodes corresponding to the current round of conversation, and the semantic association relationships between all entity knowledge nodes.
5. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 4 is characterized in that: The conversation history reconstruction module, combined with the user's emotional expression view stored in the cloud, supplements the missing information in the context summary to expand the context summary into complete historical conversation context information, including: Retrieving corresponding emotion semantic nodes from the user's emotion expression view using the context summary as an index; Use the retrieved sentiment semantic nodes to query the corresponding historical user feature vectors in other modalities; Based on the queried historical user feature vectors, the corresponding historical conversation information is predicted according to the emotion development trend, and the predicted historical conversation information is merged and constructed into the historical conversation context information according to the order of the emotion semantic nodes in the timeline.
6. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 4 is characterized in that: The knowledge extraction module constructs entity knowledge nodes that match the current round of conversation and the semantic association relationships between each entity knowledge node based on user feature vectors in different modalities and historical conversation context information, including: Extract key named entities that match the current conversation round and conversation topic from user feature vectors in different modalities, and extract semantic associations between key named entities based on historical conversation context information. Based on the extraction results, establish entity knowledge nodes corresponding to key named entities and the semantic association relationships between each entity knowledge node; The fusion module combines user emotional features and the confidence of different modalities, and integrates the newly added entity knowledge nodes into the user's emotional expression view stored in the cloud based on the message passing mechanism of the graph neural network GAT, including: Based on user emotional characteristics and the confidence of different modalities, corresponding emotional labels are added to the semantic association relationship between the newly added entity knowledge nodes; where the emotional label is used to indicate the emotional state type corresponding to the connected entity knowledge nodes; According to the propagation trend of the sentiment trend expansion subgraph, a sentiment reasoning chain including newly added entity knowledge nodes and semantic association relationships is formed; A dynamic pruning mechanism is used to remove emotion reasoning chains whose overall confidence is lower than a set threshold from the emotion expression view; The node states of historical conversations in the emotion expression view are inherited across conversation turns and transferred to the newly added entity knowledge nodes and semantic association relationships.
7. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 4 is characterized in that: The dynamic knowledge graph also includes: expert knowledge bases and knowledge graphs in various fields associated with each entity knowledge node in the emotion expression view; The knowledge enhancement process is performed based on the updated dynamic knowledge graph to construct an initial response prototype for interacting with the user in the current round, including: Extract the user's current interaction intention information in the current conversation round from the conversation subgraph of the current conversation round; Extract the user's historical interaction intention information in previous historical rounds of conversation from the conversation subgraph corresponding to the historical conversation in the emotion expression view; Perform intent analysis on the current interaction intention information and historical interaction intention information to obtain a semantic framework for representing the user's interaction intention; Using each entity knowledge node in the conversation subgraph of the current conversation round as an index, the corresponding entity knowledge information is retrieved from the expert knowledge base and / or various domain knowledge graphs, and a knowledge filling unit is generated based on the retrieved entity knowledge information. The knowledge filling unit includes at least: a conversation content template constructed based on the entity knowledge information and slots to be filled. The slots to be filled are supplemented in the user's edge terminal to improve cloud-edge transmission efficiency. Using the emotion-knowledge mapping mechanism, the emotional labels between the entity knowledge nodes in the conversation subgraph of the current conversation round are used as indexes. In the emotion association matrix corresponding to the current conversation round, the emotion change type that matches the emotion development trend is searched and used as the emotional guidance for the reply in the current conversation round. Based on the semantic framework, knowledge filling unit, and reply emotion orientation determined in the above steps, the initial reply prototype is constructed.
8. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 4 is characterized in that: The fused feature vector and the updated dynamic knowledge graph are input into the dialogue strategy model through the cloud service platform to determine the response strategy and knowledge call direction of the user's current dialogue round, including: Inputting the fused feature vector and the dialogue subgraph of the current dialogue round into the dialogue strategy model; Through the dialogue strategy model, the corresponding emotional state features are extracted from the fused feature vector according to the three dimensions of emotional type, emotional intensity, and emotional mixed components; Through a cross-modal cross-validation mechanism, the user's dominant emotion type in the current round of conversation is determined based on the emotion type and the emotional state characteristics under the emotion mixture component; Based on the preset response strategy library that maps the dominant emotion type and emotion intensity, select the response strategy that matches the current round of dialogue; Identify the knowledge type of each entity knowledge node in the conversation subgraph of the current conversation round; Based on the proportion of knowledge types occupied by entity knowledge nodes, the knowledge base type that matches the current round of dialogue is determined as the knowledge call direction of the current round of dialogue.
9. The conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement according to claim 1 is characterized in that: Generating, through the user edge terminal, real-time interactive reply information output to the user according to the initial reply prototype, the response strategy, and the knowledge call direction, includes: Based on the knowledge call direction, a locally stored knowledge subgraph is loaded; wherein, according to a preset strategy, whether the knowledge call direction of the current round of dialogue is consistent with the knowledge call direction of the historical dialogue is monitored; if not, a matching knowledge base is pre-called, and the corresponding knowledge subgraph is obtained by segmenting from the matching knowledge base; According to the knowledge filling unit and the reply emotion orientation in the initial reply prototype, the matching entity knowledge information is loaded from the knowledge subgraph and filled into the to-be-filled slot of the knowledge filling unit; Converting the filled knowledge filling unit into a complete sentence based on the response strategy, and performing a rationality check on the complete sentence based on the semantic framework in the initial reply prototype to obtain first real-time interactive reply information; Through a lightweight rendering engine, the language style and / or visual style in the first real-time interactive reply information is dynamically adjusted according to the reply emotion orientation in the initial reply prototype to obtain the final output second real-time interactive reply information, so as to avoid negative targets of the emotional change type contained in the second real-time interactive reply information.
10. A conversational interaction system based on multimodal emotion perception and knowledge graph dynamic enhancement, characterized by: The system comprises: The user edge terminal is used to obtain multimodal data generated by users during the interaction process. It uses a lightweight Transformer fusion network to convert user feature vectors under different modalities into a fused feature vector through cross-modal attention fusion and dynamic weight adjustment. The fused feature vector and historical conversations within a preset round are compressed in real time and uploaded to the cloud. The cloud service platform is configured to input the received compressed data into the dynamic knowledge graph enhancement model DKGE, mine the fused feature vector and the entity knowledge and sentiment relationships implicit in the historical conversations in real time during the conversation interaction to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial response prototype for the current round of interaction with the user, and send it to the user's edge terminal; The cloud service platform is further configured to input the fused feature vector and the updated dynamic knowledge graph into the dialogue strategy model, determine the response strategy and knowledge call direction for the current dialogue round, and send them to the user edge terminal; The user edge terminal is further configured to generate real-time interactive response information output to the user based on the initial response prototype, the response strategy, and the knowledge calling direction.
Citation Information
Patent Citations
Intelligent real-time emotion evaluation method for social media and online text data based on multi-modal knowledge graph
CN119202270A
Generating model output using a knowledge graph
US20240428787A1
Cited By
Intelligent customer service automatic reply generation method and system based on multi-modal learning
CN120975248A
Old people accompanying robot system based on multi-mode emotion cognition reinforcement learning
CN121211342A
Intelligent statistical analysis method for news transmission cross-platform data
CN121303114A
Multi-round dialogue logic optimization method in intelligent question-answering system based on knowledge graph
CN121478989A
Method for optimizing multi-turn dialogue logic in intelligent question-answering system based on knowledge graph
CN121478989B