Dialogue interaction system based on multi-modal sentiment perception and knowledge graph dynamic enhancement
By acquiring multimodal data and using dynamic knowledge graph enhancement models, the problems of capturing non-textual information and dynamically associating knowledge graphs in artificial intelligence dialogue systems have been solved, achieving efficient response and fluency in intelligent dialogue interaction.
Patent Information
- Application Number
- CN202511274967.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing AI dialogue systems struggle to capture non-textual information, resulting in stiff interactions and a lack of empathy. Furthermore, traditional knowledge graphs are unable to dynamically connect real-time information and potential knowledge needs within the dialogue context, failing to deeply respond to users' emotional states.
Multimodal data is acquired through user edge terminals, cross-modal attention fusion is performed using a lightweight Transformer fusion network, and the data is uploaded to a cloud service platform for dynamic knowledge graph augmentation model DKGE to mine entity knowledge and emotional relationships. Real-time interactive responses are generated by combining the dialogue strategy model.
It enables intelligent dialogue interaction, improves the relevance of responses and the smoothness of interaction, reduces network latency, and enhances user satisfaction and the naturalness of interaction.
Smart Images

Figure CN120745853B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of artificial intelligence, and more particularly, embodiments of the present application relate to a dialogue interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph. BACKGROUND
[0002] At present, with the development of artificial intelligence technology, artificial intelligence dialogue systems are becoming more and more popular.
[0003] In the related art, chat robots (such as customer service robots, Siri, Alexa, etc.) mainly rely on text semantic analysis, which is difficult to capture non-text information (such as emotions implied in voice intonation, psychological states represented by facial micro-expressions, and stress implied in voice) that is crucial in communication. This leads to stiff interaction, lack of empathy, and easy misunderstanding of the user's true intention and emotional state, especially in complex, emotionally charged dialogue scenarios (such as psychological counseling, complaint handling, and companion care). Moreover, in the related art, traditional knowledge graphs in dialogue are usually static or pre-set references, which are difficult to dynamically associate real-time information in the dialogue context with potential knowledge needs, lack the ability to dynamically mine relevant knowledge and make analogies based on the current emotional state, making the answers relevant but not deep enough or in line with user needs, and lacking resonance in the user interaction process.
[0004] Therefore, there is an urgent need to design a more efficient dialogue interaction scheme to solve at least one of the above technical problems. SUMMARY
[0005] In this context, embodiments of the present application aim to provide a dialogue interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph, which can realize intelligent dialogue interaction and improve the relevance of replies and the fluency of interaction.
[0006] In a first aspect of the embodiments of the present application, a dialogue interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph is provided, which at least includes a user edge terminal and a cloud service platform, the system comprising:
[0007] Through the user edge terminal, multi-modal data generated by the user in the interaction process is obtained;
[0008] Through the user edge terminal, a lightweight Transformer fusion network is used to convert user feature vectors in different modalities into fusion feature vectors through cross-modal attention fusion and dynamic weight adjustment; the fusion feature vectors are real-time compressed and uploaded to the cloud along with the historical dialogue within the preset round;
[0009] The cloud service platform inputs the received compressed data into a dynamic knowledge graph enhancement model DKGE, mines entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, performs knowledge enhancement processing based on the updated dynamic knowledge graph, constructs an initial reply prototype for the current round of interaction with the user, and delivers the initial reply prototype to the user edge terminal;
[0010] The cloud service platform inputs the received compressed data into a dynamic knowledge graph enhancement model DKGE, mines entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, performs knowledge enhancement processing based on the updated dynamic knowledge graph, constructs an initial reply prototype for the current round of interaction with the user, and delivers the initial reply prototype to the user edge terminal;
[0011] The cloud service platform inputs the received compressed data into a dynamic knowledge graph enhancement model DKGE, mines entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, performs knowledge enhancement processing based on the updated dynamic knowledge graph, constructs an initial reply prototype for the current round of interaction with the user, and delivers the initial reply prototype to the user edge terminal.
[0012] In a second aspect of the embodiments of the present application, a dialogue interaction system based on multi-modal sentiment perception and dynamic enhancement of a knowledge graph is provided, comprising:
[0013] The user edge terminal is configured to acquire multi-modal data generated by the user during the interaction process, convert user feature vectors in different modalities into a fusion feature vector through cross-modal attention fusion and dynamic weight adjustment by using a lightweight Transformer fusion network, and upload the fusion feature vector and historical dialogues in a preset round to the cloud in real time.
[0014] The cloud service platform is configured to input the received compressed data into a dynamic knowledge graph enhancement model DKGE, mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial reply prototype for the current round of interaction with the user, and deliver the initial reply prototype to the user edge terminal.
[0015] The cloud service platform is configured to input the received compressed data into a dynamic knowledge graph enhancement model DKGE, mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial reply prototype for the current round of interaction with the user, and deliver the initial reply prototype to the user edge terminal.
[0016] The cloud service platform is configured to input the received compressed data into a dynamic knowledge graph enhancement model DKGE, mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial reply prototype for the current round of interaction with the user, and deliver the initial reply prototype to the user edge terminal.
[0017] In a third aspect of the embodiments of the present application, a terminal device is provided, comprising at least one processor, a memory and an input-output unit; wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program stored in the memory to execute the dialog interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph according to any one of the first aspect.
[0018] In a fourth aspect of the embodiments of the present application, a computer readable storage medium is provided, comprising instructions which, when executed on a computer, cause the computer to execute the dialog interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph according to any one of the first aspect.
[0019] In a fifth aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the dialog interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph according to any one of the first aspect.
[0020] According to the dialog interaction system based on multi-modal sentiment perception and dynamic enhancement of knowledge graph according to the embodiments of the present application, the dialog interaction system is applied to a cloud-edge collaborative dialog interaction system, which comprises at least a user edge terminal and a cloud service platform. In the embodiments of the present application, the user edge terminal is used to acquire multi-modal data generated by the user in the interaction process; a lightweight Transformer fusion network is used to convert the user feature vectors in different modalities into fusion feature vectors through cross-modal attention fusion and dynamic weight adjustment; and the fusion feature vectors and historical dialog in a preset round are compressed in real time and uploaded to the cloud. Further, the cloud service platform is used to input the received compressed data into a dynamic knowledge graph enhancement model DKGE, to mine the entity knowledge and emotional relationship implied in the fusion feature vectors and the historical dialog in real time during the dialog interaction process, to update the dynamic knowledge graph, and to perform knowledge enhancement processing based on the updated dynamic knowledge graph, to construct an initial reply prototype for the current round of interaction with the user, and to deliver the initial reply prototype to the user edge terminal; the fusion feature vectors and the updated dynamic knowledge graph are input into a dialog strategy model to determine the response strategy and knowledge calling direction of the current round of dialog, and to deliver the response strategy and knowledge calling direction to the user edge terminal. Finally, the user edge terminal is used to generate real-time interaction reply information output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction.
[0021] In this implementation, multimodal data acquisition first overcomes the limitations of single text, comprehensively capturing multimodal interaction information such as user voice and facial expressions to understand user intent from multiple dimensions, laying a data foundation for subsequent processing. Then, a lightweight Transformer fusion network is used to fuse multimodal features through cross-modal attention and dynamic weighting mechanisms, solving the data heterogeneity problem. This reduces edge computing load while improving feature semantic richness, providing high-quality fused features for subsequent steps. Joint compression of historical dialogues and fused features reduces data transmission volume, lowers network latency, and improves system response speed, while retaining key information to support cloud processing. Next, a dynamic knowledge graph enhancement model mines entity knowledge and emotional relationships in the dialogue in real time, dynamically updating the knowledge graph to form a knowledge system tailored to user needs, providing precise knowledge support for constructing response prototypes. Based on the updated knowledge graph and fused features, natural language generation technology is used to construct an initial response prototype containing core semantics, improving the logic and knowledge content of the response. The dialogue strategy model combines user state and knowledge graph information to determine response strategies and knowledge invocation directions, achieving intelligent dialogue flow control and improving response relevance and interaction fluency. Finally, the edge terminal integrates the initial prototype, response strategy, etc., and optimizes the generation of natural language responses that are both knowledge-rich and emotionally resonant, thereby improving the naturalness of interaction and user satisfaction, and enabling the system to respond to user needs efficiently and intelligently. Attached Figure Description
[0022] Figure 1 A schematic diagram of the process of a dialogue interaction system based on multimodal emotion perception and dynamic knowledge graph enhancement provided in an embodiment of this application;
[0023] Figure 2 A schematic diagram of the structure of a dialogue interaction system based on multimodal emotion perception and dynamic knowledge graph enhancement provided in an embodiment of this application;
[0024] Figure 3 A schematic diagram of the structure of a medium according to an embodiment of this application is shown. Detailed Implementation
[0025] The following is for reference. Figure 1 , Figure 1 This is a flowchart illustrating a dialogue interaction system based on multimodal emotion perception and dynamic knowledge graph enhancement, provided as an embodiment of this application.
[0026] Figure 1 The flowchart of a dialogue interaction system based on multimodal emotion perception and dynamic knowledge graph enhancement provided in an embodiment of this application, as shown, includes:
[0027] Step S101: Obtain multimodal data generated by the user during the interaction process through the user edge terminal;
[0028] In step S102, the user edge terminal converts the user feature vectors in different modalities into a fused feature vector by using a lightweight Transformer fusion network, through cross-modal attention fusion and dynamic weight adjustment.
[0029] In step S103, the fused feature vector is compressed in real time and uploaded to the cloud together with the historical dialogue in the preset round.
[0030] In step S104, the cloud service platform inputs the received compressed data into a dynamic knowledge graph enhancement model DKGE, mines the entity knowledge and emotional relationship implied in the fused feature vector and the historical dialogue in real time during the dialogue interaction process to update the dynamic knowledge graph, performs knowledge enhancement processing based on the updated dynamic knowledge graph, and constructs an initial reply prototype for the current round of interaction with the user, and delivers it to the user edge terminal.
[0031] In step S105, the cloud service platform inputs the fused feature vector and the updated dynamic knowledge graph into a dialogue strategy model to determine the response strategy and knowledge calling direction of the current round of dialogue, and delivers them to the user edge terminal.
[0032] In step S106, the user edge terminal generates real-time interactive reply information output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction.
[0033] In the embodiments of the present application, a dialogue interaction system applied to cloud-edge collaboration is provided. Further optionally, the dialogue interaction system at least includes a user edge terminal and a cloud service platform.
[0034] Specifically, the user edge terminal is a front-end device directly contacted by the dialogue interaction system and the user, and undertakes the core functions of multi-modal data acquisition and preliminary processing. At the hardware level, it integrates microphones, cameras, various sensors and other devices, and can capture multi-modal data such as voice, text, expressions, actions and other multi-modal data generated by the user during the interaction process in real time, realizing comprehensive perception of the user's interactive information. At the algorithm level, it is equipped with a lightweight Transformer fusion network, which fuses feature vectors of different modalities into unified semantic representations through cross-modal attention mechanism and dynamic weight adjustment strategy, effectively solving the heterogeneity problem of multi-modal data. At the same time, the terminal is also responsible for real-time compression of the historical dialogue and the fused feature in the preset round, and reduces the data volume by removing redundant information, preparing for subsequent transmission to the cloud. In addition, the edge terminal receives the initial reply prototype, response strategy and other information issued by the cloud, combines the local language generation model and the emotion rendering module, and finally generates real-time interactive replies for the user, and completes the feedback through the speaker, screen and other output devices.
[0035] The cloud service platform is the core brain of the dialogue interaction system, mainly responsible for tasks such as knowledge processing, strategy decision-making, and reply framework construction. In the data receiving link, it receives compressed data from user edge terminals, including fused feature vectors and historical dialogue information, providing data support for subsequent knowledge graph updates. The core module, the dynamic knowledge graph enhancement model (DKGE), analyzes these data in real time, mines the implied entity knowledge and emotional relationships through techniques such as graph neural networks, and dynamically updates the nodes and edges of the knowledge graph, allowing the knowledge graph to evolve with the dialogue process and form a knowledge system that meets user needs. Graph neural networks, such as GAT, are neural networks that calculate the weights between nodes in a graph through attention mechanisms, suitable for handling node interactions and feature transmission in knowledge graphs. Based on the updated knowledge graph, the cloud uses natural language generation techniques to construct an initial reply prototype that contains core semantics and reply frameworks that match the current interaction with the user. At the same time, the dialogue strategy model integrates feature vectors and dynamic knowledge graph information to determine the response strategy (such as asking, answering, guiding, etc.) and knowledge invocation direction (such as domain knowledge, user preferences, etc.) for the current round through reinforcement learning and other methods, achieving intelligent control of the dialogue process. Finally, the cloud sends the initial reply prototype and strategy decision results to the edge terminal, providing key inputs for the terminal to generate the final reply.
[0036] Entity knowledge is a structured semantic unit extracted from dialogue, such as specific concepts or objects such as weather and mood, and is the basic node of the knowledge graph. Emotional relationships are associations between entities with emotional polarity, such as the "causes" relationship in "rain → causes → bad mood" with "low" emotional intensity.
[0037] In practical applications, cloud-edge collaboration sinks some data processing tasks to edge terminals, avoiding the direct upload of large amounts of raw data to the cloud. For example, in a dialogue system, the edge terminal first performs feature fusion and compression on multi-modal data, transmitting only key information to the cloud, which can significantly reduce network traffic. This edge preprocessing mode and cloud deep processing mode can reduce data transmission delay and ensure real-time interaction. For example, user voice commands can be quickly converted into feature vectors at the edge, eliminating the need to wait for the cloud to fully analyze the original audio, improving dialogue response speed.
[0038] Specifically, the cloud has strong computing power and storage resources, making it suitable for tasks such as large-scale knowledge graph updates and complex strategy decision-making. The edge terminal, on the other hand, is good at local real-time data collection and lightweight computing (such as multi-modal feature fusion). Taking a dialogue system as an example, the edge uses lightweight Transformer to process multi-modal data, avoiding the occupation of cloud resources by massive raw data, allowing the cloud to focus on deep mining of dynamic knowledge graphs and reply framework construction. This division of labor can maximize the use of cloud-edge resource advantages, reducing overall system energy consumption and computing costs.
[0039] Edge terminals can complete part of data processing locally, maintain basic interaction functions (such as generating simple replies based on local cache knowledge) even if the network is interrupted, and avoid system complete paralysis. At the same time, sensitive data (such as user expressions, voice features) can be desensitized or feature extracted at the edge, reducing the risk of uploading raw data to the cloud. For example, only the compressed feature vector is transmitted to the cloud instead of the complete audio and video, combined with edge local encryption technology, which can effectively enhance the user privacy protection capability.
[0040] The cloud can construct a global knowledge graph and continuously optimize by integrating the data of multiple edge terminals, while the edge terminal can also form personalized knowledge supplements based on local user interaction history. For example, in a dialogue system, the edge terminal records the user's emotional preferences in a specific scenario, uploads it to the cloud, and dynamically updates the knowledge graph, so that the cloud can provide more personalized reply strategies for the user, realizing the two-way cooperation of edge personalized perception and cloud global knowledge enhancement, and improving the personalization and intelligence level of interaction.
[0041] In the embodiments of the present application, first, through multi-modal data acquisition, the limitations of single text mode are broken through, and user interaction information is comprehensively captured, understanding user intent from multiple dimensions such as language expression, emotional state and behavior characteristics, laying an accurate data foundation for subsequent analysis and processing. Further, the lightweight Transformer fusion network effectively solves the problem of multi-modal data heterogeneity, reduces the computational load of the edge terminal, improves the integrity and semantic richness of feature representation, ensures real-time processing and provides high-quality fusion features for subsequent steps. Real-time compression and uploading of historical dialogue reduces data transmission volume, reduces network bandwidth occupation and delay, improves system response speed, while preserving key information, balancing data integrity and transmission efficiency, so that the cloud can work based on effective data. Next, the dynamic knowledge graph enhancement model (DKGE) allows the knowledge graph to evolve in real time with the dialogue, accumulates relevant entity knowledge and emotional associations, and forms a knowledge system that meets user needs, providing accurate and dynamic knowledge support for reply prototype construction, enhancing the system's knowledge expression and reasoning ability. Initial reply prototype construction is based on updated knowledge graph and fusion features, using natural language generation technology to generate reply frameworks with rich knowledge support and semantic accuracy, providing a basis for subsequent optimization and improving the logicality and knowledge of the reply. The dialogue strategy model determines the response strategy and knowledge calling direction according to the user's real-time state and knowledge graph information, realizes intelligent interaction process control, accurately obtains the required knowledge, improves the relevance and effectiveness of the reply, and enhances the fluency of the dialogue and user experience. Finally, real-time interactive reply information generation integrates the achievements of each part, generates a reply containing rich knowledge, meeting the requirements of the strategy and the user's emotional state, and feeds back to the user through natural language expression, improving the naturalness and satisfaction of the interaction, and making the entire system respond efficiently and intelligently to user needs.
[0042] At step S101, the user edge terminal acquires multi-modal data generated by the user in the interaction process. Specifically, the user edge terminal acquires multi-modal data through hardware devices and sensor networks. The microphone captures the voice signal and converts it into audio waveform data, the camera captures visual information such as facial expressions and body movements, the sensor (such as an accelerometer and a gyroscope) records the user's action posture, and the text input device (keyboard and touch screen) acquires the text instruction. After analog-to-digital conversion, the edge terminal performs noise reduction, normalization and other operations through the preprocessing module to form a structured multi-modal data sequence. For example, the speech data will be framed to extract the Mel frequency cepstral coefficient (MFCC), and the image data will be extracted through a convolutional neural network to extract visual features, and finally the different modal data will be stored in the form of feature vectors to provide a basis for subsequent cross-modal fusion.
[0043] For example, when the user interacts with the intelligent customer service terminal, the microphone of the terminal records the user's voice in real time (such as "help me check the weather tomorrow"), the camera captures the user's facial expression when speaking (such as a slight frown, which may indicate concern), and if the user holds a device with an accelerometer, the sensor will record the gesture action (such as pointing out the window). These data are transmitted synchronously to the preprocessing module of the edge terminal: the voice signal is converted into a spectrogram and acoustic features are extracted, the facial image is identified by key point detection to identify the expression type, and the gesture action data is standardized as a motion trajectory vector. Finally, a multi-modal data set containing voice text features, expression features, and action features is formed, which is used for subsequent analysis of user intent and emotional state.
[0044] Through the above step S101, the limitation of single text input can be broken through, and user interaction information can be obtained from dimensions such as voice tone, facial expression, and body movement, for example, when the user says "it doesn't matter", the true emotion can be judged to be negative in combination with the frown expression, avoiding semantic understanding bias. The edge terminal processes the original data locally, reducing data upload volume and delay, for example, after the voice data is extracted at the edge, only the feature vector needs to be transmitted instead of the complete audio, improving the system response speed. It is suitable for interaction requirements in complex environments, such as combining lip-reading visual features to assist speech recognition in noisy environments, or replacing text input with gesture actions in barrier-free interaction, expanding the scope of application of the system.
[0045] As an optional embodiment, it is assumed that the multi-modal data at least includes: voice data, video data, text data, and biological signal data collected by wearable devices. It is assumed that the lightweight Transformer fusion network at least includes: a feature extraction layer, a space-time alignment layer, a fusion layer, and an adaptive weight distribution layer. Based on the above assumptions, in step S102, the user edge terminal adopts the lightweight Transformer fusion network to convert the user feature vectors in different modalities into fusion feature vectors through cross-modal attention fusion and dynamic weight adjustment, including:
[0046] In the feature extraction layer, for different modal data, a sound emotion feature extraction module, a speech tone feature extraction module, a micro-expression feature recognition module, a text sentiment semantic analysis module, and a physiological change feature recognition module are constructed respectively; and the user feature vectors in different modalities are extracted from the multi-modal data through the constructed different modal processing modules; wherein the user feature vectors at least include: user sound emotion features, speech tone features, micro-expression features, text sentiment semantic features, and physiological change features;
[0047] Through the space-time alignment layer, the user feature vectors are projected into a unified dimensional space through a fully connected layer, and the user feature vectors in other modalities are positioned to each emotional semantic node with the emotional semantic nodes in the text emotional semantic features as the time axis reference, to construct a space-time consistent emotional expression view;
[0048] Through the fusion layer, combined with the emotional expression view, the emotional semantic nodes in the text emotional semantic features are taken as the time axis reference, and the different user feature vectors are used to perform cross-modal retrieval on other user feature vectors in the time order of the time axis reference to obtain corresponding matching feature vectors, and the fusion feature vectors in different branches are obtained based on the cross-modal retrieval results;
[0049] Through the adaptive weight distribution layer, the fusion feature vectors in different branches are subjected to signal quality evaluation and semantic consistency detection, and the weight parameters corresponding to different modalities are dynamically adjusted based on the evaluation results and detection results to obtain the final output fusion feature vectors, so as to eliminate data conflicts in the fusion feature vectors and improve the accuracy of the output results.
[0050] In the embodiments of the present application, the lightweight Transformer fusion network adopts a lightweight variant based on the Transformer architecture, such as MobileViT, TinyBERT, etc. The number of attention heads is reduced, and the hidden layer dimension is reduced to balance the calculation efficiency and cross-modal feature interaction capability. A lightweight version of the cross-modal attention network such as CLIP can also be introduced, which aligns the text, speech, visual and other modal features through a contrast learning mechanism, and optimizes the parameter scale for edge computing scenarios. A dynamic weight fusion model can also be constructed by referring to the Mixture of Experts (MoE) idea, which calculates the reliability score of each modality using a lightweight neural network, and realizes intelligent weighting of conflicting features.
[0051] In the above network, the sound emotion feature extraction module converts the speech into a spectrum graph through short-time Fourier transform, extracts frequency domain features in combination with Mel frequency cepstrum coefficient (MFCC), and then uses a lightweight CNN such as MobileNet to capture acoustic patterns corresponding to emotions such as anger and sadness. For example, when the user says “I’m fine”, the module can extract anxiety features from the spectrum of the voice tremor. The speech intonation feature extraction module analyzes prosodic features such as fundamental frequency and speech rate through bidirectional LSTM. For example, in the questioning tone of “Really?”, the timing feature of the rising intonation is extracted. The micro-expression feature recognition module uses a lightweight convolutional network (such as ShuffleNet) to detect facial key points and locate micro-expressions such as eyelid drooping and mouth twitching. For example, the action of the user’s eye corner pulling down when speaking is captured. The text sentiment semantic analysis module extracts semantic vectors of emotional words such as “sad” and “disappointed” through a lightweight version of BERT (such as ALBERT), and constructs semantic nodes in combination with syntactic analysis. The physiological change feature recognition module extracts stress response features from heart rate and skin electrical signals collected by wearable devices through wavelet transform, such as physiological signals of sudden heart rate increase. These modules process multi-modal data in parallel to provide multi-dimensional feature vectors for subsequent fusion, realize comprehensive feature capture from speech prosody to physiological response, and avoid semantic understanding bias of a single modality.
[0052] The spatio-temporal alignment layer projects each modality feature to a unified dimensional space through a fully connected layer, and takes the sentiment semantic node (such as the text segment of "I am very sad") in the text sentiment semantic feature as the time axis reference, and aligns the voice pause, micro-expression action and other features to the corresponding semantic node by using the dynamic time warping (DTW) algorithm. For example, when the user says the text "I am very sad", the sobbing feature extracted by the voice module and the low eyelid action captured by the visual module will be synchronously mapped to the timestamp of the text, forming a spatio-temporal matrix containing text semantics, voice pause and expression action. This alignment method anchored by text solves the problem of different sampling rates of multi-modal signals, and constructs a spatio-temporal consistent emotion expression view, so that the system can associate different modal emotion signals from the time dimension, for example, when analyzing the text "I am fine", the smile expression and the physiological signal of heart rate increase in the same period are combined to judge the conflict between the user's real emotion and the text semantics.
[0053] The fusion layer takes the time axis of the text semantic node as the reference, and realizes feature interaction through cross-modal attention mechanism. For each text semantic node (such as "I am very disappointed"), the voice feature vector will be used as a query vector to retrieve the corresponding micro-expression matching vector (such as strong and happy facial action) in the visual feature, and at the same time the text semantic vector will also retrieve the corresponding stress response vector (such as skin conductance signal mutation) in the physiological feature. This cross-modal retrieval is based on cosine similarity or dot product operation to find the semantic related feature pairs between different modalities, and then generate branch fusion features through concatenation or weighted summation. For example, when processing the text node of "I am very disappointed", the fusion layer will cross-modal correlate the pitch drop feature in the voice, the smile expression feature in the vision, and the heart rate stable feature in the physiology, find the conflict signal between the voice and the visual modalities, and provide a basis for subsequent weight adjustment. This mechanism enables different modal features to complement each other on the time axis, enhances the representation ability of complex emotions (such as disguised emotions), and avoids misjudgment caused by a single modality.
[0054] The adaptive weight allocation layer handles feature conflicts through a dual mechanism. On the one hand, the signal-to-noise ratio (SNR) and other indicators are used to evaluate the quality of signals of each modality. For example, the low signal-to-noise ratio of speech features in a noisy environment will be given a lower weight. On the other hand, a semantic consistency detection model (such as a lightweight BERT classifier) is used to calculate the matching probability of different modal features and text semantics. For example, when a user smiles and says "I am disappointed", the model will combine the historical records of the user's "social smile" in the knowledge graph, calculate the semantic conflict probability between the visual modality "smile" feature and the text "disappointment", and then reduce the weight of the visual modality. The weight adjustment is based on gradient descent dynamic optimization. The final output of the fusion feature vector will weaken the conflicting modalities (such as the visual features of social smile) and strengthen the consistent modalities (such as the features of disappointment in speech tone). Taking "smile to say disappointment" as an example, this layer will use historical data reasoning and context probability calculation to increase the weight of speech features to 0.6, reduce the weight of visual features to 0.3, and keep the weight of physiological features at 0.1. In this way, the modal conflict is eliminated, the fusion features more accurately reflect the user's true emotions, and the accuracy and robustness of emotion recognition are improved.
[0055] Further optionally, in the above steps, different modal user feature vectors are extracted from the multi-modal data through the constructed different modal processing modules, including:
[0056] The sound emotion feature extraction module extracts the user's sound emotion features from the speech data based on the Mel spectrum and CNN attention mechanism. The speech tone feature extraction module extracts the user's speech tone features from the speech data based on the mixed model of Transformer and BiLSTM. The micro-expression feature recognition module extracts the user's micro-expression features from the video data based on the VisionTransformer and spatio-temporal convolution network. The text emotional semantic analysis module extracts the user's text emotional semantic features from the text data based on BERT and attention mechanism. The physiological change feature recognition module extracts the user's physiological change features from the biological signal data using the LSTM model.
[0057] The sound emotion feature extraction module first converts the speech data into a Mel spectrum through a Mel filter bank to simulate the perception characteristics of the human ear to different frequencies of sound, and then uses the convolutional layer of CNN (such as ResNet) to capture the local acoustic patterns in the Mel spectrogram, and focuses on the key frequency band (such as the sad emotion often corresponds to the enhancement of low-frequency energy) by combining the attention mechanism. The specific process is: speech framing → short-time Fourier transform → Mel spectrum calculation → CNN feature extraction → attention weight allocation → emotion feature vector output. When the user says "I'm fine", the module converts the speech into a Mel spectrum, and the CNN detects energy fluctuations in the low-frequency region (80-300Hz), the attention mechanism gives high weight to this region, and combines the pre-trained model to identify the anxiety feature corresponding to the voice tremor, and outputs a feature vector containing a "nervous" emotion probability of 0.7. Precise stripping of emotion-related acoustic features from speech signals overcomes the ambiguity between semantics and emotion (such as the text "fine" may correspond to multiple emotions), improves the accuracy of emotion recognition, and provides acoustic dimension emotional evidence for multi-modal fusion.
[0058] The speech intonation feature extraction module uses a hybrid model of Transformer and BiLSTM to process speech prosody features. BiLSTM captures the long-term dependencies of time series features such as fundamental frequency, speech rate, and pause (such as the rising pattern of intonation at the end of a question sentence), and the self-attention mechanism of Transformer models the global association of each time step (such as the intonation fluctuation rhythm of the whole sentence). The module first extracts the prosodic parameters of the speech (fundamental frequency curve, energy entropy, etc.), and then encodes them into intonation feature vectors through the hybrid model. For example, when the user asks "Will it rain today?", the module extracts the frequency rising trend of the last syllable in the fundamental frequency curve (from 200Hz to 250Hz), BiLSTM records the timing changes of this rising pattern, and Transformer pays attention to the association between the intonation fluctuation of the whole sentence and the position of the question word "whether". Output a vector containing the "question" intonation feature, representing the intensity and rhythm of the rising intonation. Thus, it effectively captures the prosodic emotional clues in the speech, distinguishes the semantic differences between the same text with different intonations (such as the intonation difference between statement and question sentences), provides prosodic support for dialogue intent recognition, and enhances the time series expression ability of multi-modal features.
[0059] The micro-expression feature recognition module processes video frame sequences based on Vision Transformer (ViT) and spatio-temporal convolutional network. ViT captures global facial movements (e.g., the overall feature of mouth corner droop) by dividing the facial image into blocks and using self-attention mechanism. The spatio-temporal convolutional network analyzes muscle movement trajectories in consecutive frames (e.g., the time series of eyelid closure). The module first detects facial key points using MTCNN, then inputs the key point coordinate sequence into the model to extract spatio-temporal features of micro-expressions. For example, when a user speaks with a "sad" micro-expression, the module locates eye key points (e.g., eye corner, eyelid contour) in the video frames, ViT captures the global shape changes in the eye area, and the spatio-temporal convolutional network analyzes the eyelid descent trajectory in the last 3 frames (eyelid coordinates moving down in the 10th-12th frames of 25 frames per second), outputting a feature vector containing a "sad" micro-expression probability of 0.8. Extracting transient micro-expression features from the visual modality, recognizing implicit emotions that text and speech cannot express (e.g., socially masked smile), supplementing the spatio-temporal dynamic information of facial movements, and improving the subtle expression capture capability of multi-modal sentiment analysis.
[0060] Further, by detecting 68 facial key points (e.g., eye corner, nose tip) using algorithms such as Dlib or MediaPipe, the coordinate change sequence is input into LSTM or GRU to extract key point time series motion features. For example, a sad micro-expression may be accompanied by a continuous movement trajectory of the eye corner drooping. Further, texture analysis is performed on local areas (e.g., eye area, forehead), and 3D convolution is used to capture the dynamic changes of skin wrinkles. For example, an angry micro-expression may exhibit rapid contraction and relaxation of forehead horizontal wrinkles.
[0061] The text sentiment semantic analysis module uses the BERT pre-training model and attention mechanism to analyze the sentiment semantics of the text. BERT encodes the context semantics of the text (e.g., the difference in emotional intensity of "disappointment" in different contexts) through multiple layers of Transformer, and the attention mechanism focuses on emotional keywords (e.g., "sad" and "terrible"). Combined with syntactic analysis, semantic nodes are constructed. The module first inputs the text into BERT after tokenization, then extracts the sentiment semantic feature vector through the pooling layer. For example, for the text "Today's event was canceled, I'm so disappointed", BERT encodes the context association of "cancel" and "disappointment", the attention mechanism gives "disappointment" a high weight, and the output feature vector contains a "disappointment" sentiment dimension score of 0.9, as well as a semantic representation of "event cancellation". In this way, deep semantics and emotional polarity are extracted from the text, semantic nodes and emotional labels are constructed, and a semantic benchmark is provided for multi-modal fusion in the text dimension, ensuring that features from different modalities can be aligned to the semantic timeline of the text.
[0062] The physiological change feature recognition module uses an LSTM model to process the time sequence features of biological signal data (such as heart rate, skin electrical signal). The LSTM memory unit captures the long-term trend of physiological signals (such as a sustained increase in heart rate caused by stress), filters noise and retains key physiological responses through a gating mechanism. The module first denoises the biological signal for preprocessing, and then inputs the LSTM to extract a time sequence feature vector. For example, if the user's physiological signal shows a sudden increase in heart rate (from 70 beats per minute to 90 beats per minute) during a conversation, the LSTM model identifies the upward trend and its temporal correlation with the conversation content "work evaluation", and judges it to be a stress response based on historical data, outputting a vector containing stress physiological features representing the amplitude and duration of the change in heart rate. In this way, objective emotional stress features are extracted from physiological signals, making up for the masking of subjective expressions such as voice and facial expressions (such as physiological responses when pretending to be calm), providing objective evidence of the physiological dimension for multi-modal emotion analysis, and enhancing the reliability of emotion recognition.
[0063] As an optional embodiment, it is assumed that the dynamic knowledge graph at least includes: an emotional expression view. Based on the above assumption, in step S103, the received compressed data is input into the dynamic knowledge graph enhancement model through the cloud service platform, and the fusion feature vector and the entity knowledge and emotional relationship implied in the historical conversation are mined in real time during the conversation interaction to update the dynamic knowledge graph, including:
[0064] The decompression module decompresses the compressed data and restores the fusion feature vector and the context summary of the historical conversation; the extraction module extracts the user emotional features of the current round of conversation, the user feature vectors under different modalities, and the confidence of different modalities from the fusion feature vector; the conversation history reconstruction module supplements the missing information in the context summary to expand the context summary to complete historical conversation context information, in combination with the emotional expression view of the user stored in the cloud; the knowledge extraction module constructs entity knowledge nodes matched with the current round of conversation and semantic association relationships between each entity knowledge node based on the user feature vectors under different modalities and the historical conversation context information; the fusion module combines the user emotional features and the confidence of different modalities, and based on the message passing mechanism of the graph neural network GAT, fuses the newly added entity knowledge nodes into the emotional expression view of the user stored in the cloud to obtain the conversation subgraph of the current round of conversation.
[0065] The conversation subgraph of the current round of conversation at least includes: all entity knowledge nodes corresponding to the current round of conversation, and semantic association relationships between all entity knowledge nodes. The conversation subgraph is a local substructure of the knowledge graph generated by the current round of conversation, containing all entity nodes and associated edges of this round, and is an incremental update unit of the dynamic knowledge graph.
[0066] It is worth mentioning that the dynamic knowledge graph enhancement model (DKGE) is the core component of the cloud service platform, and its essence is a dynamic knowledge modeling system based on graph neural networks. Based on the sentiment expression view, the model continuously mines the entity knowledge and emotional relationship in the dialogue by real-time analysis of the fusion feature vector uploaded by the edge terminal and the historical dialogue data, and realizes the dynamic evolution of the knowledge graph. Its core function is to convert the multi-modal emotional signals and semantic information generated in the user interaction process into a structured knowledge network, enabling the knowledge graph to continuously accumulate personalized entity associations and emotional dependencies as the dialogue progresses, providing precise knowledge support for subsequent reply generation and strategy decision-making, and solving the problem that traditional static knowledge graphs cannot adapt to real-time dialogue scenarios.
[0067] The decompression module uses a semantic-based compression restoration algorithm to perform reverse processing on the compressed data uploaded by the edge terminal. By analyzing the keyword index and feature mask preserved during compression, the dimensions of the fusion feature vector and the context summary of the historical dialogue are restored to their original format. For example, a 128-dimensional fusion vector compressed by feature selection is restored to a 512-dimensional feature vector containing speech, vision, and physiological modalities, and the keyword-compressed dialogue summary is expanded to a text sequence containing timestamps.
[0068] Compressed data is lightweight data obtained by semantic compression of multi-modal features and historical dialogue by the edge terminal, containing the key dimensions of the fusion feature vector and the dialogue summary, and is used to reduce network transmission load. Further, the dynamic compression method based on reinforcement learning learns through the interaction between the agent and the data environment, dynamically adjusting the compression strategy. The agent can decide which compression algorithm and compression ratio to use in real time based on the characteristics of multi-modal data, such as the importance of the data and the transmission bandwidth status. For example, when low network bandwidth is detected, the agent increases the compression ratio of the speech data, while learning how to preserve emotional features in the speech under high compression ratio through reinforcement learning algorithms. For text data containing important semantic information, the compression ratio is appropriately reduced. Through continuous trial and error and reward feedback, the agent can find the optimal compression strategy to maximize network transmission load reduction while ensuring data effectiveness.
[0069] The compression method based on the variational autoencoder (VAE) uses the encoding-decoding structure of VAE to compress and restore data. The encoder maps the original multi-modal data to a low-dimensional latent space. In this process, by constraining the distribution of the latent space, the compressed data can not only retain the key information of the original data, but also have a certain generalization. The decoder reconstructs the original data from the low-dimensional representation of the latent space. Taking historical dialogue data as an example, VAE can compress the lengthy dialogue text into a short vector containing core semantics and sentiment orientation, and restore it when needed. This compression method not only preserves the integrity of the data semantics, but also greatly reduces the data volume, and has good robustness when facing noisy data.
[0070] The context summary is the key information of the compressed dialogue history, which retains entity keywords and semantic fragments, but needs to be completed with the history reconstruction.
[0071] For example, when receiving the compressed data uploaded by the edge terminal "[user features: compressed vector {voice 0.7, vision 0.5}, dialogue summary: 'bad weather -> mood']", the decompression module will restore the voice feature dimension to 20-dimensional Mel Frequency Cepstral Coefficients (MFCC) according to the pre-defined compression dictionary, restore the visual feature to 68-dimensional coordinates of facial key points, and expand the dialogue summary to the complete text "user said 'today the weather is very bad, and the mood is not good'".
[0072] Thus, lossless or approximately lossless restoration of compressed data is achieved, ensuring that the fusion features and historical dialogue information obtained by the cloud have complete semantic details, providing accurate data basis for subsequent knowledge extraction, and avoiding the influence of feature loss caused by compression on the accuracy of knowledge graph update.
[0073] The extraction module separates the user emotion features and modality-specific features from the fusion feature vector through multi-dimensional feature analysis algorithms. It uses a pre-trained sentiment classifier (such as a BERT-based sentiment analysis model) to extract the sentiment polarity (such as "sadness" "joy") and intensity value of the current round from the fusion vector; through the modality mask matrix, it separates the feature vectors of different modalities such as voice tone and micro-expression, and calculates the confidence based on indicators such as variance and entropy value of each modality feature (such as the confidence of voice features in a noisy environment).
[0074] For example, for the fusion feature vector "voice features [flat tone 0.3, slow speech speed 0.8], visual features [drooping corners of the mouth 0.6, drooping eyelids 0.9]", the extraction module will judge the overall emotion to be "low" through the sentiment classifier, and calculate the voice modality confidence 0.6 (affected by environmental noise) and the visual modality confidence 0.9 (clear image).
[0075] Thus, the precise deconstruction of multi-modal features is realized, the core emotional features and modal details are separated, the knowledge graph is provided with the annotation information of the emotional dimension, the confidence evaluation provides the weight basis for subsequent modal fusion, and the reliability of knowledge extraction is improved.
[0076] The dialogue history reconstruction module matches the current context summary with the user emotional expression view stored in the cloud based on a time sequence association algorithm. By identifying entity keywords (such as “weather” and “mood”) in the summary, the relevant emotional expression view segments in the historical dialogue are retrieved to supplement the missing timestamp, previous dialogue semantics and other information in the summary. For example, if the current summary is “user mentions weather → low mood”, the module will find the record of “user discusses weather → happy” in the previous three rounds in the historical view, and complete the context logic.
[0077] For example, if the context summary of the current compressed data is “user says ‘don’t want to go out’”, the reconstruction module will retrieve the previous two rounds of dialogue “user mentions ‘rain’ → low voice tone → increased heart rate” in the cloud-stored emotional expression view, complete the causal relationship between the missing “rain” entity and “don’t want to go out”, and form a complete historical context “rain → don’t want to go out → low mood”.
[0078] In this way, the problem of missing context information caused by compressed data is solved, the complete dialogue logic chain is constructed through knowledge reuse of the historical emotional expression view, and knowledge extraction can be based on full context, avoiding entity relationship misjudgment caused by information fragmentation.
[0079] The knowledge extraction module identifies entities and relationships from multi-modal features and historical dialogue based on a cross-modal semantic association algorithm. The named entity recognition (NER) model extracts entities such as “weather” and “mood” from text; visual features are used to supplement visual entities (such as umbrella images captured by a camera); the co-occurrence analysis and time sequence association are used to construct the relationship between entities (such as “rain → impact → mood”), and the emotional signal in the multi-modal feature (such as low voice tone) is used as the emotional weight between entities.
[0080] For example, in the user dialogue “It’s raining today, and the mood is good or bad”, the knowledge extraction module extracts entities “rain” and “mood” from the text, identifies the entity “umbrella” corresponding to the umbrella image from the visual features, constructs the semantic relationship “rain → causes → bad mood” through time sequence association, and uses the low voice tone (0.8 intensity) in the voice feature as the emotional weight of the relationship, forming an entity association with emotional annotation.
[0081] In this way, the conversion from multi-modal data to structured knowledge is realized, and the semantic information and emotional signals in user interaction are abstracted as nodes and edges in the knowledge graph, so that the knowledge graph can express entity semantic association and store dynamic characteristics such as emotional intensity, thereby improving the semantic richness and emotional expression ability of the knowledge graph.
[0082] The fusion module fuses the new entity knowledge node into the existing emotional expression view based on the message passing mechanism of the graph attention network (GAT). The message passing mechanism is the process of feature propagation between nodes in GAT, which allows nodes to aggregate information of neighbor nodes according to attention weights, realizing semantic fusion of the knowledge graph. The entity knowledge node is a node representing a specific concept or object in the knowledge graph, which contains semantic attributes and emotional characteristics (such as the “rain” node with the “wet” attribute and the emotional association of “affecting mood”). The emotional expression view is a multi-modal emotional association graph based on a time axis, which is used to store the spatio-temporal alignment relationship between emotional signals and semantic nodes in the dialogue. The semantic similarity between the new node and the existing node is calculated through the attention mechanism (such as the association degree between “rain” and the historical node “weather”), the features of the existing nodes (such as the emotional tendency in the historical weather dialogue) are passed to the new node through the message passing mechanism, and the weight of the edge is adjusted according to the user emotional characteristics and modal confidence. For example, when a new “mood bad” node is added, GAT will calculate the association strength between it and the “rain” node, and combine the high-confidence “frown” feature of the visual modality to enhance the weight of the “rain→mood bad” edge.
[0083] For example, when the entity “umbrella” and the relationship “rain→use→umbrella” are added, the fusion module calculates the semantic similarity between “umbrella” and the historical node “rain” (0.7) through GAT, and assigns the historical emotional features stored in the “rain” node (such as the user's 2 times of emotional decline in the past 3 times of rain dialogue) to the “umbrella” node through message passing, and sets the weight of the relationship edge to 0.7x0.9=0.63 according to the recognition confidence of the visual modality to the umbrella 0.9, forming an enhanced association with emotional transmission.
[0084] Therefore, the dynamic attention mechanism of GAT is used to realize the incremental update of the knowledge graph, so that the new knowledge can interact with the historical knowledge semantically, and through the weighting of emotional features and modal confidence, the emotional relationship in the knowledge graph can accurately reflect the emotional intensity and modal reliability in the user's real-time interaction, thereby improving the dynamic evolution ability of the knowledge graph.
[0085] Further optionally, in the above steps, the missing information in the context summary is supplemented by the dialogue history reconstruction module in combination with the emotional expression view of the user stored in the cloud to expand the context summary into complete historical dialogue context information, including:
[0086] retrieve corresponding emotional semantic nodes from the user's emotional expression view indexed by the context summary; query corresponding historical user feature vectors in other modalities using the retrieved emotional semantic nodes; predict corresponding historical dialogue information according to the emotional development trend based on the queried historical user feature vectors, and merge and construct the predicted historical dialogue information into the historical dialogue context information according to the order of the emotional semantic nodes in the time axis.
[0087] Further optionally, in the above step, the knowledge extraction module constructs entity knowledge nodes matched with the current round of dialogue and semantic association relationships between the entity knowledge nodes based on the user feature vectors in different modalities and the historical dialogue context information, including:
[0088] extract key named entities matched with the current round of dialogue and the dialogue topic from the user feature vectors in different modalities, and extract semantic association relationships between the key named entities in combination with the historical dialogue context information; based on the extraction results, establish entity knowledge nodes corresponding to the key named entities and semantic association relationships between the entity knowledge nodes.
[0089] For example, the knowledge extraction module uses cross-modal collaborative perception and semantic deep mining technology to break through the limitations of traditional single-modal analysis. The core principle is to construct a joint representation model based on multi-modal Transformer. This model simultaneously processes acoustic features of speech, semantic vectors of text, visual features of video, and fluctuation data of physiological signals through a multi-head attention mechanism. For example, when a user excitedly says, "Just got tickets to someone's concert!" in a video call, the system not only identifies named entities such as "someone", "concert", and "tickets" from the text, but also detects the user's excited features such as rising intonation and accelerated speech rate through the voice module, captures the user's facial expressions such as raised eyebrows and bright eyes using the camera, and combines the heart rate acceleration data monitored by the wearable device to verify and enhance the accuracy of entity extraction from multiple dimensions.
[0090] In the relationship extraction stage, the module innovatively introduces a temporal graph convolutional network (TGCN) to model historical dialogue data as a temporal knowledge graph. By analyzing the co-occurrence patterns and semantic evolution of entities in the time dimension, the module can mine deep semantic associations between entities. For example, if a user has expressed multiple times that they love a certain person's music, the system can infer a stable relationship of "user-loves-someone" and further derive potential semantic connections such as "user-anticipates-someone's concert." In addition, the module uses an adversarial training mechanism to enhance the model's robustness to noisy data and ambiguous expressions through a generative adversarial network (GAN). This allows the model to accurately identify referential expressions such as "some Dong's performance" and "that amazing Live."
[0091] This multi-modal collaborative knowledge extraction approach improves the accuracy and recall rates of entity and relationship extraction in practical applications. Compared to traditional methods, this module improves entity recognition accuracy and relationship extraction accuracy in complex dialogue scenarios, enabling accurate capture of implicit information and potential associations in user expressions and providing rich and accurate knowledge reserves for the construction of dynamic knowledge graphs.
[0092] Furthermore, through the fusion module, the user's emotional features and the confidence of different modalities are combined, and based on the message passing mechanism of graph neural networks, the newly added entity knowledge nodes are fused into the user's emotional expression view stored in the cloud, including:
[0093] Based on user emotional features and the confidence of different modalities, corresponding emotional labels are added to the semantic association relationships between newly added entity knowledge nodes. The emotional labels represent the common emotional state types of the connected entity knowledge nodes. According to the propagation trend of the subgraph expansion, an emotional reasoning chain is formed, which includes the newly added entity knowledge nodes and semantic association relationships. A dynamic pruning mechanism is used to delete emotional reasoning chains with overall confidence below a certain threshold from the emotional expression view. The node states of historical dialogues in the emotional expression view are inherited across dialogue turns and transferred to the newly added entity knowledge nodes and semantic association relationships.
[0094] For example, the fusion module realizes real-time updating and emotional evolution of the knowledge graph based on a dynamic graph neural network and a sentiment reasoning mechanism. The principle is to use the message passing mechanism of the graph attention network (GAT), combine the user's emotional feature vector and the confidence score of each modal data, and perform sentiment labeling and dynamic fusion on the newly added entity knowledge nodes and semantic association relationships. For example, when the user says with a tinge of regret: "It's a pity that the concert is on a weekday, and I might not be able to go", the system first adds the "regret" sentiment label to the "concert-time-weekday" and "user-attitude-regret" relationships based on the low tone in the voice (confidence score 0.85), negative words in the text, and the frown action in the micro-expression (confidence score 0.8).
[0095] Next, based on the sentiment label and semantic association, the module uses the Spreading Activation Algorithm to construct and extend the emotional reasoning chain in the knowledge graph. Taking "concert-time-weekday" as the starting point, combined with the user's complaints about being busy at work in the historical dialogue, the complete logical chain of "weekday-causes-busy at work-makes-user unable to attend the concert" is inferred. At the same time, to ensure the quality of the knowledge graph, the module introduces a dynamic pruning mechanism driven by reinforcement learning, calculates the overall confidence of the reasoning chain based on the confidence product of each node and edge, automatically deletes weakly associated paths below the threshold (such as 0.6), and avoids the accumulation of invalid information.
[0096] In addition, the module also sets up a cross-dialog turn memory transfer mechanism, which transfers the emotional state and knowledge experience related to the current topic in the historical dialogue to the new nodes through attention weighting. For example, if the user has expressed regret about time conflicts in discussing other activities before, the system will transfer this emotional tendency and coping strategy to the knowledge graph of the current dialogue, enhancing the emotional coherence and reasoning ability of the knowledge graph.
[0097] In practical applications, the fusion module enables the dynamic knowledge graph to quickly adapt to changes in the dialogue scene, and updates the knowledge structure and emotional expression in real time. The dialogue system using this module improves the accuracy of sentiment understanding and performs better in complex semantic reasoning tasks, effectively realizing dynamic accumulation of knowledge and accurate expression of emotions, enhancing the naturalness and intelligence of dialogue interaction.
[0098] As an optional example, suppose the dynamic knowledge graph also includes an expert knowledge base associated with each entity knowledge node in the emotional expression view and each domain knowledge graph. Based on this, based on the updated dynamic knowledge graph, knowledge enhancement processing is performed to construct an initial reply prototype for the user's current turn interaction, including:
[0099] extract, from the dialogue subgraph of the current turn dialogue, current interaction intent information of the user in the current dialogue turn; extract, from the dialogue subgraph corresponding to the historical dialogue of the emotion expression view, historical interaction intent information of the user in the previous historical turn dialogue; perform intent analysis on the current interaction intent information and the historical interaction intent information to obtain a semantic framework for representing the user interaction intent; take each entity knowledge node in the dialogue subgraph of the current turn dialogue as an index to search for corresponding entity knowledge information from an expert knowledge base and / or each domain knowledge graph, and generate a knowledge filling unit based on the searched entity knowledge information; the knowledge filling unit at least includes a dialogue content template constructed based on the entity knowledge information and a to-be-filled slot; the to-be-filled slot is supplemented in the user edge terminal to improve cloud edge transmission efficiency; adopt an emotion-knowledge mapping mechanism, take the emotion labels between each entity knowledge node in the dialogue subgraph of the current turn dialogue as an index, query, in the emotion correlation matrix corresponding to the current turn dialogue, an emotion change type matching the emotion development trend as a reply emotion orientation of the current turn dialogue; based on the semantic framework, the knowledge filling unit, and the reply emotion orientation determined in the foregoing steps, construct the initial reply prototype.
[0100] Firstly, in the above steps, the depth understanding of the user intent is realized through cross-turn dialogue intent association analysis and dynamic semantic modeling. The principle is to use a Transformer-based multi-turn dialogue encoder to jointly model the entity nodes in the current dialogue subgraph and the intent sequence of the historical dialogue subgraph, and capture the intent evolution clues through the attention mechanism.
[0101] Among them, the dialogue subgraph of the current turn dialogue is essentially a local instantiation product of the dynamic knowledge graph in a specific dialogue scenario. Taking the current user input as the core, key entities and semantic relationships are extracted from multi-modal features (text, speech, expressions, etc.) to form a micrograph structure focusing on the current interaction theme. The construction logic is similar to "close-up" in the global knowledge graph to crop the part strongly related to the current dialogue. Through the triple structure of entity knowledge nodes (such as named entities), semantic association edges (such as "belongs to" and "association"), and emotion labels (such as "positive" and "question"), the user's immediate intent is converted into a graph data model understandable by machines. This subgraph does not exist independently, but forms a "graph-graph association" with the historical dialogue subgraph and the global emotion expression view, and realizes the semantic coherence of the dialogue context through the cross-turn node state inheritance mechanism.
[0102] For example, in a medical consultation scenario, the user says in the current turn: "Taking aspirin makes my stomach uncomfortable" (the current intent is drug side effect consultation), the system extracts "asked about ibuprofen contraindications last week" (the historical intent is drug selection) from the historical dialogue subgraph, and finds that the user has a persistent concern about "non-steroidal anti-inflammatory drug side effects" through comparative analysis. Technically, the dynamic time warping (DTW) algorithm is used to align the intent semantic space of different turns, so that "allergy history" "dose" and other slots in the historical intent can be automatically associated with the current intent, forming an intent chain across turns. This approach can improve the accuracy of intent understanding in multi-turn dialogue compared to traditional single-turn intent recognition, especially in scenarios that require long-term tracking (such as chronic disease management), avoiding response bias caused by ignoring historical information.
[0103] The intent analysis module uses frame semantics and dynamic slot generation technology to convert free text intent into structured semantic framework. The principle is to identify the core predicate (such as "query" "side effect") in the intent through a pre-trained frame semantic model (such as BERT-Framenet), and generate corresponding argument slots (such as "drug name" "symptom performance") according to the domain knowledge graph. For example, in the education field, the user says: "Recommend a physics primer for middle school students", the system analyzes the "resource recommendation" framework, which includes "education stage = middle school students" "discipline = physics" "resource type = primer" and other slots, and extracts the "previously recommended mathematics books" framework from the historical dialogue, automatically supplementing the implicit slot "user preference = picture and text". The process introduces a dynamic slot pruning mechanism, which calculates the information gain rate of the slot to the intent, and deletes non-key slots such as "publication date", so that the framework dimension is reduced. The technical effect is to form a standardized semantic representation, providing clear structured guidance for subsequent knowledge retrieval, which can improve the efficiency of knowledge matching in intelligent customer service scenarios.
[0104] Further, the dynamic knowledge graph and the expert knowledge base are combined for graph-text retrieval. The principle is to use the embedding vector of the entity knowledge node (such as the semantic representation generated by the TransE model) as the query key to perform approximate nearest neighbor search in the expert knowledge base, and to combine the relationship path of the domain knowledge graph for reasoning expansion. For example, in legal consultation, the user's current dialogue subgraph contains "contract breach" and "compensation amount" entity nodes. The system not only retrieves the provisions of Article 577 of the Civil Code, but also generates a template containing "evidence type" and "calculation method" filling slots through the relationship chain of "contract breach → compensation calculation → evidence responsibility" in the domain graph. The knowledge filling unit adopts a hierarchical structure of "core content + variable slots", such as: "According to Article 577 of the Civil Code, the breaching party shall bear the responsibility of continuing performance, taking remedial measures, or compensating for losses (fixed content), among which the calculation of compensation amount needs to consider __ loss type __ and __ evidence materials __ (to be filled in slots)". This setting makes the cloud only need to transmit the template framework (about 1 KB) instead of the complete text (about 10 KB), and the edge terminal fills in the slots according to the local user data, reducing the cloud-edge transmission volume, while ensuring the accuracy and scene adaptability of the knowledge.
[0105] Through graph propagation and emotion trend prediction of the sentiment expression view, the joint modeling of knowledge and emotion is realized. The principle is to use the emotion label (such as "anxiety" and "doubt") in the dialogue subgraph as the query signal of the graph neural network, and perform diffusion activation in the emotion correlation matrix to calculate the transition probability of each emotion type. For example, the user said in the financial investment consultation: "Recently the stock market has fallen sharply, and the principal has lost 20%" (emotional label "anxiety"), the system finds through the message passing mechanism of GAT that the "anxiety" emotion often accompanies the knowledge demand of "risk tolerance assessment" in historical dialogue, and the emotional weight of the "stop-loss strategy" entity node in the current dialogue subgraph is 0.8, so the reply emotion orientation is determined as "calm + rational suggestion". The emotion entropy reduction optimization objective is innovatively introduced to preferentially select the reply style that can reduce the user's emotional uncertainty, such as when the user's emotional entropy value is higher than the threshold, the proportion of soothing statements is automatically increased. In the psychological counseling scene experiment, this mechanism improves the emotional matching degree of the reply, and the user's subsequent interaction participation is improved.
[0106] The reply prototype construction module adopts a multi-source information fusion generation framework to structurally integrate semantic frameworks, knowledge filling units, and emotion guidance. The principle is to first align the slots of the semantic framework with the to-be-filled items of the knowledge filling unit (such as matching the "drug name" slot with the "aspirin" entry in the knowledge base), and then adjust the expression of the content according to the emotion guidance (such as converting "side effects include stomach pain" to "some users may experience mild stomach discomfort, and it is recommended to take with meals"). Taking intelligent health management as an example, the current user intent is "hypertension medication consultation", and the historical intent includes "exercise recommendation". The prototype structure constructed by the system is as follows:
[0107] Emotion guidance layer: "Your concern about blood pressure control is very important, and we will answer you in detail" (soothing tone)
[0108] Knowledge filling layer: "The mechanism of action of commonly used antihypertensive drugs such as __drug name__ is __mechanism of action__, and attention should be paid to monitoring __monitoring indicators__ during medication" (slot from knowledge base retrieval)
[0109] Historical association layer: "Combined with your previous consultation of aerobic exercise, it is recommended to avoid intense exercise __time__ after taking medication" (cross-turn intent association)
[0110] The prototype generation process introduces an adversarial rewriting mechanism to ensure that the generated content not only meets the knowledge accuracy (such as the correctness of the mechanism of action of the drug), but also meets the emotional fluency (such as tone consistency). Here, the generated reply prototype has both knowledge depth and emotional temperature. In the medical question and answer scenario, compared with traditional template generation methods, user satisfaction with the reply is improved, and information retention rate is improved.
[0111] As an optional embodiment, when the cloud constructs an initial reply prototype through a dynamic knowledge graph enhancement model (DKGE), it needs to be differentiated according to the differences in the capabilities of edge terminals.
[0112] Specifically, by quantifying the terminal computing power, memory, and emotion rendering capability, for example, high-performance terminals (such as smartphones, ) can receive complete prototypes (including 3D animations); low-power devices (such as smart watches, ) only receive text summaries. Among them, GFLOPS is the computing power of the central processing unit of the edge node, measured in billions of floating-point operations per second (GFLOPS). For example, the CPU computing power of a smartphone is usually 10-100 GFLOPS, while a home gateway may reach more than 200 GFLOPS. Measuring the ability of a node to handle complex tasks, such as 3D emotional animation rendering, requires high While text summarization requires less computing power.
[0113] Memory capacity (GB), the size of the node's random access memory, directly affects data caching and multitasking capabilities. For example, smartwatch memory is usually 1-2 GB, while high-performance edge servers can reach more than 16 GB. The number of tasks that a node can handle simultaneously, such as loading large knowledge sub-graphs, requires sufficient memory support.
[0114] Emotional rendering capability (0-10), measured by quantitative indicators to measure the output capability of the node in emotional interaction, including: emotional expressiveness of speech synthesis (such as tone range); emotional adaptability of visual rendering (such as image tone adjustment, animation smoothness); synchronization of multi-modal fusion (such as the matching degree of voice and expression animation).
[0115] Furthermore, the matching degree formula is used to calculate the adaptability of semantic units and terminals, ensuring that high-emotion-weight units (such as =0.9) are preferentially distributed to devices with strong emotional rendering capabilities ( =0.8), avoiding emotional distortion caused by low terminals.
[0116] where and represent the importance proportion of data type matching degree and emotional adaptability, and sum to 1. In high-emotion-demand scenarios (such as psychological counseling), the can be dynamically increased to 0.6 to strengthen the matching priority of emotional rendering capability. Data type matching degree. Data type of the current semantic unit (such as text, link, structured data), is the set of data types supported by the edge node . Emotional weight of semantic unit, indicating the emotional intensity or orientation of the unit. Here, the emotional dimension can correspond to specific emotional types (such as anxiety, comfort) or intensity (such as weak to strong), output by the upstream emotional analysis module. Node's emotional rendering preference. Indicates the node's tendency in emotional expression, calibrated by device capability. By calculating the difference between emotional needs and node capabilities, the smaller the difference, the higher the matching degree. For example, =0.9 strong emotional units and The node difference is 0.1 when the matching degree is 0.8, and the corresponding matching degree term is 1-0.1=0.9.
[0117] After constructing the initial reply prototype for the current round of interaction with the user, the process of delivering it to the user's edge terminal faces challenges such as low latency, resource adaptation, and emotional fidelity.
[0118] Here, the hierarchical distribution architecture is centered on semantic driving and edge node adaptation, aiming to solve the problem of terminal heterogeneity. The principle is to disassemble the initial reply prototype into independent units according to semantic logic, and distribute it differently according to data sensitivity and terminal capability. Through dynamic semantic segmentation, the prototype is divided into "problem attribution" and "solution" modules, and data types and emotional weights are labeled. In the medical consultation scenario, "drug efficacy description" and "psychological support link" are split into different units, with the former marked as text type and emotional weight 0.7, and the latter as link type and emotional weight 0.9. At the same time, based on the security grading strategy, medical records containing user privacy are only delivered to trusted edge nodes, and general suggestions are directed to all terminals. In the edge node dynamic matching link, each terminal maintains a capability profile containing information such as computing power and memory, and the cloud assigns tasks accordingly: high-performance devices receive complete prototypes for 3D emotional animation rendering, and low-power devices only get text summaries and emotional symbols. If the device resources are insufficient, neighboring nodes can cooperate to complete the rendering. For example, when a user uses a smart watch to consult health problems, the cloud distributes complex graphic text replies to the home gateway, which proxies the rendering and pushes simple text to the watch, while the phone displays the complete content. This architecture improves the resource utilization rate of edge devices and effectively avoids rendering failures or resource waste caused by terminal performance differences.
[0119] Further, the emotion weight-based incremental streaming focuses on ensuring the timeliness and accuracy of emotional responses. The principle is to prioritize the transmission data according to the emotion label, and establish an emotion priority queue. For high anger, sadness and other emotions that may trigger crisis intervention (P0 level), the highest transmission priority is given, and low priority tasks can be interrupted; neutral or low intensity emotions (P2 level) are used for daily question and answer. For example, when a user expresses suicidal thoughts in an emotional state, the P0 level data containing the psychological hotline immediately occupies the regular information transmission channel, ensuring delivery within 200ms. At the same time, streaming progressive rendering is adopted, and the terminal gradually displays the reply content in the order of reception, first presenting the core conclusion, and then loading supporting data and emotional suggestions. Text units are embedded with emotion micro-labels to guide the terminal to adjust speech synthesis parameters, such as soft speed and low pitch for sad emotions. For example, when a user complains about product problems, the terminal first displays the soothing conclusion "very understand your distress", and then supplements the solution, while the voice is output in a concerned tone. This transmission method greatly shortens the response time in crisis situations, and reduces the waiting perception through progressive rendering, making data transmission more in line with user emotional needs.
[0120] In addition, further, context-aware delivery optimization dynamically adjusts the delivery strategy by real-time sensing of user state and network environment. In principle, biological signals (such as heart rate, skin conductance) collected by wearable devices and terminal positioning information are used to determine user emotions and the scene in which they are located. When an abnormal increase in heart rate is detected, the system determines that the user is in an anxious state, and at this time the knowledge unit is compressed and emotional comfort content is added. If the user is in a sensitive environment such as a hospital, medical terminology is replaced with euphemistic expressions. In terms of network adaptation, the multi-modal delivery strategy is dynamically adjusted according to the bandwidth: a complete prototype with text and pictures is provided in high bandwidth, and only text summary and emotional symbols are sent in low bandwidth, and combined with local cache to pre-fetch commonly used knowledge. For example, when a user asks for navigation while hiking outdoors, the phone only receives the text route in poor network conditions, while pre-caching nearby attraction information. When the network is restored at home, the attraction pictures and introduction are automatically loaded. This optimization reduces bandwidth requirements, ensuring that responses meet user emotional needs and stable transmission in complex network and user state change scenarios, avoiding distortion of emotional expression due to network fluctuations.
[0121] The security and privacy enhancement setting is aimed at the risk of data leakage in knowledge transmission, and adopts differential privacy injection and edge lightweight decryption technology. The principle is to add semantic noise in the knowledge unit to blur the original data, such as changing the accurate "targeted drug remission rate 85%" to "most patients feedback significant remission", which preserves the core semantics while protecting the data details. For high-sensitive data, use the SM4 encryption algorithm, and the key is hosted by the terminal trusted execution environment (TEE), ensuring that the data is only decrypted in a secure environment and immediately destroyed after use. For example, the diagnosis report in medical consultation is encrypted before transmission, and after reaching the home health terminal, it is decrypted and displayed by the terminal TEE, preventing data from being stolen or tampered with throughout the process. The above embodiments meet the strict compliance requirements of GDPR, HIPAA, etc., ensuring the security of knowledge transmission on the premise of protecting user privacy, and improving the trust of users in the system.
[0122] As an optional embodiment, in the above steps, the fusion feature vector and the updated dynamic knowledge graph are input into a dialogue strategy model through a cloud service platform to determine the response strategy of the current round of dialogue and the knowledge calling direction, including:
[0123] The fusion feature vector and the dialogue subgraph of the current round of dialogue are input into the dialogue strategy model. Through the dialogue strategy model, the corresponding emotional state features are extracted from the fusion feature vector according to the three dimensions of emotion type, emotion intensity, and emotion mixed components. Through the cross-modal cross-validation mechanism, the dominant emotion type of the user in the current round of dialogue is determined based on the emotional state features under the emotion type and the emotion mixed components. Based on the preset response strategy library of the collaborative mapping of the dominant emotion type and the emotion intensity, the response strategy matching the current round of dialogue is selected. The knowledge type to which each entity knowledge node in the dialogue subgraph of the current round of dialogue belongs is identified. Based on the proportion of the knowledge type occupied by the entity knowledge node, the knowledge base type matching the current round of dialogue is determined as the knowledge calling direction of the current round of dialogue.
[0124] For example, the dialogue strategy model realizes accurate interaction strategy decision through multi-dimensional sentiment analysis and knowledge graph reasoning. The dialogue strategy model can be a hybrid expert model (MoE) and an attention mechanism, which dynamically routes different types of dialogue requests to specialized expert modules through a gating mechanism. For example, in medical consultation, the "symptom analysis" expert handles diagnostic problems, and the "psychological support" expert handles emotional soothing. The gating network determines which experts to activate based on input features such as emotional intensity and keywords. In the gating function, add emotional features, such as activating the "empathy expert" when the anxiety emotion is strong; according to the knowledge type of the dialogue subgraph (such as drugs, surgery), assign it to the corresponding knowledge base module, and a certain tumor consultation system improves the knowledge matching accuracy rate in this way; through parameter efficient fine-tuning, only update the expert module parameters, keep the shared layer unchanged, and greatly reduce the cross-domain adaptation cost.
[0125] The principle is to construct a three-dimensional emotional deconstruction framework to extract emotion types (such as anger, joy), emotional intensity (numerical quantification from weak to strong), and emotional mixed components (such as 70% anxiety and 30% expectation of composite emotions) from the fused feature vector. For example, a user expresses excitement in intelligent financial consulting: "Stocks plummeted by 20%, what should I do?", The model analyzes the emotion type as "anxiety", the intensity reaches 8 levels (full score 10 levels), and the mixed components include "worry about financial security (60%)" and "seek solutions (40%)". The cross-modal cross-verification mechanism confirms that "anxiety" is the dominant emotion by comparing the consistency of multi-modal signals, and then retrieves the response strategy "soothing priority + professional advice" from the preset strategy library, such as sending the soothing statement "understand your worries, we will analyze for you immediately" first.
[0126] Knowledge calling direction is to determine the guiding path of extracting specific type of knowledge from knowledge base according to current dialogue content, user demand and context information in the process of dialogue interaction. Its core is to accurately locate the category of the required knowledge by analyzing entities, semantics and emotions in the dialogue, so as to realize efficient and accurate knowledge retrieval and application. Combined with the emotional guidance in response strategy (such as giving priority to calling solution type knowledge when the user is angry), it can avoid the disconnection between knowledge calling and emotional demand.
[0127] In the determination of the knowledge calling direction, the model classifies the entity knowledge nodes in the dialogue subgraph by type (such as financial concepts, operation processes, and risk cases). Taking a financial planning dialogue as an example, nodes such as "stock crash" and "capital loss" all point to the "investment risk" knowledge type, and by calculating the proportion of nodes of this type to reach a certain value, the "investment risk response" knowledge base is determined to be called. This combined decision mechanism of emotion and knowledge dimensions improves the accuracy of dialogue strategies compared to traditional single rule matching, especially in dealing with complex emotions (such as anxiety implying anger) and professional knowledge calling, which can effectively avoid strategy misjudgment and knowledge mismatch problems.
[0128] As an optional embodiment, in step S106, through the user edge terminal, the real-time interactive reply information output to the user is generated according to the initial reply prototype, the response strategy, and the knowledge calling direction, including:
[0129] Based on the knowledge calling direction, load the locally pre-stored knowledge subgraph; wherein, according to a preset strategy, monitor whether the knowledge calling direction of the current round of dialogue is consistent with the knowledge calling direction of the historical dialogue; if not consistent, call the matching knowledge base in advance, and segment the corresponding knowledge subgraph from the matching knowledge base; according to the knowledge filling unit and the reply emotion orientation in the initial reply prototype, load the matching entity knowledge information from the knowledge subgraph, and fill it into the to-be-filled slot of the knowledge filling unit; based on the response strategy, convert the filled knowledge filling unit into a complete sentence, and based on the semantic framework in the initial reply prototype, perform rationality detection on the complete sentence to obtain first real-time interactive reply information; through a lightweight rendering engine, dynamically adjust the language style and / or visual style in the first real-time interactive reply information according to the reply emotion orientation in the initial reply prototype to obtain second real-time interactive reply information finally output, so as to avoid that the second real-time interactive reply information contains negative target points of the emotion change type.
[0130] Specifically, the user edge terminal uses dynamic knowledge loading and strategy content generation technology to convert cloud instructions into natural interactive replies. The core principle is to establish a dynamic monitoring mechanism for knowledge calling. When it is detected that the knowledge calling direction of the current dialogue (such as "investment risk") is inconsistent with that of the historical dialogue (such as "product yield"), the corresponding knowledge subgraph is segmented from the locally stored knowledge base in advance to avoid delays caused by real-time requests. For example, when the user shifts from consulting fund yield to stock risk, the terminal immediately calls the "stock risk control" knowledge subgraph, which contains entity knowledge such as "stop-loss strategy" and "market volatility analysis".
[0131] Based on the filling process of the initial reply prototype, the reply emotion guidance and the knowledge subgraph are combined for content matching. For example, for the user's anxiety emotion, soothing professional content such as "diversified investment reduces risk" is selected from the knowledge subgraph and filled into the "solution" slot of the prototype. The response strategy drives the sentence conversion, and the filled knowledge unit is converted from "diversified investment, risk control" to "it is recommended that you effectively reduce the risk of a single stock by diversifying your portfolio", and the semantic framework is used to detect the logical rationality of the sentence.
[0132] The lightweight rendering engine adjusts the language and visual style through the emotion mapping algorithm. For example, when detecting the user's anxiety emotion, the text reply is adjusted to a short sentence and an exclamation word is added ("Don't worry! We have a way!"), and if it is a visual interactive interface, the background is changed to a soft color tone. This mechanism avoids using negative expressions such as "heavy losses" and "difficult to recover" through positive emotion guidance, so that the emotional matching degree of the reply content is improved. In the intelligent customer service scene, the user's satisfaction with the reply is improved, effectively alleviating negative emotional conflicts in the interaction, achieving the dual goals of emotional resonance and knowledge transmission.
[0133] In the embodiments of the present application, first, through multi-modal data acquisition, the limitation of single text is broken through, and multi-modal interaction information such as user voice and expression is comprehensively captured, so as to understand the user's intention from multiple dimensions and lay a data foundation for subsequent processing. Then, a lightweight Transformer fusion network is used to fuse multi-modal features through cross-modal attention and dynamic weight mechanism, solve the problem of data heterogeneity, reduce the load of edge computing while improving the semantic richness of features, and provide high-quality fused features for subsequent steps. Through joint compression of historical dialogue and fused features, the data transmission amount is reduced, the network delay is reduced, the system response speed is improved, and the key information is retained to support cloud processing. Next, the dynamic knowledge graph enhancement model mines the entity knowledge and emotion relationship in the dialogue in real time, dynamically updates the knowledge graph, forms a knowledge system that meets the user's needs, and provides accurate knowledge support for the construction of the reply prototype. Based on the updated knowledge graph and fused features, an initial reply prototype containing core semantics is constructed through natural language generation technology, improving the logicality and knowledge of the reply. The dialogue strategy model determines the response strategy and knowledge calling direction in combination with the user state and knowledge graph information, realizes intelligent dialogue process control, and improves the relevance of the reply and the smoothness of the interaction. Finally, the edge terminal integrates the initial prototype, the response strategy, etc., optimizes the generation of natural language replies with knowledge richness and emotional adaptability, improves the naturalness of the interaction and the user's satisfaction, and realizes the efficient and intelligent response of the system to the user's needs.
[0134] After introducing the system of the example embodiments of the present application, next, with reference to Figure 2An exemplary embodiment of the present application is described below. A dialog interaction system based on multi-modal sentiment perception and dynamic knowledge graph enhancement includes:
[0135] A user edge terminal is configured to acquire multi-modal data generated by a user during an interaction process, convert user feature vectors in different modalities into a fusion feature vector by using a lightweight Transformer fusion network through cross-modal attention fusion and dynamic weight adjustment, and compress and upload the fusion feature vector and historical dialog within a preset round to a cloud in real time.
[0136] A cloud service platform is configured to input the received compressed data into a dynamic knowledge graph enhancement model (DKGE), mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialog in real time during the dialog interaction process to update a dynamic knowledge graph, perform knowledge enhancement processing based on the updated dynamic knowledge graph, construct an initial reply prototype for the current round of interaction with the user, and deliver the initial reply prototype to the user edge terminal.
[0137] The cloud service platform is further configured to input the fusion feature vector and the updated dynamic knowledge graph into a dialog strategy model, determine a response strategy and a knowledge calling direction for the current round of dialog, and deliver the response strategy and the knowledge calling direction to the user edge terminal.
[0138] The user edge terminal is further configured to generate real-time interaction reply information output to the user based on the initial reply prototype, the response strategy, and the knowledge calling direction.
[0139] The above system can implement each step described in the above embodiments, and the specific implementation of each step will not be repeated here.
[0140] After introducing the system of the exemplary embodiments of the present application, next, a terminal device of the exemplary embodiments of the present application is described, which can be implemented as a user edge terminal. The user edge terminal is configured to acquire multi-modal data generated by a user in an interaction process; convert a user feature vector in different modalities into a fusion feature vector through cross-modal attention fusion and dynamic weight adjustment by using a lightweight Transformer fusion network; and compress and upload the fusion feature vector and historical dialogues in a preset round to the cloud in real time. Thus, the compressed data received by the cloud service platform is input into a dynamic knowledge graph enhancement model DKGE, the entity knowledge and emotional relationship implied in the fusion feature vector and the historical dialogues are mined in real time during the dialog interaction process to update the dynamic knowledge graph, and knowledge enhancement processing is performed based on the updated dynamic knowledge graph to construct an initial reply prototype for the current round of interaction with the user and deliver it to the user edge terminal; the fusion feature vector and the updated dynamic knowledge graph are input into a dialog strategy model to determine the response strategy and knowledge calling direction of the current round of dialog and deliver them to the user edge terminal. Further, the terminal device is further configured to generate real-time interaction reply information output to the user according to the initial reply prototype, the response strategy, and the knowledge calling direction. The above system can implement the steps described in the above embodiments, and the specific implementation of each step is not repeated here.
[0141] After introducing the system and the terminal device of the exemplary embodiments of the present application, next, reference is made to Figure 3 The computer readable storage medium of the exemplary embodiments of the present application is described, please refer to Figure 3 The computer readable storage medium shown is an optical disc 30, and a computer program (i.e. program product) is stored on the optical disc 30, which, when executed by a processor, will implement the steps described in the above embodiments. The specific implementation of each step is not repeated here.
[0142] It should be noted that examples of the computer-readable storage medium can also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical, magnetic storage medium, and the like, which are not listed one by one here. The above-described embodiments are only specific embodiments of the present application, used to illustrate the technical solutions of the present application, rather than limit them. The protection scope of the present application is not limited to this, although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily think of changes to the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A dialog interaction system based on multi-modal sentiment perception and knowledge graph dynamic enhancement, characterized in that, The system at least includes a user edge terminal and a cloud service platform, and the system comprises: Through the user edge terminal, multi-modal data generated by the user in the interaction process is acquired; Through the user edge terminal, a lightweight Transformer fusion network is used to convert the user feature vectors in different modalities into a fusion feature vector through cross-modal attention fusion and dynamic weight adjustment; the fusion feature vector is compressed in real time with the historical dialogue in the preset round and uploaded to the cloud; Through the cloud service platform, the received compressed data is input into a dynamic knowledge graph enhancement model DKGE, and the entity knowledge and emotional relationship implied in the fusion feature vector and the historical dialogue are mined in real time in the dialogue interaction process to update the dynamic knowledge graph, and knowledge enhancement processing is performed based on the updated dynamic knowledge graph, and an initial reply prototype for the current round of interaction with the user is constructed and delivered to the user edge terminal; The dynamic knowledge graph at least includes an emotional expression view; through the cloud service platform, the received compressed data is input into the dynamic knowledge graph enhancement model DKGE, and the entity knowledge and emotional relationship implied in the fusion feature vector and the historical dialogue are mined in real time in the dialogue interaction process to update the dynamic knowledge graph, which comprises: decompressing the compressed data through a decompression module to restore the context summary of the fusion feature vector and the historical dialogue; extracting the user emotional features of the current round of dialogue, the user feature vectors in different modalities, and the confidence of different modalities from the fusion feature vector through an extraction module; combining the emotional expression view of the user stored in the cloud through a dialogue history reconstruction module to supplement the missing information in the context summary to expand the context summary to complete historical dialogue context information; based on the user feature vectors in different modalities and the historical dialogue context information, constructing entity knowledge nodes matched with the current round of dialogue and semantic association relationships between the entity knowledge nodes through a knowledge extraction module; combining the user emotional features and the confidence of different modalities, and based on the message passing mechanism of the graph neural network GAT, the newly added entity knowledge nodes are fused into the emotional expression view of the user stored in the cloud to obtain the dialogue subgraph of the current round of dialogue through a fusion module; wherein the dialogue subgraph of the current round of dialogue at least includes all entity knowledge nodes corresponding to the current round of dialogue and semantic association relationships between all entity knowledge nodes; Through the cloud service platform, the fusion feature vector and the updated dynamic knowledge graph are input into a dialogue strategy model to determine the response strategy and knowledge calling direction of the current round of dialogue, and are delivered to the user edge terminal; Through the user edge terminal, real-time interaction reply information output to the user is generated according to the initial reply prototype, the response strategy and the knowledge calling direction.
2. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The multi-modal data at least includes: voice data, video data, text data, and biological signal data collected by wearable devices; The lightweight Transformer fusion network at least comprises: a feature extraction layer, a space-time alignment layer, a fusion layer, and an adaptive weight distribution layer. The user edge terminal adopts the lightweight Transformer fusion network to convert the user feature vectors in different modalities into fusion feature vectors through cross-modal attention fusion and dynamic weight adjustment, including: In the feature extraction layer, sound emotion feature extraction modules, speech tone feature extraction modules, micro-expression feature recognition modules, text sentiment semantic analysis modules, and physiological change feature recognition modules are constructed for different modalities of data; and user feature vectors in different modalities are extracted from the multi-modal data through the constructed different modal processing modules; wherein the user feature vectors at least include: user sound emotion features, speech tone features, micro-expression features, text sentiment semantic features, and physiological change features. Through the space-time alignment layer, the user feature vectors are projected into a unified dimensional space through a full connection layer, and the user feature vectors in other modalities are positioned to each sentiment semantic node with the sentiment semantic nodes in the text sentiment semantic features as the time axis reference, to construct a sentiment expression view consistent in space-time. Through the fusion layer, the different user feature vectors are used to perform cross-modal retrieval on the other user feature vectors in the time sequence of the time axis reference according to the sentiment semantic nodes in the text sentiment semantic features as the time axis reference, to obtain corresponding matching feature vectors, and the fusion feature vectors in different branches are obtained based on the cross-modal retrieval results. Through the adaptive weight distribution layer, the fusion feature vectors in different branches are subjected to signal quality evaluation and semantic consistency detection, and the weight parameters corresponding to different modalities are dynamically adjusted based on the evaluation results and detection results, to obtain the final output fusion feature vector, so as to eliminate data conflicts in the fusion feature vector and improve the accuracy of the output result.
3. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 2, characterized in that, The user feature vectors in different modalities are extracted from the multi-modal data through the constructed different modal processing modules, including: The sound emotion feature extraction module is used to extract user sound emotion features from speech data based on a Mel spectrum and a CNN attention mechanism; The speech tone feature extraction module is used to extract user speech tone features from speech data based on a Transformer and a BiLSTM hybrid model; The micro-expression feature recognition module is used to extract user micro-expression features from video data based on a Vision Transformer and a space-time convolution network; The text sentiment semantic analysis module is used to extract user text sentiment semantic features from text data based on a BERT and an attention mechanism; The physiological change feature recognition module is used to extract user physiological change features from the biological signal data using an LSTM model.
4. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The dialogue history reconstruction module is used to supplement missing information in the context summary by combining the cloud-stored user sentiment expression view, to expand the context summary into complete historical dialogue context information, including: retrieve a corresponding emotional semantic node from the user's emotional expression view indexed by the context summary; query a corresponding historical user feature vector in other modalities by using the retrieved emotional semantic node; predict corresponding historical dialogue information according to the emotional development trend based on the queried historical user feature vector, and merge and construct the predicted historical dialogue information into the historical dialogue context information according to the order of the emotional semantic nodes in the time axis.
5. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The knowledge extraction module constructs entity knowledge nodes matched with the current round of dialogue and semantic association relationships between the entity knowledge nodes based on the user feature vectors in different modalities and the historical dialogue context information, including: extract key named entities matched with the current round of dialogue and the dialogue topic from the user feature vectors in different modalities, and extract semantic association relationships between the key named entities in combination with the historical dialogue context information; establish entity knowledge nodes corresponding to the key named entities and semantic association relationships between the entity knowledge nodes based on the extraction results; The fusion module combines the user emotional features and the confidence of different modalities, and fuses the newly added entity knowledge nodes into the user's emotional expression view stored in the cloud based on the message passing mechanism of the graph neural network GAT, including: add corresponding emotional labels in the semantic association relationships between the newly added entity knowledge nodes based on the user emotional features and the confidence of different modalities; wherein the emotional labels are used to represent the emotional state types corresponding to the connected entity knowledge nodes; form an emotional reasoning chain containing the newly added entity knowledge nodes and semantic association relationships according to the propagation trend of the subgraph expanded by the emotional trend; adopt a dynamic pruning mechanism to delete emotional reasoning chains with overall confidence lower than a set threshold from the emotional expression view; inherit the node states of the historical dialogue in the emotional expression view across dialogue rounds and pass them to the newly added entity knowledge nodes and semantic association relationships.
6. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The dynamic knowledge graph further includes an expert knowledge base and various domain knowledge graphs associated with each entity knowledge node in the emotional expression view. The knowledge enhancement processing based on the updated dynamic knowledge graph constructs an initial reply prototype for the user's current round of interaction, including: extract the user's current interaction intent information in the current dialogue round from the dialogue subgraph of the current round of dialogue; extract the user's historical interaction intent information in the previous historical round of dialogue from the dialogue subgraph corresponding to the historical dialogue in the emotional expression view; perform intent analysis on the current interaction intent information and the historical interaction intent information to obtain a semantic framework representing the user's interaction intent; use each entity knowledge node in the dialogue subgraph of the current round of dialogue as an index to retrieve corresponding entity knowledge information from the expert knowledge base and / or various domain knowledge graphs, and generate a knowledge filling unit based on the retrieved entity knowledge information; the knowledge filling unit at least includes a dialogue content template constructed based on the entity knowledge information and a to-be-filled slot; the to-be-filled slot is supplemented in the user's edge terminal to improve the cloud-edge transmission efficiency; Adopting an emotion-knowledge mapping mechanism, taking the emotion labels between each entity knowledge node in the dialogue subgraph of the current round of dialogue as an index, querying the emotion change type matching the emotion development trend in the emotion correlation matrix corresponding to the current round of dialogue as the reply emotion orientation of the current round of dialogue; Based on the semantic framework, knowledge filling unit and reply emotion orientation determined in the foregoing steps, an initial reply prototype is constructed.
7. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The cloud service platform inputs the fusion feature vector and the updated dynamic knowledge graph into the dialogue strategy model to determine the response strategy and knowledge calling direction of the user's current round of dialogue, including: The fusion feature vector and the dialogue subgraph of the current round of dialogue are input into the dialogue strategy model; Through the dialogue strategy model, the corresponding emotion state features are extracted from the fusion feature vector according to the emotion type, emotion intensity and emotion mixed component three dimensions; Through the cross-modal cross-validation mechanism, the dominant emotion type of the user in the current round of dialogue is determined based on the emotion state features under the emotion type and emotion mixed component; Based on the preset response strategy library cooperatively mapped by the dominant emotion type and the emotion intensity, a response strategy matching the current round of dialogue is selected; The knowledge type to which each entity knowledge node in the dialogue subgraph of the current round of dialogue belongs is identified. Based on the proportion of the knowledge type occupied by the entity knowledge node, the knowledge base type matching the current round of dialogue is determined as the knowledge calling direction of the current round of dialogue.
8. The dialog interaction system based on multimodal sentiment perception and knowledge graph dynamic enhancement according to claim 1, characterized in that, The user edge terminal generates real-time interactive reply information output to the user according to the initial reply prototype, the response strategy and the knowledge calling direction, including: Based on the knowledge calling direction, a locally pre-stored knowledge subgraph is loaded; wherein, according to a preset strategy, it is monitored whether the knowledge calling direction of the current round of dialogue is consistent with the knowledge calling direction of the historical dialogue; if not, the matching knowledge base is called in advance, and the corresponding knowledge subgraph is segmented from the matching knowledge base; According to the knowledge filling unit and the reply emotion orientation in the initial reply prototype, the matching entity knowledge information is loaded from the knowledge subgraph and filled into the to-be-filled slot of the knowledge filling unit; Based on the response strategy, the filled knowledge filling unit is converted into a complete sentence, and the complete sentence is reasonably detected based on the semantic framework in the initial reply prototype to obtain first real-time interactive reply information; Through a lightweight rendering engine, the language style and / or visual style in the first real-time interactive reply information are dynamically adjusted according to the reply emotion orientation in the initial reply prototype to obtain the second real-time interactive reply information finally output, so as to avoid that the second real-time interactive reply information contains negative target points of the emotion change type.
9. A dialog interaction system based on multi-modal sentiment perception and knowledge graph dynamic enhancement, characterized in that, The system comprises: The user edge terminal is used to acquire multi-modal data generated by the user in the interaction process; a lightweight Transformer fusion network is adopted to convert the user feature vectors in different modalities into a fusion feature vector through cross-modal attention fusion and dynamic weight adjustment; and the fusion feature vector and historical dialogues in a preset round are compressed in real time and uploaded to the cloud; The cloud service platform is used to input the received compressed data into a dynamic knowledge graph enhancement model DKGE, to mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogues in real time during the dialog interaction process, to update the dynamic knowledge graph, and to perform knowledge enhancement processing based on the updated dynamic knowledge graph, to build an initial reply prototype for the current round of interaction with the user, and to deliver it to the user edge terminal; The dynamic knowledge graph at least includes an emotional expression view; when the cloud service platform inputs the received compressed data into the dynamic knowledge graph enhancement model DKGE to mine entity knowledge and emotional relationships implied in the fusion feature vector and the historical dialogues in real time during the dialog interaction process to update the dynamic knowledge graph, it is specifically used for: decompressing the compressed data through a decompression module to restore the fusion feature vector and the context summary of the historical dialogues; extracting user emotional features of the current round of dialogues, user feature vectors in different modalities, and confidence degrees of different modalities from the fusion feature vector through an extraction module; supplementing missing information in the context summary to expand the context summary into complete historical dialogue context information through a dialog history reconstruction module in combination with the emotional expression view of the user stored in the cloud; constructing entity knowledge nodes matched with the current round of dialogues and semantic association relationships between the entity knowledge nodes based on the user feature vectors in different modalities and the historical dialogue context information through a knowledge extraction module; combining the user emotional features and the confidence degrees of different modalities, and based on the message passing mechanism of the graph neural network GAT, the fusion module fuses the newly added entity knowledge nodes into the emotional expression view of the user stored in the cloud to obtain a dialog subgraph of the current round of dialogues; wherein the dialog subgraph of the current round of dialogues at least includes all entity knowledge nodes corresponding to the current round of dialogues and semantic association relationships between all entity knowledge nodes; The cloud service platform is also used to input the fusion feature vector and the updated dynamic knowledge graph into a dialog strategy model to determine the response strategy and knowledge calling direction of the current round of dialogues, and to deliver them to the user edge terminal; The user edge terminal is also used to generate real-time interaction reply information output to the user according to the initial reply prototype, the response strategy, and the knowledge calling direction.
Citation Information
Patent Citations
Intelligent real-time emotion evaluation method for social media and online text data based on multi-modal knowledge graph
CN119202270A
Generating model output using a knowledge graph
US20240428787A1