Intelligent customer service automatic reply generation method and system based on multi-modal learning

The intelligent customer service automatic response generation method based on multimodal learning solves the problem of insufficient multimodal feature fusion in existing technologies, achieves accurate identification of user intent and emotions, and improves the personalization and consistency of intelligent responses.

CN120975248AInactive Publication Date: 2025-11-18NANTONG BEIRUISMAN INFORMATION TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511505102.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2025-11-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of deep fusion of multimodal features and dynamic context modeling capabilities in existing technologies makes it difficult to accurately identify users' true intentions and emotions in complex interactions, affecting the continuous tracking of user needs and the personalized effect of intelligent responses in multi-turn dialogues.

Method used

By using a multimodal learning-based intelligent customer service automatic response generation method, we read and encode customers' multimodal data, perform feature fusion and time alignment, establish a contextual graph structure, perform multi-label intent recognition and emotion classification, and generate natural language responses.

Benefits of technology

It improves the accuracy of intent recognition, enhances the comprehensiveness of emotion understanding, and enables personalized and contextually coherent natural language responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975248A_ABST
    Figure CN120975248A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent customer service automatic reply generation method and system based on multi-modal learning, and relates to the technical field of intelligent customer service, and the method comprises the steps: carrying out the feature coding of multi-modal data inputted by a customer, and constructing a modal feature vector; after time alignment is carried out on the modal feature vectors, feature fusion of the modal feature vectors is carried out based on a cross-modal attention mechanism, and fusion features are established; mapping the fusion feature to an entity and relationship set of the business knowledge spectrogram, establishing a situation node, and establishing a situation map structure at the situation node; performing multi-label intention recognition by using the situation map structure, executing emotion classification, and establishing a recognition result; and inputting the identification result into a strategy planner, establishing a business strategy template according to the customer characteristics and the identification result, and outputting a natural language reply. The technical problem of customer service intelligent reply efficiency in the prior art can be solved, and the technical effect of improving the customer service intelligent reply efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent customer service technology, and in particular to an intelligent customer service automatic response generation method and system based on multimodal learning. Background Technology

[0002] In the development of intelligent human-computer interaction and intelligent customer service, significant progress has indeed been made in understanding user intent and providing responses, thanks to the continuous improvement of related technologies such as multimodal data processing, contextual understanding, and natural language generation.

[0003] Currently, existing technologies still have significant shortcomings. In terms of intent recognition and emotion recognition, most are single-modal processing methods, often relying solely on text content and lacking comprehensive analysis of speech prosody, non-verbal information in images and videos, and structured business data, leading to biases in understanding the user's true context. Secondly, existing multi-turn dialogue management methods generally lack dynamic update mechanisms, typically matching only within limited contexts and failing to form a complete contextual graph, causing the system to easily lose the user's core needs during long dialogues.

[0004] In summary, existing technologies suffer from a lack of deep fusion of multimodal features and dynamic context modeling capabilities, which makes it difficult to accurately identify users' true intentions and emotions in complex interactions. This further affects the continuous tracking of user needs and the personalized effect of intelligent responses in multi-turn dialogues. Summary of the Invention

[0005] The purpose of this application is to provide a method and system for generating intelligent customer service automatic responses based on multimodal learning, in order to solve the technical problem in the prior art that the lack of deep fusion of multimodal features and dynamic situation modeling capabilities makes it difficult to accurately identify the user's true intentions and emotions in complex interactions, which further affects the continuous tracking of user needs and the personalized effect of intelligent responses in multi-turn dialogues.

[0006] In view of the above problems, this application provides a method and system for generating intelligent customer service automatic replies based on multimodal learning.

[0007] Firstly, this application provides a method for generating automatic responses to intelligent customer service based on multimodal learning, implemented through an intelligent automatic response generation system for customer service based on multimodal learning. The method includes: reading multimodal data input by the customer; encoding the features of the multimodal data to construct modal feature vectors; the multimodal data including text, voice, images, videos, and structured business data; aligning the modal feature vectors temporally; fusing the modal feature vectors based on a cross-modal attention mechanism to establish fused features; mapping the fused features to an entity and relation set in a business knowledge graph to establish context nodes; and establishing a context graph structure within the context nodes; using the context graph structure to perform multi-label intent recognition and emotion classification to establish recognition results; inputting the recognition results into a strategy planner; establishing a business strategy template based on customer characteristics and the recognition results; and outputting a natural language response.

[0008] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: extracting semantic units of different modalities from the fusion features and establishing context nodes, wherein the context nodes include intent nodes, emotion nodes, event nodes, and state nodes, wherein: the intent nodes are constructed by performing semantic analysis of text features, speech-to-text features, and video-to-text features in the fusion features; analyzing the speech prosody of speech features and video features to establish a first emotion association; identifying and analyzing emotion words in text features to establish a second emotion association; using the first emotion association and the second emotion association to establish an emotion node; performing content recognition on images and videos to establish event nodes; performing business structured data analysis to establish state nodes; and establishing a context graph structure at the context nodes.

[0009] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: establishing cross-modal fusion edges for the context nodes, wherein the cross-modal fusion edges include temporal consistency edges, semantic association edges, and emotion transmission edges; the temporal consistency edges are constructed by calculating the association between any two nodes through time windows, and the weights of the temporal consistency edges are constructed through a time similarity function; the semantic association edges are constructed through cross-modal feature similarity, and the weights of the semantic association edges are constructed through the cosine similarity of the fused features; the emotion transmission edges are constructed through emotion propagation relationships, and the weights of the emotion transmission edges are constructed through an emotion intensity coefficient and a business sensitivity coefficient; a context graph structure is established using the context nodes and cross-modal fusion edges, and the context graph structure is dynamically updated in multi-turn dialogues based on changes in node semantic similarity and business status.

[0010] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: setting a node similarity merging threshold; performing node similarity calculation during each round of dialogue update; if the node similarity calculation result meets the similarity merging threshold, then performing corresponding node merging, and performing weighted fusion of nodes according to weighted reliability; configuring a node decay index for the context graph, and using the node decay index to perform node aging and elimination management within the context graph; and completing the context graph structure update using node merging and node aging and elimination management.

[0011] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: using the context graph structure as a dynamic context semantic space, using the context path matching channel to perform historical context subgraph matching search in the context graph structure to establish a candidate intent set; obtaining associated emotion tags in the context graph structure, using the associated emotion tags to establish emotion-intent constraints; using the emotion-intent constraints to filter the candidate intent set to establish a recognition result.

[0012] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: performing contextual backtracking on the recognition results to establish a contextual backtracking set; and using the contextual backtracking set to perform temporal smoothing processing on the recognition results to establish updated recognition results.

[0013] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: generating a source tracing record instruction, using the source tracing record instruction to save the modality source identifier, timestamp, and node confidence field in the context node; and performing interpretive source tracing management based on the data saved in the context node.

[0014] Preferably, the intelligent customer service automatic response generation method based on multimodal learning further includes: the strategy planner constructs strategy nodes based on the recognition results and the customer characteristics, calculates strategy scores according to the weights of the strategy nodes and emotion-intent coupling; after filtering the strategy scores, a business strategy template is constructed; and after adaptive adjustment of the natural language output on the business strategy template, a natural language response is output.

[0015] Preferably, the intelligent customer service automatic reply generation method based on multimodal learning further includes: reading the user's historical reply preference features; establishing a preference influence factor based on the historical reply preference features; and updating the natural language reply after using the preference influence factor to compensate for the natural language reply.

[0016] Secondly, this application also provides an intelligent customer service automatic response generation system based on multimodal learning, used to execute the intelligent customer service automatic response generation method based on multimodal learning as described in the first aspect, including: a modal feature vector construction module, used to read multimodal data input by the customer, perform feature encoding on the multimodal data, and construct modal feature vectors, wherein the multimodal data includes text, voice, image video, and business structured data; a fusion feature establishment module, used to perform time alignment on the modal feature vectors, and then perform feature fusion of the modal feature vectors based on a cross-modal attention mechanism to establish fusion features; a context node establishment module, used to map the fusion features to the entity and relation set of the business knowledge graph, establish context nodes, and establish a context graph structure on the context nodes; a recognition result establishment module, used to perform multi-label intent recognition using the context graph structure, and perform emotion classification to establish recognition results; and a natural language response output module, used to input the recognition results into a strategy planner, establish a business strategy template based on customer characteristics and recognition results, and then output a natural language response.

[0017] The technical solution provided in this application has at least the following technical effects or advantages: by achieving the technical goal of building an intelligent interactive system based on multimodal fusion and context graph, it achieves the technical effects of improving the accuracy of intent recognition, enhancing the comprehensiveness of emotion understanding, and realizing the personalization and contextual coherence of natural language responses.

[0018] The above description is merely an overview of the technical solution of this application. To better understand the technical means of this application and to facilitate its implementation according to the description, and to make the above and other objects, features, and advantages of this application more apparent, specific embodiments of this application are described below. It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent through the following description. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 This is a flowchart illustrating the intelligent customer service automatic response generation method based on multimodal learning proposed in this application.

[0021] Figure 2This is a schematic diagram of the structure of the intelligent customer service automatic response generation system based on multimodal learning in this application.

[0022] Figure labeling: Modal feature vector construction module 1, fusion feature establishment module 2, context node establishment module 3, recognition result establishment module 4, natural language response output module 5. Detailed Implementation

[0023] This application provides a method and system for generating intelligent customer service responses based on multimodal learning. It addresses the technical problem in existing technologies where the lack of deep fusion of multimodal features and dynamic context modeling capabilities makes it difficult to accurately identify users' true intentions and emotions in complex interactions, further impacting the continuous tracking of user needs and the personalization of intelligent responses in multi-turn dialogues. The application aims to build an intelligent interaction system driven by multimodal fusion and contextual graphs, achieving improved accuracy in intent recognition, enhanced comprehensiveness in emotion understanding, and personalized and contextually coherent natural language responses.

[0024] The technical solutions of this application will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. It should be understood that this application is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. It should also be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all of them.

[0025] Example 1, please refer to the appendix. Figure 1 This application provides a method for generating automatic responses to intelligent customer service based on multimodal learning, which is applied to an intelligent customer service automatic response generation system based on multimodal learning. The method specifically includes the following steps: S1: Read the multimodal data input by the customer, perform feature encoding on the multimodal data, and construct a modal feature vector. The multimodal data includes text, voice, image and video, and business structured data.

[0026] Specifically, after obtaining customer authorization, the system reads multimodal data input by the customer, acquiring information from different sources, such as text typed by the customer, voice recordings, uploaded images or videos, and customer-related business database data. Multimodal data refers to multiple ways of representing information within a single scenario. For example, in customer service conversations, text can directly express needs, voice recordings include tone and emotion, images or videos may show products or malfunctions, and structured business data includes customer identity information, transaction records, or work order status.

[0027] Next, feature encoding is performed on multimodal data, transforming data from different sources into a digital representation that computers can understand and process uniformly. For example, text can be transformed into a set of numerical values ​​through word vector models, speech can be transformed into a combination of frequency and energy values ​​through acoustic feature extraction, and image and video can have feature vectors of shape, color, or object category extracted through convolutional neural networks. Business structured data is directly encoded into numerical values ​​using existing table fields.

[0028] After encoding data from different modalities, modal feature vectors are constructed, representing a numerical sequence or vector that can represent the information of that modality. For example, a piece of text can be encoded into a floating-point vector of length 300, a piece of speech can be converted into an acoustic feature vector containing 200 dimensions, an image may be converted into a visual feature vector containing 512 dimensions, and business structured data may be a vector composed of dozens of numerical fields.

[0029] S2: After time alignment of the modal feature vectors, feature fusion of the modal feature vectors is performed based on the cross-modal attention mechanism to establish fused features.

[0030] Specifically, time alignment is performed on modal feature vectors to find the correspondence between data from different modalities along the time dimension. For example, if the word "damaged" appears at 10 seconds in a customer's speech segment, and the customer uploads a photo of damaged packaging at the same time, then time alignment ensures that the speech features and image features correspond to the same point in time, thus preventing information misalignment during analysis. Time alignment relies on timestamp or sequence alignment algorithms; for example, dynamic time warping methods can reasonably match speech and text information of varying lengths along the timeline.

[0031] Next, feature fusion of modal feature vectors is performed based on a cross-modal attention mechanism to determine the importance of information between different modalities and perform weighted combination. The cross-modal attention mechanism is a deep learning method that dynamically allocates weights based on contextual relationships. For example, when a customer uploads a picture of a damaged product, the weight of image features in the fusion process is automatically increased; conversely, when a customer describes the product model in text, textual features are emphasized, thus establishing connections between different modalities. After completing temporal alignment and attention weighting, a comprehensive feature representation is obtained, and a fused feature is established. This fused feature represents not only the semantic information of the text, the tone and emotion of the speech, and the visual information of the image or video, but also the background information of the business structured data.

[0032] S3: Map the fused features to the entity and relation set of the business knowledge graph, establish context nodes, and establish a context graph structure in the context nodes.

[0033] Specifically, the fused features are mapped to the entity and relation set of the business knowledge graph, transforming the unified representation after multimodal fusion into specific elements within the business knowledge graph. A business knowledge graph is a network structure containing entities and relations. Entities are objects involved in the business, such as users, products, and orders; relations are connections between entities, such as a user purchasing a product or an order containing goods. The mapping process associates fused features with entities and relations. For example, when a user says, "I want to return a piece of clothing," the semantic features are mapped to entities such as "user," "return request," and "product," connected through the relationship "user requests to return a product." The mapping results are organized into new nodes, establishing context nodes. A context graph structure is built within these context nodes, illustrating how nodes are interconnected through edges, forming a dynamic graph. The context graph is updated in real-time based on the conversation, recording changes in user intent, emotions, events, and states.

[0034] S4: Utilize the aforementioned contextual graph structure to perform multi-label intent recognition and emotion classification, and establish recognition results.

[0035] Specifically, multi-label intent recognition is performed using a contextual graph structure. This combines user input data with the contextual graph to identify multiple related intents simultaneously. Multi-label intent recognition refers to the fact that a single message may contain multiple needs; for example, a sentence might express both an intent to check balance and an intent to complain. Emotion classification is then performed to analyze the user's emotional state in their expression, such as anger, anxiety, calmness, or happiness. Emotion classification is achieved by analyzing tone of voice, emotional words in text, or facial expressions in video, helping to adjust response strategies—for example, focusing on soothing when the user is angry and on interaction when the user is happy. Finally, the results of multi-label intent recognition and emotion classification are combined to form a complete recognition output, establishing the recognition results to guide the next step of response generation.

[0036] S5: Input the recognition results into the strategy planner, and after establishing a business strategy template based on customer characteristics and recognition results, output a natural language response.

[0037] Specifically, the identification results are input into the strategy planner, which receives the results of multi-label intent recognition and emotion classification as the basis for formulating a response plan. The identification results represent a complete understanding of the user's current needs and emotions, while the strategy planner is the module responsible for translating this understanding into actionable decisions.

[0038] Based on customer characteristics and identification results, a business strategy template is created. The strategy planner integrates individual user characteristics, such as spending power, service history, and preferences, with identified intent and emotions to generate a targeted business strategy template. This template serves as a response framework, containing specific processing paths and execution steps, such as reassuring the user, providing billing results, and recommending personalized offers. Guided by the business strategy template, the decision content is translated into natural language answers that align with the user's expression habits, outputting a natural language response.

[0039] Furthermore, this application also includes: extracting semantic units of different modalities from the fusion features, establishing context nodes, wherein the context nodes include intent nodes, emotion nodes, event nodes, and state nodes, wherein: the intent nodes are constructed by performing semantic analysis of text features, speech-to-text features, and video-to-text features in the fusion features; analyzing the speech prosody of speech features and video features to establish a first emotion association; identifying and analyzing emotion words in text features to establish a second emotion association, and using the first emotion association and the second emotion association to establish an emotion node; performing content recognition on images and videos to establish event nodes; performing business structured data analysis to establish state nodes; and establishing a context graph structure in the context nodes.

[0040] Specifically, semantic units of different modalities are extracted from the fused features. Meaningful information fragments are then separated from the fused feature vector. These information fragments represent the specific semantics expressed by the user, such as keywords in text, intonation features in speech, object categories in images, and field values ​​in business data. Semantic units decompose complex multimodal vectors into finer-grained elements, facilitating the further construction of contextual nodes.

[0041] Different types of nodes are created based on semantic units, with each node representing a contextual element. Contextual nodes mainly include four categories: intent nodes, emotion nodes, event nodes, and state nodes. Together, these contextual nodes form a complete framework for expressing user questions, mapping raw multimodal information onto a structured knowledge graph.

[0042] Intent nodes are constructed through semantic analysis of fused features, including text features, speech-to-text features, and video-to-text features. This process converts speech and video content into text, which is then combined with the original text features for semantic analysis to extract the user's true needs or purposes. For example, when a customer says "I want to return the goods" and a damaged product is shown in the video, the intent node would be constructed as "return request".

[0043] By analyzing speech and video features, including prosodic characteristics such as pitch, speed, and pauses in speech, and matching facial expressions or lip movements in video with the speech, we can determine the user's emotional tendency and establish a primary emotional association. For example, when a user speaks quickly and in a high-pitched voice, we can establish an emotional association of "anxiety."

[0044] The analysis of emotion words in text features involves identifying words with emotional connotations, such as "very angry" and "very satisfied," to establish a second emotion association. This association is then combined with prosodic results from audio and video to form more reliable emotion nodes. For example, if the audio shows anxiety and the text contains "angry," the emotion node would be labeled as "angry."

[0045] Content recognition is performed on images and videos to establish event nodes. Events that occur can be identified from visual modalities. For example, if a damaged express box is identified in an image or a falling action is identified in a video, an event node "defective express box" can be established.

[0046] Perform structured data analysis on business operations and establish status nodes. These nodes represent the status information generated using customer data or order data in the database. For example, if a customer is found to have an item marked "in transit" in their order, then the status node "Order in transit" will be generated.

[0047] A context graph structure is built at the context nodes, connecting the relationships between intent nodes, emotion nodes, event nodes, and state nodes through edges to form a context graph. For example, an intent node of "return request" can be connected to an emotion node of "anger," an event node of "damaged package," and a state node of "order in transit," thus forming a complete contextual semantic representation.

[0048] Furthermore, this application also includes: establishing cross-modal fusion edges for the context nodes, wherein the cross-modal fusion edges include temporal consistency edges, semantic association edges, and emotion transmission edges; the temporal consistency edges are constructed by calculating the temporal window association between any two nodes, and the weights of the temporal consistency edges are constructed using a temporal similarity function; the semantic association edges are constructed using cross-modal feature similarity, and the weights of the semantic association edges are constructed using the cosine similarity of the fusion features; the emotion transmission edges are constructed using emotion propagation relationships, and the weights of the emotion transmission edges are constructed using an emotion intensity coefficient and a business sensitivity coefficient; a context graph structure is established using the context nodes and cross-modal fusion edges, and the context graph structure is dynamically updated in multi-turn dialogues based on changes in node semantic similarity and business status.

[0049] Specifically, establishing cross-modal fusion edges for context nodes can organize information from different modalities into a network structure. Cross-modal fusion edges include three types: temporally consistent edges, semantically related edges, and sentiment-transmitting edges. Each type of edge represents a relationship between nodes, thus making context nodes no longer isolated points, but an interconnected whole.

[0050] Temporally consistent edges are constructed by calculating the time window association between any two nodes. This determines whether the two nodes are close or overlap in time. If they appear within similar time periods, a temporally consistent edge is established to indicate that the two nodes may belong to the same context. For example, if a customer uploads a damaged photo at 5 seconds and says "This is broken" at 6 seconds, a temporally consistent edge is established between these two nodes. The weight of the temporally consistent edge is constructed using a temporal similarity function; that is, similarity is calculated based on the time difference. The smaller the time difference, the higher the similarity, and the larger the weight.

[0051] Semantic association edges are constructed using cross-modal feature similarity. Nodes from different modalities are linked when they are semantically similar. For example, if the text contains "return goods" and the speech recognition content also contains "want to return goods," then there is a strong semantic association between the two. The weights of semantic association edges are constructed using fused feature cosine similarity. Cosine similarity is a similarity metric with a value between 0 and 1. The closer the value is to 1, the closer the semantics are, and therefore the higher the edge weight.

[0052] Emotional transmission edges are constructed based on emotion propagation relationships, representing how a user's emotions may spread between different nodes. For example, if a user's voice expresses anger while their text uses phrases like "That's terrible," a transmission edge will be established between the emotion-related nodes. The weight of the emotional transmission edge is constructed using an emotion intensity coefficient and a business sensitivity coefficient. The emotion intensity coefficient represents the strength of the user's emotion, and the business sensitivity coefficient represents the sensitivity of the business type to emotions. For example, the weight of the emotional transmission edge is obtained by summing and averaging the emotion intensity coefficient and the business sensitivity coefficient.

[0053] A context graph structure is established by utilizing context nodes and cross-modal fusion edges. The context graph structure is dynamically updated in multi-turn dialogues based on node semantic similarity and business state changes. A context graph is constructed at the initial dialogue and continuously updated in subsequent dialogues.

[0054] Furthermore, this application also includes: setting a similarity merging threshold for nodes; performing node similarity calculation during each round of dialogue update; if the similarity calculation result of a node meets the similarity merging threshold, then performing corresponding node merging, and performing weighted fusion of nodes according to weighted reliability; configuring a node decay index for the context graph, and using the node decay index to perform node aging and elimination management within the context graph; and completing the context graph structure update using node merging and node aging and elimination management.

[0055] Specifically, a similarity merging threshold is set for nodes. The similarity merging threshold is a preset value, such as 0.8, which means that if the semantic similarity between two nodes exceeds 80%, the two nodes can be considered to express the same concept.

[0056] During each round of dialogue updates, the similarity of nodes is calculated. The similarity calculation is based on the semantic similarity function or the distance metric between feature vectors, such as cosine similarity or Euclidean distance. The similarity between the new node and the existing nodes is calculated every time the user speaks or inputs new data.

[0057] If the similarity calculation results of nodes meet the similarity merging threshold, it means that the two nodes are highly similar. Therefore, the corresponding nodes are merged, and a weighted fusion of nodes is performed according to their confidence levels to merge them into a single node, avoiding duplicate information in the graph. Weighted confidence level refers to the importance or credibility of a node itself. For example, a node supported by multimodal information might have a confidence level of 0.9, while a node constructed from only a single modality might have a confidence level of 0.6. The weighted fusion process involves weighting the node features according to their confidence levels to retain more reliable information. For example, a return intention node supported by both speech and text has a higher confidence level and will occupy a larger proportion during fusion.

[0058] Configure a node decay index for the context graph. This index is used to manage node aging and elimination within the context graph, thereby controlling the decline in node importance over time. The node decay index can be a value between 0 and 1, such as 0.95, which means the node's weight decreases by 5% in each round of dialogue. If a node is not used for a long time and decays below a certain elimination threshold, such as 0.2, it will be deleted, ensuring the context graph does not expand indefinitely while still reflecting the latest dialogue status. For example, if a user asked "When will the package arrive?" three days ago, and hasn't mentioned it again, the node's weight will gradually decrease until it is eliminated.

[0059] The context graph structure is updated using node merging and node aging / elimination management. By merging similar nodes and deleting outdated nodes, the context graph remains concise and efficient during dialogue. Updates not only reduce redundant information but also ensure that the context graph always reflects the most relevant context. As dialogue rounds increase, the number of nodes in the context graph may decrease from the initial 10 to 8 due to node merging, or it may grow back to 12 as new nodes are added, ensuring the flexibility and real-time performance of the context graph.

[0060] Furthermore, this application also includes: using the context graph structure as a dynamic context semantic space, using the context path matching channel to perform historical context subgraph matching search in the context graph structure to establish a candidate intent set; obtaining associated emotion tags in the context graph structure, using the associated emotion tags to establish emotion-intent constraints; using the emotion-intent constraints to filter the candidate intent set to establish a recognition result.

[0061] Specifically, the context graph structure is viewed as a dynamic contextual semantic space, a semantic background that is constantly updated as the dialogue changes. All nodes and edges record the semantic relationships of the user's intentions, emotions, events, and states. The dynamic context emphasizes that it is not fixed but is adjusted according to new input information. For example, if the user previously said "I want to return the item," there is already a "return" intention node in the context graph. If the user later says "because of quality issues," then a "quality issues" node will be added to the contextual semantic space and connected to the return intention.

[0062] Utilizing contextual path matching channels within the contextual graph structure, historical contextual subgraphs are matched and searched, meaning similar contexts that have appeared in past conversations are sought within the graph. A contextual path refers to a semantic link from one node to another in the graph, such as "user → order → return request". Subgraph matching searches for whether such a link has appeared before, inferring the user's possible intent. If the path "user returns goods → quality issue" has been seen before, the possibility of "return due to quality issue" will be added to the candidate intent set when the user mentions a similar context again.

[0063] Obtaining associated emotion tags from the context graph structure refers to extracting emotional information related to user dialogue, such as anger, anxiety, and satisfaction, from the context graph. Emotion tags originate from speech prosody, text emotion words, or video facial expression recognition, providing additional clues to understanding the user's true needs. The relevance of emotion tags is reflected in the connection between emotions and nodes such as intentions and events. For example, if a return request is accompanied by "anger," it means the user has a low tolerance for the problem.

[0064] By using associated emotion tags to establish emotion-intention constraints, the connection between emotion and intention is set as a filtering condition. For example, if a user expresses "anger," the filter is more likely to identify intentions related to complaints or refunds, rather than recommendations or consultations. Emotion-intention constraints ensure that intention recognition results are more context-aware and avoid misjudgments based solely on semantics.

[0065] Using emotion-intention constraints to filter candidate intent sets and establish recognition results means that within an existing candidate intent set, the coupling relationship between emotion and intention is used for filtering and narrowing, ultimately retaining only the intents that best fit the context and emotional state. For example, if the candidate set contains "return," "exchange," and "logistics inquiry," but considering the emotion of anger, the final recognition result may only include "return."

[0066] Furthermore, this application also includes: performing contextual backtracking on the recognition results to establish a contextual backtracking set; and using the contextual backtracking set to perform temporal smoothing processing on the recognition results to establish updated recognition results.

[0067] Specifically, contextual backtracking is performed on the recognition results to establish a contextual backtracking set. This means that after obtaining a recognition result from a certain round, the previous dialogue content is traced back to find the historical context related to the current result, and these historical fragments are collected into a set. Contextual backtracking combines relevant parts of the dialogue chain. For example, if a user says "I want to return the product" in the first round and adds "because of quality issues" in the third round, then the contextual backtracking set will contain two pieces of content, thus taking into account the correlation between the preceding and following parts during recognition.

[0068] Temporal smoothing is applied to the recognition results using contextual backtracking sets to establish and update the results. This involves comparing and fusing the current recognized intent and emotion with the historical contextual backtracking sets, eliminating sudden biases through the continuity of the time series. Temporal smoothing is a dynamic adjustment method that prevents the recognition results from drastically changing due to a temporary, ambiguous expression. For example, if a user has consistently expressed the intent to "return the product" in previous rounds, but only says "the product is bad" in the latest round, temporal smoothing will combine the current round's result with the previous rounds, ultimately stabilizing it as "return the product" rather than arbitrarily shifting to other intents.

[0069] Furthermore, this application also includes: generating a traceability record instruction, using the traceability record instruction to save the modality source identifier, timestamp, and node confidence field at the context node; and performing interpretive traceability management based on the data saved at the context node.

[0070] Specifically, a source tracing record instruction is generated, with additional record information added to track the origin and reliability of nodes. The source tracing record instruction stores a modality source identifier, timestamp, and node confidence field at the context node. The modality source identifier refers to the modality source of the node information, such as text, voice, or image. The timestamp refers to the specific time when the information was captured or generated. The node confidence field is a numerical evaluation of the system's accuracy for that node, ranging from 0 to 1; for example, 0.9 indicates very reliable.

[0071] Explanatory attribution management, based on data stored in context nodes, allows for tracing how each result was arrived at, thereby achieving explanatory power. The purpose of explanatory attribution management is to make the decision-making process more transparent; users or developers can judge the reasonableness of a response by viewing the modality source, time, and confidence level of a node.

[0072] Furthermore, this application also includes: the strategy planner constructs strategy nodes based on the recognition results and the customer characteristics, calculates strategy scores according to the weights of the strategy nodes and the emotion-intent coupling; after filtering the strategy scores, a business strategy template is constructed; and after adaptively adjusting the natural language output on the business strategy template, a natural language response is output.

[0073] Specifically, the strategy planner constructs strategy nodes based on the recognition results and customer characteristics. That is, it generates strategy nodes based on the intent and emotion recognition results, combined with customer characteristics such as age, consumption habits, and historical behavior records. Strategy nodes are decision points for responding to the needs of a certain type of customer. For example, when customer characteristics indicate a high-value user, the node may be biased towards providing priority services.

[0074] A strategy score is calculated by coupling the weights of strategy nodes with emotional intent. This means that each strategy node has a weight indicating its importance, while also combining intent and emotion. For example, if a customer intends to check their balance but is also angry, a strategy score is calculated for this situation. A higher score indicates that the task requires priority. This combination of weights and coupling ensures that the strategy score considers not only business value but also the urgency of the customer's emotions.

[0075] After filtering the strategy scores, a business strategy template is constructed. The highest-scoring or most suitable subset of all strategy nodes is selected to form a processing solution. The business strategy template is an executable response framework; for example, a template might include steps such as calming the user, providing billing details, and recommending promotional activities.

[0076] After adaptively adjusting the natural language output on the business strategy template, a natural language response is output. This means fine-tuning the expression based on the user's language habits and the context; for example, the response is more conversational when addressing younger users and more formal when addressing older users. This adaptive adjustment of the natural language output ensures that the response is both business-logical and human-centered, and the final natural language response is the answer presented to the user.

[0077] Furthermore, this application also includes: reading the user's historical response preference characteristics; establishing a preference influence factor based on the historical response preference characteristics; and updating the natural language response after using the preference influence factor to compensate for the natural language response.

[0078] Specifically, this involves reading users' historical response preferences to extract communication habits from their past interactions, such as whether they prefer brief or detailed answers, whether they prefer a formal or relaxed tone, and frequently used keywords or sentence structures. These historical response preferences essentially form a user's language usage profile, helping the system better adapt to personalized needs.

[0079] Preference influence factors are established based on historical response preference characteristics. These characteristics are transformed into calculable factors to guide the style and structure of generated responses. For example, a user's preference factor might increase the weight of conciseness or the weight of colloquial expressions, making the response more in line with user habits.

[0080] After compensating for the initial response using preference influence factors, the natural language response is updated. This involves fine-tuning the generated natural language response based on preference influence factors. For example, if the original response is a detailed explanation, but the user historically prefers shorter answers, redundant parts are automatically removed to make the response more in line with expectations. The process of updating the natural language response is equivalent to personalized optimization of the standard answer.

[0081] In summary, the intelligent customer service automatic response generation method based on multimodal learning provided in this application has the following technical effects: by achieving the technical goal of constructing an intelligent interactive system based on multimodal fusion and context graph driving, it achieves the technical effects of improving the accuracy of intent recognition, enhancing the comprehensiveness of emotion understanding, and realizing the personalization and contextual coherence of natural language responses.

[0082] Example 2: Based on the same inventive concept as the intelligent customer service automatic response generation method based on multimodal learning in the foregoing examples, this application also provides an intelligent customer service automatic response generation system based on multimodal learning. Please refer to the appendix. Figure 2The system includes: a modal feature vector construction module 1, used to read multimodal data input by the customer, perform feature encoding on the multimodal data, and construct modal feature vectors, wherein the multimodal data includes text, speech, images, videos, and business structured data; a fusion feature establishment module 2, used to perform time alignment on the modal feature vectors, and then perform feature fusion based on a cross-modal attention mechanism to establish fused features; a context node establishment module 3, used to map the fused features to the entity and relation set of the business knowledge graph, establish context nodes, and establish a context graph structure on the context nodes; a recognition result establishment module 4, used to perform multi-label intent recognition using the context graph structure, perform emotion classification, and establish recognition results; and a natural language response output module 5, used to input the recognition results into the strategy planner, establish a business strategy template based on customer characteristics and recognition results, and then output a natural language response.

[0083] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used to: extract semantic units of different modalities from the fusion features, and establish context nodes, wherein the context nodes include intent nodes, emotion nodes, event nodes, and state nodes, wherein: the intent nodes are constructed by semantic analysis of text features, speech-to-text features, and video-to-text features in the fusion features; analyze the speech prosody of speech features and video features to establish a first emotion association; identify and analyze emotion words in text features to establish a second emotion association, and establish emotion nodes using the first emotion association and the second emotion association; perform content recognition on images and videos to establish event nodes; perform business structured data analysis to establish state nodes; and establish a context graph structure at the context nodes.

[0084] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used to: establish cross-modal fusion edges for the context nodes, wherein the cross-modal fusion edges include temporal consistency edges, semantic association edges, and emotion transmission edges; the temporal consistency edges are constructed by calculating the association between any two nodes through time windows, and the weights of the temporal consistency edges are constructed through a time similarity function; the semantic association edges are constructed through cross-modal feature similarity, and the weights of the semantic association edges are constructed through the cosine similarity of the fusion features; the emotion transmission edges are constructed through emotion propagation relationships, and the weights of the emotion transmission edges are constructed through emotion intensity coefficients and business sensitivity coefficients; a context graph structure is established using the context nodes and cross-modal fusion edges, and the context graph structure is dynamically updated in multi-turn dialogues based on changes in node semantic similarity and business status.

[0085] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used for: setting a similarity merging threshold for nodes; performing node similarity calculation during each round of dialogue update; if the similarity calculation result of a node meets the similarity merging threshold, then performing corresponding node merging, and performing weighted fusion of nodes according to weighted reliability; configuring a node decay index for the context graph, and using the node decay index to perform node aging and elimination management within the context graph; and completing the context graph structure update using node merging and node aging and elimination management.

[0086] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used to: use the context graph structure as a dynamic context semantic space, use the context path matching channel to perform historical context subgraph matching search in the context graph structure to establish a candidate intent set; obtain the associated emotion tags in the context graph structure, use the associated emotion tags to establish emotion-intent constraints; use the emotion-intent constraints to filter the candidate intent set and establish recognition results.

[0087] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used to: perform contextual backtracking on the recognition results to establish a contextual backtracking set; and use the contextual backtracking set to perform temporal smoothing processing on the recognition results to establish updated recognition results.

[0088] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used to: generate traceability record instructions, and use the traceability record instructions to save the modality source identifier, timestamp, and node confidence field in the context node; and perform interpretive traceability management based on the data saved in the context node.

[0089] Furthermore, the intelligent customer service automatic response generation system based on multimodal learning is also used for: the strategy planner constructing strategy nodes based on the recognition results and customer characteristics, calculating strategy scores according to the weights of the strategy nodes and emotion-intent coupling; after filtering the strategy scores, constructing a business strategy template; and after adaptively adjusting the natural language output on the business strategy template, outputting a natural language response.

[0090] Furthermore, the intelligent customer service automatic reply generation system based on multimodal learning is also used to: read the user's historical reply preference features; establish a preference influence factor based on the historical reply preference features; and update the natural language reply after using the preference influence factor to compensate for the natural language reply.

[0091] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The intelligent customer service automatic reply generation method and specific examples based on multimodal learning in the foregoing embodiment one are also applicable to the intelligent customer service automatic reply generation system based on multimodal learning in this embodiment. Through the foregoing detailed description of the intelligent customer service automatic reply generation method based on multimodal learning, those skilled in the art can clearly understand the intelligent customer service automatic reply generation system based on multimodal learning in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here.

[0092] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0093] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for generating automatic responses to intelligent customer service based on multimodal learning, characterized in that, The method includes: Read multimodal data input by the customer, perform feature encoding on the multimodal data, and construct a modal feature vector. The multimodal data includes text, voice, images and videos, and business structured data. After temporal alignment of the modal feature vectors, feature fusion of the modal feature vectors is performed based on a cross-modal attention mechanism to establish fused features; The fused features are mapped to the entity and relation set of the business knowledge graph, context nodes are established, and a context graph structure is established on the context nodes; The aforementioned context graph structure is used to perform multi-label intent recognition and emotion classification to establish recognition results. The identification results are input into the strategy planner, and a business strategy template is created based on customer characteristics and identification results. Then, a natural language response is output.

2. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 1, characterized in that, The step of mapping the fused features to the entity and relation set of the business knowledge graph, establishing context nodes, and establishing a context graph structure at the context nodes includes: Semantic units of different modalities are extracted from the fused features to establish context nodes, which include intent nodes, emotion nodes, event nodes, and state nodes, wherein: The intent node is constructed through semantic analysis of fused features, including text features, speech-to-text features, and video-to-text features. Analyze the speech prosody of speech features and video features to establish primary emotion associations; The text features are analyzed to identify and establish a second emotion association. The first and second emotion associations are then used to establish emotion nodes. Perform content recognition on images and videos to establish event nodes; Perform structured data analysis of business operations and establish status nodes; Establish a context graph structure at the context node.

3. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 2, characterized in that, The process of establishing a context graph structure at the context node includes: Cross-modal fusion edges are established for the context nodes, including temporal consistency edges, semantic association edges, and emotion transmission edges; The temporally consistent edge is constructed by calculating the association between any two nodes through time windows, and the weight of the temporally consistent edge is constructed using a time similarity function. The semantic association edges are constructed using cross-modal feature similarity, and the weights of the semantic association edges are constructed using fused feature cosine similarity. The emotion transmission edge is constructed through the emotion propagation relationship, and the weight of the emotion transmission edge is constructed through the emotion intensity coefficient and the business sensitivity coefficient; A context graph structure is established using the context nodes and cross-modal fusion edges, and the context graph structure is dynamically updated based on the semantic similarity of nodes and changes in business status during multi-turn dialogues.

4. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 3, characterized in that, The dynamic updating of the context graph structure based on node semantic similarity and business state changes in multi-turn dialogues includes: Set a similarity merging threshold for nodes, and perform node similarity calculation during each round of dialogue update; If the similarity calculation result of the nodes meets the similarity merging threshold, then the corresponding node merging is performed, and the weighted fusion of nodes is performed according to the weighted reliability. Configure the node decay index of the scenario graph, and use the node decay index to perform node aging and elimination management within the scenario graph; The context graph structure is updated by using node merging and node aging and elimination management.

5. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 1, characterized in that, The process of using the context graph structure to perform multi-label intent recognition, emotion classification, and establish recognition results includes: The context graph structure is used as a dynamic context semantic space. The context path matching channel is used to match and search the historical context subgraphs in the context graph structure to establish a candidate intent set. Obtain the associated emotion tags in the context graph structure, and use the associated emotion tags to establish emotion-intention constraints; The candidate intent set is filtered using the emotion-intent constraint to establish the recognition result.

6. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 5, characterized in that, The establishment of the identification result also includes: The identification results are used to perform contextual backtracking to establish a contextual backtracking set; The recognition results are then processed using the context backtracking set to establish and update the recognition results.

7. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 1, characterized in that, The process of establishing a context graph structure at the context node also includes: Generate a source tracing record instruction, and use the source tracing record instruction to save the modality source identifier, timestamp, and node confidence field at the context node; Interpretive traceability management is performed based on the data stored in the context nodes.

8. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 1, characterized in that, The process of inputting the recognition results into the strategy planner, establishing a business strategy template based on customer characteristics and the recognition results, and then outputting a natural language response includes: The strategy planner constructs strategy nodes based on the identification results and customer characteristics, and calculates strategy scores according to the weights of the strategy nodes and the emotion-intent coupling. After filtering by strategy scores, a business strategy template is constructed; After adaptively adjusting the natural language output on the business strategy template, a natural language response is output.

9. The intelligent customer service automatic response generation method based on multimodal learning as described in claim 1, characterized in that, After establishing a business strategy template based on customer characteristics and identification results, the output of a natural language response also includes: Read the user's historical response preference characteristics; Establish a preference influence factor based on the historical response preference characteristics; After compensating the natural language response using the aforementioned preference influence factor, the natural language response is updated.

10. An intelligent customer service automatic response generation system based on multimodal learning, characterized in that, The steps for implementing the intelligent customer service automatic response generation method based on multimodal learning according to any one of claims 1 to 9 include: The modal feature vector construction module is used to read multimodal data input by the customer, perform feature encoding on the multimodal data, and construct modal feature vectors. The multimodal data includes text, voice, images and videos, and business structured data. The fusion feature establishment module is used to perform feature fusion of the modal feature vectors based on a cross-modal attention mechanism after temporal alignment of the modal feature vectors, and establish fusion features. The context node establishment module is used to map the fused features to the entity and relation set of the business knowledge graph, establish context nodes, and establish a context graph structure in the context nodes; The recognition result establishment module is used to perform multi-label intent recognition using the context graph structure, perform emotion classification, and establish recognition results; The natural language response output module is used to input the recognition results into the strategy planner, and after establishing a business strategy template based on customer characteristics and recognition results, output a natural language response.

Citation Information

Patent Citations

  • Language data processing method and system, terminal equipment and storage medium

    CN120144768A

  • Digital intelligent customer service system based on multiple modes

    CN120654819A

  • Intelligent customer service system

    CN120707157A

  • Dialogue interaction system based on multi-modal emotion perception and knowledge graph dynamic enhancement

    CN120745853A

  • KR20250073758A