Intelligent customer service method and system based on user intention recognition

Through multimodal data fusion and interactive recognition mechanism, the intelligent customer service system can more accurately understand user intentions and emotions, solve the problem of incomplete intentions and emotions recognition in the existing technology, and provide a more accurate response strategy.

CN120430404AActive Publication Date: 2025-08-05HENAN YUEBAO NETWORK TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510497173.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-14
Filing Date
2025-04-21
Publication Date
2025-08-05
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

The existing intelligent customer service system fails to effectively integrate multimodal data when handling user consultations, resulting in insufficient comprehensive and accurate understanding of user intentions and emotions. Especially when faced with vague intentions or complex emotions, it is difficult to provide accurate response strategies.

Method used

Using an intelligent customer service method based on user intention recognition, we use the Bert, wav2vec and Faster RCNN models to extract features by receiving multimodal conversation data (including text, audio and video), and generate multimodal features through timing alignment and fusion. Combining the attention mechanism and the Transformer encoder to achieve interactive recognition of emotions and intentions, iteratively adjust the recognition results until the preset conditions are met.

Benefits of technology

It improves the accuracy of user intention and emotion recognition, can have a deeper understanding of user needs, provide response interaction results that are more in line with user expectations, and avoids information loss or misunderstanding caused by single modal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430404A_ABST
    Figure CN120430404A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent customer service method and system based on user intention recognition, and belongs to the technical field of intelligent customer service, and the method comprises the steps: S1, receiving multi-modal dialogue data of a user; s2, analyzing the multi-modal dialogue data based on a user intention recognition model to obtain a user intention recognition result; and S3, a response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy. According to the invention, interactive guidance of user intention recognition and emotion recognition can be realized, and the accuracy of user intention recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent customer service, and in particular relates to an intelligent customer service method and system based on user intent recognition. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, intelligent customer service has become a crucial link between users and services. With its powerful automated processing capabilities and efficient service response speeds, it is gradually replacing traditional manual customer service models. Intelligent customer service systems can receive user inquiries in real time and, leveraging advanced natural language processing technology and machine learning algorithms, quickly analyze user intent and provide relevant answers or services. While existing intelligent customer service solutions have achieved some success in intelligent response, they still exhibit numerous shortcomings in practical applications. For example, most intelligent customer service systems still rely primarily on single text data when processing user inquiries, insufficiently utilizing other modalities such as voice and images, resulting in an incomplete and inaccurate understanding of user intent.

[0003] Sentiment analysis is a key research area in intelligent customer service. A user's emotional state may influence their acceptance and satisfaction with intelligent customer service responses. Through sentiment analysis, intelligent customer service systems can identify users' emotional states, such as anger, satisfaction, and anxiety, in real time. Based on these emotional states and the results of user intent recognition, they can adjust service strategies to provide more accurate, attentive, and personalized service. In reality, user intent and emotions are closely linked. Intent often reflects user emotions, and accurately understanding user emotions is crucial for understanding their true intent. However, in existing intelligent customer service systems, intent recognition and sentiment analysis technologies are often implemented separately, without considering the effective integration and mutual guidance of sentiment and intent analysis. This is particularly inadequate when handling complex and emotionally rich user inquiries. In particular, when faced with inquiries involving ambiguous intent, complex emotions, or diverse expressions, existing intelligent customer service systems often struggle to accurately capture users' true needs, leading to inappropriate response strategies or inaccurate answers, severely impacting user experience and service quality.

[0004] Therefore, providing an intelligent customer service method and system that can integrate multimodal data, realize interactive recognition of emotions and intentions, and improve the accuracy of user intention and emotion recognition has become an urgent problem to be solved. Summary of the Invention

[0005] In view of the above-mentioned defects in the prior art, the present invention provides an intelligent customer service method based on user intention recognition, comprising the following steps:

[0006] S1: receiving multimodal conversation data of a user, where the multimodal conversation data includes text data;

[0007] S2: Analyzing the multimodal conversation data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module;

[0008] The shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively;

[0009] The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module;

[0010] The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module;

[0011] The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result;

[0012] S3: The response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy.

[0013] Furthermore, in step S1, the multimodal conversation data also includes audio data and video data; in step S2, the shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively, including:

[0014] Using the Bert model as a text encoder to extract text features of the text data;

[0015] Using the wav2vec model as an acoustic encoder to extract acoustic features of the audio data;

[0016] Using the Faster RCNN model as a visual encoder to extract visual features of the video data;

[0017] A temporal alignment module is used to align the text features, the acoustic features, and the visual features, wherein the temporal alignment module is composed of an LSTM module and a Softmax function;

[0018] Connecting the aligned text features with the acoustic features to obtain acoustic enhancement features;

[0019] Connecting the aligned text features with the visual features to obtain visual enhancement features;

[0020] Fusing the acoustic enhancement feature with the visual enhancement feature to obtain a non-text feature;

[0021] Calculating a fusion weight between the text feature and the non-text feature, and obtaining the multimodal feature based on the fusion weight;

[0022] The multimodal features are sent to the emotion recognition module and the intention recognition module.

[0023] Furthermore, the emotion recognition module includes an attention module and an emotion recognition model;

[0024] In step S2, the emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs the user's emotion recognition result, and inputs the emotion category into the intention recognition module, including:

[0025] In a first iteration, the emotion recognition module performs emotion recognition based on the multimodal features to obtain a first emotion recognition result, and inputs the first emotion recognition result into the intention recognition module;

[0026] At the kth iteration, the emotion recognition module implements intention-guided emotion recognition using an attention mechanism based on the multimodal features and the k-1th user intention recognition result, obtains the kth user emotion recognition result, and inputs the kth emotion recognition result into the intention recognition module, where 2≤k≤K, where K is the preset iteration number threshold;

[0027] The method of using the attention mechanism to implement intention-guided emotion recognition includes:

[0028] The attention module receives the multimodal features and the user intention recognition result as input, obtains the attention weight of each element according to the correlation between each element in the multimodal features and the user intention recognition result, inputs the multimodal features and the attention weight into the emotion recognition model, and obtains the kth user emotion recognition result.

[0029] Further, the intention recognition module includes an intention recognition model implemented based on a Transformer encoder; in step S2, the intention recognition module realizes emotion-driven intention recognition based on the multi-modal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module, including:

[0030] At the t-th iteration, the intention recognition model takes the multi-modal features and the t-th emotion recognition result as inputs, and obtains the t-th user intention recognition result, where 1 ≤ t ≤ K;

[0031] When t < K and the change of the joint loss function is not less than a preset threshold, the t-th user intention recognition result is input into the emotion recognition module;

[0032] When t = K, or when the change of the joint loss function is less than the preset threshold, the t-th user intention recognition result is used as the final user intention recognition result, and the t-th user emotion recognition result is used as the final user emotion recognition result.

[0033] Further, in step S3, the response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy, including:

[0034] The response module determines a response strategy based on the final user emotion recognition result, determines response content based on the final user intention recognition result, and feeds back the response content to the user according to the response strategy.

[0035] The present invention also provides an intelligent customer service system based on user intention recognition, and the system includes:

[0036] An input module, configured to receive multi-modal dialogue data of a user, where the multi-modal dialogue data includes text data;

[0037] A user intention recognition module, configured to analyze the multi-modal dialogue data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module;

[0038] Among them, the shared underlying module is configured to analyze the multi-modal dialogue data to obtain multi-modal features, and input the multi-modal features into the emotion recognition module and the intention recognition module respectively;

[0039] The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module;

[0040] The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module;

[0041] The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result;

[0042] The response module is used to determine a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feed back the response content to the user according to the determined response strategy.

[0043] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned intelligent customer service method based on user intention recognition.

[0044] By receiving and integrating multimodal user data, the present invention captures user information more comprehensively, enabling a more accurate understanding of the user's true intentions. This avoids information loss or misunderstandings that may arise from single-modal data, and improves the accuracy of user intent recognition. Furthermore, the emotion recognition module in the present invention utilizes an attention mechanism to implement intention-guided emotion recognition. Simultaneously, the emotion recognition results serve as input information to assist the intention recognition module in making more accurate intention judgments. This interactive guidance mechanism of intention recognition and emotion recognition enables the system to more deeply understand the user's needs and emotions. When faced with inquiries where the user's intentions are ambiguous and their emotions are complex, the system can accurately capture the user's true needs and provide interactive responses that better meet the user's expectations. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the following detailed description with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0046] Figure 1The present invention is a flowchart illustrating an intelligent customer service method based on user intent recognition according to an embodiment of the present invention.

[0047] Figure 2 FIG. 4 is a diagram illustrating the architecture of a user intent recognition model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0048] To make the objectives, technical solutions, and advantages of the present invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making creative efforts shall fall within the scope of protection of the present invention.

[0049] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two.

[0050] It should be understood that although the terms "first," "second," "third," etc. may be used to describe "...," these "..." should not be limited to these terms. These terms are merely used to distinguish "...." For example, "first..." could also be referred to as "second...", and similarly, "second..." could also be referred to as "first..." without departing from the scope of the present invention.

[0051] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0052] Depending on the context, the words "if" and "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting." Similarly, depending on the context, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)."

[0053] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the product or device comprising the element.

[0054] like Figure 1 As shown, an embodiment of the present invention discloses an intelligent customer service method based on user intent recognition, the method comprising:

[0055] S1: receiving multimodal conversation data of a user, where the multimodal conversation data includes text data;

[0056] S2: Analyzing the multimodal conversation data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module;

[0057] The shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively;

[0058] The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module;

[0059] The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module;

[0060] The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result;

[0061] S3: The response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy.

[0062] When implementing intelligent customer service, the embodiment of the present invention adopts an identification mechanism of intention recognition and emotion recognition interactive guidance. The emotion recognition module uses the attention mechanism to realize intention-guided emotion recognition. At the same time, the emotion recognition results serve as input information to assist the intention recognition module to make more accurate intention judgments, so that the system can understand the user's needs and emotions more deeply. When faced with consultations with vague user intentions and complex emotions, it can accurately capture the user's real needs and provide response interaction results that are more in line with the user's expectations.

[0063] In step S1, the multimodal conversation data further includes audio data and video data; in step S2, the shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively, including:

[0064] Using the Bert model as a text encoder to extract text features of the text data;

[0065] Using the wav2vec model as an acoustic encoder to extract acoustic features of the audio data;

[0066] Using the Faster RCNN model as a visual encoder to extract visual features of the video data;

[0067] A temporal alignment module is used to align the text features, the acoustic features, and the visual features, wherein the temporal alignment module is composed of an LSTM module and a Softmax function;

[0068] Connecting the aligned text features with the acoustic features to obtain acoustic enhancement features;

[0069] Connecting the aligned text features with the visual features to obtain visual enhancement features;

[0070] Fusing the acoustic enhancement feature with the visual enhancement feature to obtain a non-text feature;

[0071] Calculating a fusion weight between the text feature and the non-text feature, and obtaining the multimodal feature based on the fusion weight;

[0072] The multimodal features are sent to the emotion recognition module and the intention recognition module.

[0073] In this example, the Bert language model, which has excellent performance in the field of natural language processing, is used to extract text features of text data. i , the text features are obtained from the last hidden layer of the Bert encoder as follows:

[0074]

[0075] Where TextEncoder represents the Bert encoder, Represents the input text sentence t i The text features obtained after passing the Bert encoder, l S Represents the text sentence t i The length of the text sentence t i feature dimension.

[0076] For audio data, the wav2vec model is used to extract the acoustic features of the audio data. The wav2vec model is a self-supervised pre-trained acoustic recognition model implemented through a convolutional neural network. It is shown below:

[0077]

[0078] Among them, AcousticEncoder represents the wav2vec model, Represents the input audio segment a i The acoustic features obtained by the wav2vec model, l A Represents audio segment a i The length, d A Represents the acoustic feature dimension of audio segment a.

[0079] For video data, Faster RCNN is used to extract the visual features of the video data.

[0080] Among them, Faster RCNN is a typical representative of the two-stage target detection model, as shown below:

[0081]

[0082] Among them, VisualEncoder represents the Faster RCNN model, Represents the input video segment v i The visual features obtained by the Faster RCNN model, l V Represents a video clip v i The length, d V Represents the visual feature dimension of audio segment a.

[0083] After obtaining the above text features, acoustic features, and visual features, they need to be aligned. In this embodiment of the present invention, a temporal alignment module is used to align the text features, acoustic features, and visual features. The temporal alignment module consists of an LSTM network and a Softmax function. The alignment operation is represented as follows:

[0084]

[0085] in, They represent the alignment features of text modality data, audio modality data, and video modality data respectively. It can be seen that the alignment operation actually uses text as a reference to achieve alignment with audio and video.

[0086] After alignment, the text features are connected with the acoustic features and the visual features to obtain acoustic enhancement features and visual enhancement features:

[0087]

[0088] in, and Represent acoustic enhancement features and visual enhancement features respectively, ReLU represents the activation function, f * represents a linear layer, and || represents a connection.

[0089] Then, the acoustic enhancement feature and the visual enhancement feature are fused to obtain the non-text feature h i , as shown below:

[0090]

[0091] in l represents the non-text feature h i The length, f * Represents a linear layer. Finally, the fusion weight between the text features and the non-text features is calculated to fuse the text features and the non-text features to obtain multimodal features:

[0092]

[0093] Among them, ||〃||2 represents the L2 norm, ε is a hyperparameter, and f represents a normalization block containing a dropout layer.

[0094] Through the above method, multimodal features of user input data can be obtained, avoiding the information loss or misunderstanding that may be caused by single modal data, and improving the accuracy of user intention recognition. In addition, the unique method of obtaining multimodal features in the embodiment of the present invention can more accurately understand contextual information and better capture the user's true intention compared to the existing method of obtaining multimodal features. After obtaining the above multimodal features, it is necessary to send the multimodal features to the emotion recognition module and the intention recognition module respectively to identify user emotions and user intentions.

[0095] In an embodiment of the present invention, the emotion recognition module includes an attention module and an emotion recognition model. There are various implementation schemes for emotion recognition models in the art, such as LSTM-based emotion recognition models EF-LSTM, LF-LSTM, and M-BERT based on BERT. In an embodiment of the present invention, the EF-LSTM emotion recognition model is used as an example to describe how to implement intent-guided emotion recognition. Since user intent recognition is not performed during the first emotion recognition, emotion recognition needs to be divided into two stages: the first emotion recognition and multiple emotion recognitions after the first emotion recognition.

[0096] Specifically, in step S2, the emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs the user's emotion recognition result, and inputs the emotion category into the intention recognition module, including:

[0097] During the first iteration, the emotion recognition module performs emotion recognition based on the multimodal features, obtains a first emotion recognition result, and inputs this first emotion recognition result into the intent recognition module. Since user intent recognition is not performed at this point, the multimodal features are directly used as input to the EF-LSTM emotion recognition model, and the output of the EF-LSTM emotion recognition model is the first emotion recognition result.

[0098] At the kth iteration, the emotion recognition module uses an attention mechanism to implement intent-guided emotion recognition based on the multimodal features and the k-1th user intent recognition result, obtaining the kth user emotion recognition result, and inputting the kth emotion recognition result into the intent recognition module, where 2≤k≤K, where K is the preset iteration number threshold. At this point, to achieve intent-guided emotion recognition, the attention mechanism is used to fuse the multimodal features and the k-1th intent recognition result, and the fusion result is used as input to the EF-LSTM emotion recognition model to obtain the kth emotion recognition result.

[0099] The method of using the attention mechanism to implement intention-guided emotion recognition includes: the attention module receives the multimodal features and the user intention recognition result as input, obtains the attention weight of each element in the multimodal features based on the correlation between each element and the user intention recognition result, inputs the multimodal features and the attention weight into the emotion recognition model, and obtains the kth user emotion recognition result. The detailed process is as follows:

[0100] The above multimodal features Expressed as where z n Represents the multimodal feature Contains the nth element. Then, the attention weight is calculated: w i =Attention(z n ,I), where Attention() is used to calculate element z n The correlation between the intent recognition result I and the result is output as a weight. Here, the cosine similarity can be used to implement Attention(). Through the above calculation, the attention weight vector W = [w1, w2, ..., w n ]. Therefore, the weighted multimodal features are expressed as Finally, the weighted multimodal feature z' is input into the EF-LSTM emotion recognition model to obtain the intention-guided emotion recognition result.

[0101] It can be seen from this that in an embodiment of the present invention, the importance of each element in the multimodal feature is adjusted through the attention mechanism, so as to pay more attention to the features related to the user intention and make full use of the emotional information reflected by the user intention to improve the accuracy of emotion recognition.

[0102] In an embodiment of the present invention, the intent recognition module includes an intent recognition model implemented based on a Transformer encoder. Using a Transformer encoder to implement user intent recognition is a well-known technology in the art and will not be described in detail here.

[0103] In step S2, the intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module, including:

[0104] At the tth iteration, the intention recognition model takes the multimodal features and the tth emotion recognition result as input to obtain the tth user intention recognition result, 1≤t≤K. The detailed process is described as follows:

[0105] First, the multimodal features and the emotion recognition result e t For fusion, splicing or weighted fusion can be used to form a unified input representation Calculation is performed using the self-attention mechanism and feedforward neural network in the Transformer encoder: Let the input of the lth layer be V t (l-1) , for the first layer, V t (D) =X t , then the output of the lth layer is:

[0106] V t (l) =FFN(V t (l-1) +Attention(V t (l-1) ))

[0107] Among them, Attention represents the self-attention mechanism, and FFN represents the feedforward neural network. Through a fully connected layer, V t (L) Mapped to the label space of user intent recognition, the predicted user intent y is obtained t =W O V t (L) , where W O is a learnable weight matrix, L is the number of encoder layers, V t (L) is the final output of the encoder.

[0108] From this process, we can see that the emotion recognition result e t This affects the calculation of attention scores, and thus the generation of the encoder's final output representation. For example, if emotion recognition results indicate that the user is currently angry, the model may pay more attention to anger-related features when processing the multimodal data. This allows the model to better capture the emotional component of user intent, thereby improving the accuracy of user intent recognition.

[0109] In an embodiment of the present invention, user intent recognition and user emotion recognition are interactively guided, so the problem of stopping iteration needs to be considered jointly. On the one hand, a maximum number of iterations K is set for the user intent recognition model. By limiting the number of iterations, the model is prevented from overfitting and the stability of the model is ensured. On the other hand, during the iteration process, the loss function value after each iteration is calculated. If the difference between the results of two consecutive iterations is very small and less than a preset threshold, it can be considered that the model has converged, that is, the change in the loss function value tends to be stable, and the iteration can be stopped at this time. The iterative process of the intention recognition module is as follows:

[0110] When t < K and the change of the joint loss function is not less than a preset threshold, input the t-th user intention recognition result into the emotion recognition module;

[0111] When t = K or when the change of the joint loss function is less than the preset threshold, use the t-th user intention recognition result as the final user intention recognition result and use the t-th user emotion recognition result as the final user emotion recognition result.

[0112] Wherein, the joint loss function is the joint loss function of the emotion recognition module and the intention recognition module. As follows:

[0113]

[0114] Wherein, y t is the sample category of intention recognition, is the predicted category of intention recognition, e t is the sample category of emotion recognition, is the predicted category of emotion recognition, and t is the number of iterations.

[0115] In an embodiment of the present invention, after obtaining the final user intention recognition result and user emotion recognition result, the response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy, specifically including:

[0116] The response module determines a response strategy based on the final user emotion recognition result, determines a response content based on the final user intention recognition result, and feeds back the response content to the user according to the response strategy. For example, when it is recognized that the user is in different emotions such as anger, positivity, etc., the response content corresponding to the user intention is presented to the user accordingly. Or, according to the different degrees of the user's emotion, different processing methods are selected, such as directly transferring to a human customer service. The relevant content all belongs to the known calculations in the art and will not be elaborated here.

[0117] According to another embodiment of the present invention, the present invention also provides an intelligent customer service system based on user intention recognition, and the system includes:

[0118] An input module, configured to receive multi-modal dialogue data of a user, and the multi-modal dialogue data includes text data;

[0119] A user intention recognition module, configured to analyze the multimodal conversation data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module;

[0120] The shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively;

[0121] The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module;

[0122] The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module;

[0123] The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result;

[0124] The response module is used to determine a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feed back the response content to the user according to the determined response strategy.

[0125] According to another embodiment of the present invention, the present invention also provides a computer-readable storage medium, which stores a computer program, and is characterized in that when the computer program is executed by a processor, it implements the above-mentioned intelligent customer service method based on user intent recognition.

[0126] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0127] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0128] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0131] The above introduces the preferred embodiments of the present invention, which is intended to make the spirit of the present invention clearer and easier to understand, and is not intended to limit the present invention. Any modifications, replacements, and improvements made within the spirit and principles of the present invention should be included in the scope of protection outlined by the claims attached to the present invention.

Claims

1. An intelligent customer service method based on user intent recognition, characterized in that: The following steps are involved: S1: receiving multimodal conversation data of a user, where the multimodal conversation data includes text data; S2: Analyzing the multimodal conversation data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module; The shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively; The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module; The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module; The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result; S3: The response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy.

2. The intelligent customer service method based on user intention recognition according to claim 1, characterized in that: In step S1, the multimodal conversation data also includes audio data and video data; in step S2, the shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively, including: Using the Bert model as a text encoder to extract text features of the text data; Using the wav2vec model as an acoustic encoder to extract acoustic features of the audio data; Using the Faster RCNN model as a visual encoder to extract visual features of the video data; A temporal alignment module is used to align the text features, the acoustic features, and the visual features, wherein the temporal alignment module is composed of an LSTM module and a Softmax function; Connecting the aligned text features with the acoustic features to obtain acoustic enhancement features; Connecting the aligned text features with the visual features to obtain visual enhancement features; Fusing the acoustic enhancement feature with the visual enhancement feature to obtain a non-text feature; Calculating a fusion weight between the text feature and the non-text feature, and obtaining the multimodal feature based on the fusion weight; The multimodal features are sent to the emotion recognition module and the intention recognition module.

3. The intelligent customer service method based on user intention recognition according to claim 2, characterized in that: The emotion recognition module includes an attention module and an emotion recognition model; In step S2, the emotion recognition module, based on the multi-modal features, uses an attention mechanism to achieve intention-guided emotion recognition, outputs the emotion recognition result of the user, and inputs the emotion category into the intention recognition module, including: In the first iteration, the emotion recognition module performs emotion recognition according to the multi-modal features, obtains the first emotion recognition result, and inputs the first emotion recognition result into the intention recognition module; In the k-th iteration, the emotion recognition module, according to the multi-modal features and the (k - 1)-th user intention recognition result, uses an attention mechanism to achieve intention-guided emotion recognition, obtains the k-th user emotion recognition result, and inputs the k-th emotion recognition result into the intention recognition module, where 2 ≤ k ≤ K, and K is the preset iteration number threshold; Among them, the use of the attention mechanism to achieve intention-guided emotion recognition includes: The attention module receives the multi-modal features and the user intention recognition result as inputs, obtains the attention weight of each element according to the correlation between each element in the multi-modal features and the user intention recognition result, and inputs the multi-modal features and the attention weight into the emotion recognition model to obtain the k-th user emotion recognition result.

4. The intelligent customer service method based on user intention recognition according to claim 3, characterized in that: The intention recognition module includes an intention recognition model implemented based on a Transformer encoder; in step S2, the intention recognition module, based on the multi-modal features and the emotion recognition result, achieves emotion-driven intention recognition, and inputs the user intention recognition result into the emotion recognition module, including: In the t-th iteration, the intention recognition model takes the multi-modal features and the t-th emotion recognition result as inputs to obtain the t-th user intention recognition result, where 1 ≤ t ≤ K; When t < K and the change of the joint loss function is not less than the preset threshold, the t-th user intention recognition result is input into the emotion recognition module; When t = K, or when the change of the joint loss function is less than the preset threshold, the t-th user intention recognition result is used as the final user intention recognition result, and the t-th user emotion recognition result is used as the final user emotion recognition result.

5. The intelligent customer service method based on user intention recognition according to claim 1, characterized in that: In step S3, the response module of the intelligent customer service determines a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feeds back the response content to the user according to the determined response strategy, including: The response module determines a response strategy based on the final user emotion recognition result, determines response content based on the final user intention recognition result, and feeds back the response content to the user according to the response strategy.

6. An intelligent customer service system based on user intention recognition, characterized in that: The system includes: An input module for receiving multi-modal dialogue data of the user, and the multi-modal dialogue data includes text data; A user intention recognition module, configured to analyze the multimodal conversation data based on a user intention recognition model to obtain a user intention recognition result; the user intention recognition model includes a shared underlying module, an emotion recognition module, and an intention recognition module; The shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively; The emotion recognition module implements intention-guided emotion recognition based on the multimodal features using an attention mechanism, outputs an emotion recognition result of the user, and inputs the emotion recognition result into the intention recognition module; The intention recognition module implements emotion-driven intention recognition based on the multimodal features and the emotion recognition result, and inputs the user intention recognition result into the emotion recognition module; The emotion recognition module and the intention recognition module perform multiple iterations. When the number of iterations reaches a preset iteration number threshold, or when the change in the joint loss function of the emotion recognition module and the intention recognition module is less than a preset threshold, the user intention recognition result currently output by the intention recognition module is used as the final user intention recognition result, and the user emotion recognition result currently output by the emotion recognition module is used as the final user emotion recognition result; The response module is used to determine a response strategy and response content based on the final user intention recognition result and the final user emotion recognition result, and feed back the response content to the user according to the determined response strategy.

7. The intelligent customer service system based on user intention recognition according to claim 6, characterized in that: The multimodal conversation data also includes audio data and video data; the shared underlying module is used to analyze the multimodal conversation data to obtain multimodal features, and input the multimodal features into the emotion recognition module and the intention recognition module respectively, including: Using the Bert model as a text encoder to extract text features of the text data; Using the wav2vec model as an acoustic encoder to extract acoustic features of the audio data; Using the Faster RCNN model as a visual encoder to extract visual features of the video data; A temporal alignment module is used to align the text features, the acoustic features, and the visual features, wherein the temporal alignment module is composed of an LSTM module and a Softmax function; Connecting the aligned text features with the acoustic features to obtain acoustic enhancement features; Connecting the aligned text features with the visual features to obtain visual enhancement features; Fusing the acoustic enhancement feature with the visual enhancement feature to obtain a non-text feature; Calculating a fusion weight between the text feature and the non-text feature, and obtaining the multimodal feature based on the fusion weight; The multimodal features are sent to the emotion recognition module and the intention recognition module.

8. The intelligent customer service system based on user intention recognition according to claim 7, characterized in that: The emotion recognition module includes an attention module and an emotion recognition model; The emotion recognition module, based on the multi-modal features, uses an attention mechanism to achieve intention-guided emotion recognition, outputs the emotion recognition result of the user, and inputs the emotion category into the intention recognition module, including: In the first iteration, the emotion recognition module performs emotion recognition according to the multi-modal features, obtains the first emotion recognition result, and inputs the first emotion recognition result into the intention recognition module; In the k-th iteration, the emotion recognition module uses the attention mechanism to achieve intention-guided emotion recognition according to the multi-modal features and the k-1-th user intention recognition result, obtains the k-th user emotion recognition result, and inputs the k-th emotion recognition result into the intention recognition module, where 2 ≤ k ≤ K, and K is the preset iteration number threshold; Among them, the use of the attention mechanism to achieve intention-guided emotion recognition includes: The attention module receives the multi-modal features and the user intention recognition result as inputs, obtains the attention weight of each element according to the correlation between each element in the multi-modal features and the user intention recognition result, inputs the multi-modal features and the attention weight into the emotion recognition model, and obtains the k-th user emotion recognition result.

9. The intelligent customer service system based on user intention recognition according to claim 8, characterized in that: The intention recognition module includes an intention recognition model implemented based on a Transformer encoder; the intention recognition module, based on the multi-modal features and the emotion recognition result, realizes emotion-driven intention recognition, and inputs the user intention recognition result into the emotion recognition module, including: In the t-th iteration, the intention recognition model takes the multi-modal features and the t-th emotion recognition result as inputs, and obtains the t-th user intention recognition result, where 1 ≤ t ≤ K; When t < K and the change of the joint loss function is not less than the preset threshold, the t-th user intention recognition result is input into the emotion recognition module; When t = K, or when the change of the joint loss function is less than the preset threshold, the t-th user intention recognition result is used as the final user intention recognition result, and the t-th user emotion recognition result is used as the final user emotion recognition result.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements an intelligent customer service method based on user intention recognition according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Robot oriented multimodal emotion data interaction method and device

    CN106773923A

  • Fine-grained video emotion content question and answer method and system based on multi-modal data

    CN116226347A

  • Intelligent emotion recognition method based on multi-modal data fusion

    CN118709146A

  • Emotional evolution method and terminal for virtual avatar in educational metaverse

    US20250014470A1

Cited By

  • Customer service emotional intention accurate recognition method based on artificial intelligence

    CN121256021A