Intelligent customer feedback system based on natural language processing
By designing an intelligent customer feedback system based on natural language processing, the problem that traditional systems cannot fully utilize multimodal data and accurately identify complex emotional polarity and multiple intentions is solved, and efficient processing and accurate response to customer feedback is achieved.
Patent Information
- Application Number
- CN202510380849.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional customer feedback systems cannot make full use of multimodal data, are difficult to accurately identify complex emotional polarity and multiple intentions, and lack accurate search functions based on semantic similarity, resulting in inaccuracy and inefficiency of feedback responses.
An intelligent customer feedback system based on natural language processing is designed. Multimodal feedback data is obtained through the modal fusion unit and mapped to a shared high-dimensional feature space. The intention extraction unit combines the joint probability model to extract emotional polarity and operation intention. The knowledge search unit uses the knowledge graph for intelligent search. The entity screening unit filters entities based on semantic similarity. The feature encoding unit generates enhanced context features. The answer generation unit generates answer text for customers.
It realizes efficient processing of various types of inputs, accurately extracts emotional polarity and operational intentions in customer feedback, provides intelligent search functions based on knowledge graphs, improves the accuracy and efficiency of feedback, and ensures the accuracy of system response and the effectiveness of information.
Smart Images

Figure CN120216652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of customer feedback, and particularly to an intelligent customer feedback system based on natural language processing. Background Art
[0002] Traditional systems usually can only process single-modal data (such as only processing text or voice), and cannot make full use of data from multiple modalities (such as text, voice, images, etc.). Therefore, these systems cannot comprehensively understand and analyze customer feedback, ignoring the diverse information that customers may show in different situations; moreover, traditional systems often rely on a single algorithm or rule for sentiment analysis and intent extraction, and may not be able to accurately identify complex sentiment polarities (such as neutral sentiment, weak sentiment) or operation intents. In addition, traditional systems have weak processing capabilities for complex sentiment changes or multiple intents, which may lead to misunderstandings of customer needs, thus affecting the accuracy of feedback responses; and when traditional systems handle customer problems, they usually do not use knowledge graphs or related entity embeddings for intelligent retrieval, lacking a precise retrieval function based on semantic similarity, making these systems often inaccurate or incomplete when providing professional knowledge or related information, thus affecting the response quality of the system and the effectiveness of information; and in traditional systems, the screening of entities often relies on simple keyword matching or rule judgment, lacking in-depth semantic analysis. Therefore, the system may process a large number of irrelevant or low-related entities, thus affecting the accuracy and efficiency of feedback, which will result in slow system responses, information redundancy, and even misunderstandings. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to overcome the above-mentioned disadvantages of the prior art and provide an intelligent customer feedback system based on natural language processing.
[0004] The technical solution adopted to solve the above technical problem is: an intelligent customer feedback system based on natural language processing, comprising:
[0005] A modality fusion unit, which is used to obtain multi-modal feedback data of a target customer, encode the multi-modal feedback data according to an encoder to obtain multi-modal feedback features, and map the multi-modal feedback features to a shared high-dimensional feature space according to a preset alignment loss function to obtain fusion features;
[0006] An intent extraction unit, which is used to train the fusion features according to a joint probability model to obtain synchronous operation intents corresponding to the sentiment polarity of the multi-modal feedback data;
[0007] A knowledge retrieval unit, which is used to obtain a feedback knowledge graph, embed each entity in the feedback knowledge graph to obtain an entity embedding matrix corresponding to the feedback knowledge graph, and calculate the semantic similarity between the fusion feature and the entity embedding matrix;
[0008] An entity screening unit, which is used to screen out a relevant entity set in the feedback knowledge graph according to the semantic similarity, and filter the relevant entity set according to the operation intention to obtain a strongly relevant entity set;
[0009] A feature encoding unit, which is used to encode the fusion feature and the entity embedding matrix corresponding to the strongly relevant entity set to obtain an enhanced context feature;
[0010] An answer generation unit, which is used to decode the enhanced context feature to obtain an answer word sequence, and generate an answer text for the target customer according to the answer word sequence.
[0011] Preferably, the multimodal feedback data includes feedback text, feedback voice and feedback image, and the multimodal feedback features include feedback text features, feedback voice features and feedback image features.
[0012] Preferably, encoding the multimodal feedback data by an encoder to obtain multimodal feedback features, and mapping the multimodal feedback features to a shared high-dimensional feature space according to a preset alignment loss function, including:
[0013] Segment the feedback text into a word embedding sequence, and generate a context-aware vector of the word embedding sequence according to the bidirectional Transformer of the BERT-based model, that is, obtain feedback text features;
[0014] Segment the feedback voice into a frame sequence according to the Wav2Vec framework, extract local features of the frame sequence according to a convolutional neural network, and capture long-term temporal dependencies of the local features of the frame sequence according to a Transformer to obtain feedback voice features;
[0015] Segment the feedback image into multiple pixel blocks, linearly project the multiple pixel blocks into a pixel block sequence, and generate feedback image features of the pixel block sequence through the global attention mechanism of the Vision Transformer;
[0016] Map the feedback text features, the feedback voice features and the feedback image features to a shared high-dimensional feature space according to a preset alignment loss function, where the alignment loss function is as follows:
[0017] Lalign = ∑ (x,y,z)∈D (λ(||E T (x) - E A (y)|| 2 + ||E T (x) - E I (z)|| 2 ) + (1 - λ)InfoNCE(E T (x), E A (y), E I (z)));
[0018] where, L align represents the alignment loss function, λ represents the hyperparameter, and the hyperparameter is used to balance the alignment strength and modality specificity, E T (x), E A (y) and E I (z) represent the feedback text feature, the feedback speech feature, and the feedback image feature, D represents the shared high-dimensional feature space, and InfoNCE( ) represents the calculation function for enhancing the inter-modal discrimination through negative sample sampling.
[0019] Preferably, the fusion feature is trained according to the joint probability model to obtain the emotional polarity and synchronous operation intention of the multi-modal feedback data, including:
[0020] Performing emotion extraction on the fusion feature according to the emotion analysis channel to obtain the emotion hidden state of the fusion feature, where the emotion analysis channel uses LSTM;
[0021] Performing emotion classification on the emotion hidden state of the fusion feature according to Softmax to obtain the emotional polarity probability distribution of the fusion feature, and determining the emotional polarity label of the fusion feature according to the emotional polarity probability distribution of the fusion feature;
[0022] Performing label embedding on the emotional polarity label to obtain the label embedding vector of the emotional polarity label;
[0023] Defining an emotion intention coupling matrix, where the emotion intention coupling matrix is used to represent the association strength of the emotional polarity label with the intention, and the association strength is a learnable parameter;
[0024] Performing joint modeling on the emotion intention coupling matrix, the fusion feature, and the corresponding label embedding vector to obtain the emotion intention joint distribution, where the calculation formula of the emotion intention joint distribution is as follows:
[0025]
[0026] where, P(y, s|Hfusion ) represents the joint distribution of emotional intentions, I s,y represents the emotional intention coupling matrix, e s represents the label embedding vector, H fusion represents the fused feature, W[] represents the projection matrix, S represents the set of emotional labels, and Y represents the set of intentions.
[0027] Preferably, the fused feature is trained according to the joint probability model to obtain the emotional polarity and synchronized operation intention of the multimodal feedback data, and further includes:
[0028] Annotators independently annotate the emotional polarity labels and intentions, and generate the fuzzy emotional intention joint distribution between the emotional polarity labels and intentions through vote counting. Among them, the calculation formula of the fuzzy emotional intention joint distribution is as follows:
[0029]
[0030] Among them, P true (s|X s ) represents the fuzzy emotional intention joint distribution, N(s) represents the number of votes for the emotional label, N(y) represents the number of votes for the intention, and X s represents the multimodal feedback data sample;
[0031] Define the overall loss function according to the emotional intention joint distribution and the fuzzy emotional intention joint distribution, and make the emotional intention joint distribution approximate the fuzzy emotional intention joint distribution according to the overall loss function. Among them, the overall loss function is as follows:
[0032]
[0033] Among them, L total represents the overall loss function, X represents the training data set of multimodal feedback data samples, D KL represents P true (y,s|X s ) and P(y,s|H fusion ) between the differences, λ1 and λ2 represent the hyperparameters of the regularization term, represents the measurement of the difference in intention distribution between different emotions;
[0034] Obtain the operation intention corresponding to the emotional polarity label according to the improved emotional intention joint distribution.
[0035] Preferably, the feedback knowledge graph includes product entities, service entities, and fault entities, and the calculation formula of the semantic similarity is as follows:
[0036]
[0037] Among them, Sim represents semantic similarity, and E KG represents the entity embedding matrix.
[0038] Preferably, a relevant entity set in the feedback knowledge graph is screened according to the semantic similarity, and the relevant entity set is filtered according to the operation intention to obtain a strongly relevant entity set, including:
[0039] Compare the semantic similarity with a preset similarity threshold. If it is greater than the preset similarity threshold, add the entity in the feedback knowledge graph corresponding to the semantic similarity to the relevant entity set, and repeat the above operation until all entities in the feedback knowledge graph are compared;
[0040] Calculate the correlation degree between the operation intention and each relevant entity in the relevant entity set. Among them, the calculation process of the correlation degree includes word embedding of the operation intention and the relevant entity to obtain an intention word embedding sequence corresponding to the operation intention and an entity word embedding sequence corresponding to the relevant entity, and calculating the sequence similarity between the intention word embedding sequence and the entity word embedding sequence, where the sequence similarity is the correlation degree;
[0041] Compare the correlation degree with a preset correlation degree threshold. If it is greater than the preset correlation degree threshold, add the relevant entity in the relevant entity set corresponding to the correlation degree to the strongly relevant entity set, and repeat the above operation until all relevant entities in the relevant entity set are compared.
[0042] Preferably, encode the fusion feature and the entity embedding matrix corresponding to the strongly relevant entity set to obtain enhanced context features, including:
[0043] Encode the fusion feature according to a bidirectional LSTM to obtain context features, where the calculation formula of the context features is as follows:
[0044]
[0045] Among them, represents the context feature at the t-th step, represents the fusion feature at the t-th step, and BiLSTM represents a bidirectional LSTM;
[0046] Concatenate the sentiment polarity label of the fusion feature with the entity embedding matrix corresponding to the strongly relevant entity set to obtain a context conditional vector, where the calculation formula of the context conditional vector is as follows:
[0047]
[0048] Among them, c cond represents the context condition vector, represents the entity embedding matrix corresponding to the strongly related entity set, and MeanPooling( ) represents the average pooling operation;
[0049] Feature enhancement is performed on the context features according to the context condition vector to obtain enhanced context features. Among them, the calculation formula of the enhanced context features is as follows:
[0050]
[0051] Among them, represents the enhanced context feature, and W c represents the weight matrix.
[0052] Preferably, decoding is performed on the enhanced context features to obtain an answer word sequence, and an answer text of the target customer is generated according to the answer word sequence, including:
[0053] Obtain the state at the last moment of the enhanced context feature, use the state at the last moment of the enhanced context feature as the initial decoding state, and obtain the decoder states of each step according to the initial decoding state;
[0054] Calculate the entity attention of the strongly related entity set through the decoder states of each step according to the attention mechanism. Among them, the calculation formula of the entity attention is as follows:
[0055]
[0056] Among them, represents the entity attention of the i-th related entity in the related entity set, represents the decoder state at step, and W k represents the entity attention weight matrix;
[0057] Calculate the fusion attention of the fusion features through the decoder states of each step according to the attention mechanism. Among them, the calculation formula of the fusion attention is as follows:
[0058]
[0059] Among them, represents the fusion attention of the fusion features, and W d represents the attention weight matrix of the fusion features.
[0060] Preferably, decoding is performed on the enhanced context features to obtain an answer word sequence, and an answer text of the target customer is generated according to the answer word sequence, further including:
[0061] Calculate a first weighted context vector of the fused feature and a second weighted context vector of the strongly relevant entity set according to the entity attention and the fused attention;
[0062] Calculate a generation probability according to the decoder state of each step size, the first weighted context vector, and the second weighted context vector, where the calculation formula of the generation probability is as follows:
[0063]
[0064] Where, P gen represents the generation probability, σ represents the activation function, and represent the first weighted context vector and the second weighted context vector;
[0065] Calculate an answer word generation probability according to the generation probability and the generation probability of the vocabulary, where the calculation formula of the answer word generation probability is as follows:
[0066]
[0067] Where, P(w) represents the generation probability of the answer word w, P(w) represents, P vocab (w) represents the generation probability of the answer word w in the vocabulary, represents the attention distribution of the first weighted context vector and the second weighted context vector for the answer word w.
[0068] The beneficial effects of the present invention are as follows: (1) The present invention obtains and fuses customer feedback data from different modalities (such as text, speech, image, etc.) through a modality fusion unit, and maps this data to a shared high-dimensional feature space through an encoder and an alignment loss function, thereby enhancing the system's processing ability for various types of inputs. Multimodal fusion helps to extract more comprehensive customer sentiment information, enabling the system to more accurately understand customer feedback. Moreover, through the intention extraction unit, combined with a joint probability model, the fused features of multimodal feedback are trained, and the system can effectively extract the sentiment polarity and corresponding operation intentions in customer feedback. This means that the system can not only perceive the customer's sentiment tendency (such as positive, negative, or neutral), but also accurately understand the specific operations that the customer hopes the system to perform, thus improving the accuracy of feedback response; (2) The present invention can, through the knowledge retrieval unit, embed relevant entities based on the feedback knowledge graph and calculate the semantic similarity between the fused features and the entity embedding matrix, thereby providing the system with an intelligent retrieval function based on the knowledge graph. Through this process, the system can obtain professional knowledge and information related to customer feedback, enhancing the system's knowledge processing ability and the rationality of feedback. Moreover, through the entity screening unit, relevant entities are screened based on semantic similarity and filtered in combination with operation intentions, thereby obtaining a strongly relevant entity set. This process helps to filter out irrelevant or weakly relevant entities, ensuring that the system can focus on entities that are truly related to the customer's problem or needs, improving the accuracy and efficiency of feedback; (3) Through the feature encoding unit, the present invention encodes the fused features and the strongly relevant entity set to generate enhanced context features. These enhanced context features contain more customer feedback information and relevant entity knowledge, making the subsequent answer generation more comprehensive and efficient. Moreover, an answer text for the target customer is generated. Thanks to the previous multimodal fusion, intention extraction, knowledge retrieval, and entity screening, the generated answer can more accurately meet the customer's needs and be more in line with the customer's sentiment and operation intentions. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 FIG. is a schematic diagram of the system architecture of the overall system in an embodiment proposed by the present invention.
[0070] Reference numerals: 1, modality fusion unit; 2, intention extraction unit; 3, knowledge retrieval unit; 4, entity screening unit; 5, feature encoding unit; 6, answer generation unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] Embodiment 1, as Figure 1 shown, an intelligent customer feedback system based on natural language processing proposed by the present invention includes:
[0072] The modality fusion unit 1 is used to obtain the multi-modal feedback data of the target customer, encode the multi-modal feedback data according to the encoder to obtain multi-modal feedback features, and map the multi-modal feedback features to a shared high-dimensional feature space according to a preset alignment loss function to obtain fusion features;
[0073] The intention extraction unit 2 is used to train the fusion features according to the joint probability model to obtain the synchronous operation intention corresponding to the sentiment polarity of the multi-modal feedback data;
[0074] The knowledge retrieval unit 3 is used to obtain the feedback knowledge graph, embed each entity in the feedback knowledge graph to obtain the entity embedding matrix corresponding to the feedback knowledge graph, and calculate the semantic similarity between the fusion features and the entity embedding matrix;
[0075] The entity screening unit 4 is used to screen out the relevant entity set in the feedback knowledge graph according to the semantic similarity, and filter the relevant entity set according to the operation intention to obtain the strongly relevant entity set;
[0076] The feature encoding unit 5 is used to encode the fusion features and the entity embedding matrix corresponding to the strongly relevant entity set to obtain enhanced context features;
[0077] The answer generation unit 6 is used to decode the enhanced context features to obtain an answer word sequence, and generate the answer text of the target customer according to the answer word sequence.
[0078] In the present invention, the alignment loss function is an objective function for optimizing multi-modal data fusion. Its task is to map features of different modalities (such as text features and image features) to a shared high-dimensional feature space to ensure that data of different modalities can be aligned and have consistency in this space; the sentiment polarity is a measure of the sentiment analysis result, indicating the direction of sentiment (for example, positive, negative or neutral). Sentiment polarity analysis is usually used to understand the user's sentiment attitude; the operation intention refers to the specific behavior or operation that the customer hopes to achieve through interacting with the system, which reflects the customer's needs, such as "canceling an order" or "querying logistics", etc.; the knowledge graph is a graphical structure composed of nodes (representing entities) and edges (representing the relationships between entities). In this system, the feedback knowledge graph represents the entities related to the customer feedback (such as "order", "customer service", etc.) and the relationships between them; the semantic similarity is a measurement method for measuring the meaning similarity between two texts or vectors, and usually uses methods such as cosine similarity and Euclidean distance to calculate.
[0079] Example 2. An intelligent customer feedback system based on natural language processing proposed by the present invention. Compared with Example 1, this example further includes: The multimodal feedback data includes feedback text, feedback voice, and feedback images, and the multimodal feedback features include feedback text features, feedback voice features, and feedback image features.
[0080] In an optional embodiment, the encoder is used to encode the multimodal feedback data to obtain multimodal feedback features, and the multimodal feedback features are mapped to a shared high-dimensional feature space according to a preset alignment loss function, including:
[0081] The feedback text is segmented and then converted into a word embedding sequence, and the context-aware vector of the word embedding sequence is generated according to the bidirectional Transformer of the BERT-based model, that is, the feedback text features are obtained;
[0082] The feedback voice is segmented into a frame sequence according to the Wav2Vec framework, the local features of the frame sequence are extracted according to the convolutional neural network, and the long-term temporal dependencies of the local features of the frame sequence are captured according to the Transformer to obtain the feedback voice features;
[0083] The feedback image is segmented into multiple pixel blocks, the multiple pixel blocks are linearly projected into a pixel block sequence, and the feedback image features of the pixel block sequence are generated according to the global attention mechanism of the Vision Transformer;
[0084] The feedback text features, feedback voice features, and feedback image features are mapped to a shared high-dimensional feature space according to a preset alignment loss function, where the alignment loss function is as follows:
[0085] L align =∑ (x,y,z)∈D (λ(||E T (x)-E A (y)|| 2 +‖E T (x)-E I (z)|| 2 )+(1-λ)InfoNCE(E T (x),E A (y),E I (z)));
[0086] Wherein, L align represents the alignment loss function, λ represents a hyperparameter, and the hyperparameter is used to balance the alignment strength and modality specificity, E T (x), E A (y) and E I(z) represents the feedback text feature, feedback voice feature, and feedback image feature, D represents the shared high-dimensional feature space, and InfoNCE( ) represents the calculation function that enhances the inter-modal discrimination through negative sample sampling.
[0087] It should be noted that the BERT-based model is a pre-trained language model based on the Transformer structure. It adopts a bidirectional encoder (i.e., considering both the left and right context information simultaneously), can handle the context relationship of language well, and has been widely used in tasks such as text classification, question answering systems, and named entity recognition; the bidirectional Transformer means that the model not only considers the context information of the previous text (from left to right), but also considers the information of the subsequent text (from right to left) at the same time to enhance the understanding ability; Wav2Vec is a framework for speech recognition, which uses deep neural networks (especially convolutional neural networks and Transformers) to learn effective representations from raw audio signals; the convolutional neural network is a neural network specifically used to process grid-structured data (such as images, audio signals). It can extract local features through convolutional operations and gradually obtain more advanced features at multiple levels; long-term temporal dependence means that when processing time series data (such as speech, video, etc.), the model can capture the dependence relationship over a longer time span; Vision Transformer is an image processing model based on the Transformer architecture. Different from the traditional convolutional neural network (CNN), ViT first divides the image into multiple fixed-size patches, and then linearly projects these image patches into a series of vector sequences, using the Transformer to capture the global features of the image.
[0088] In an optional embodiment, the fusion feature is trained according to the joint probability model to obtain the sentiment polarity and synchronous operation intention of the multi-modal feedback data, including:
[0089] Extract the sentiment from the fusion feature according to the sentiment analysis channel to obtain the sentiment hidden state of the fusion feature, where the sentiment analysis channel uses LSTM;
[0090] Perform sentiment classification on the sentiment hidden state of the fusion feature according to Softmax to obtain the sentiment polarity probability distribution of the fusion feature, and determine the sentiment polarity label of the fusion feature according to the sentiment polarity probability distribution of the fusion feature;
[0091] Perform label embedding on the sentiment polarity label to obtain the label embedding vector of the sentiment polarity label;
[0092] Define the sentiment intention coupling matrix, where the sentiment intention coupling matrix is used to represent the association strength of the sentiment polarity label with the intention, and the association strength is a learnable parameter;
[0093] Jointly model the sentiment-intent coupling matrix, the fused features, and the corresponding label embedding vectors to obtain the sentiment-intent joint distribution. The calculation formula for the sentiment-intent joint distribution is as follows:
[0094]
[0095] where P(y, s|H fusion ) represents the sentiment-intent joint distribution, I s,y represents the sentiment-intent coupling matrix, e s represents the label embedding vector, H fusion represents the fused features, W[ ] represents the projection matrix, S represents the set of sentiment labels, and Y represents the set of intents.
[0096] It should be noted that the sentiment-intent coupling matrix is a matrix representing the relationship between sentiment labels and intents. It is obtained through learning, and each element in it represents the association strength of a certain sentiment polarity label with a specific intent. This matrix is a learnable parameter and is usually optimized through training data to describe the connection between sentiment labels and different intents; the projection matrix is a matrix used to map different features (such as label embeddings, fused features, etc.) into a common space.
[0097] In an alternative embodiment, training the fused features according to the joint probability model to obtain the sentiment polarity and synchronous operation intent of the multimodal feedback data further includes:
[0098] Annotators independently annotate the sentiment polarity labels and intents, and generate a fuzzy sentiment-intent joint distribution between the sentiment polarity labels and intents through vote counting. The calculation formula for the fuzzy sentiment-intent joint distribution is as follows:
[0099]
[0100] where P true (s|X s ) represents the fuzzy sentiment-intent joint distribution, N(s) represents the number of votes for the sentiment label, N(y) represents the number of votes for the intent, and X s represents the multimodal feedback data sample;
[0101] Define the overall loss function according to the sentiment-intent joint distribution and the fuzzy sentiment-intent joint distribution, and make the sentiment-intent joint distribution approximate the fuzzy sentiment-intent joint distribution according to the overall loss function. The overall loss function is as follows:
[0102]
[0103] where L totaldenotes the overall loss function, X denotes the training data set of multimodal feedback data samples, D KL denotes P true (y, s|X s ) and the difference between P(y, s|H fusion ), λ1 and λ2 denote the hyperparameters of the regularization terms, denotes the measure of the difference in the intention distribution between different emotions;
[0104] Obtain the operation intention corresponding to the sentiment polarity label according to the improved joint distribution of sentiment intention.
[0105] It should be noted that the fuzzy joint distribution of sentiment intention is a probability distribution, which represents the fuzzy relationship between sentiment labels and intentions. This distribution takes into account the uncertainty or ambiguity between sentiment and intention. For example, in some cases, a sentiment label may not exactly correspond to a single intention, but may be related to multiple intentions. Therefore, a fuzzy distribution is used to describe this relationship; the vote count of a sentiment label refers to the summary of opinions given by multiple annotators (or annotation systems) for the sentiment label of a certain sample. When each annotator annotates a sample, they will give a sentiment label (for example, positive, negative or neutral). After these labels are summarized, a "vote count" is formed, that is, the frequency of each sentiment label being selected; the vote count of an intention refers to the summary of the intention categories annotated by multiple annotators for the same multimodal feedback data sample. Multiple annotators may give different intention labels, and the frequencies of these labels constitute the vote count; the improved joint distribution of sentiment intention refers to the joint distribution of sentiment intention obtained after optimizing the loss function (such as the above overall loss function).
[0106] In an optional embodiment, the feedback knowledge graph includes product entities, service entities and fault entities, and the calculation formula of semantic similarity is as follows:
[0107]
[0108] where Sim denotes semantic similarity, and E KG denotes the entity embedding matrix.
[0109] In an optional embodiment, filter out the relevant entity set in the feedback knowledge graph according to the semantic similarity, and filter the relevant entity set according to the operation intention to obtain the strongly relevant entity set, including:
[0110] Compare the semantic similarity with a preset similarity threshold. If it is greater than the preset similarity threshold, add the entity in the feedback knowledge graph corresponding to the semantic similarity to the relevant entity set, and repeat the above operation until all entities in the feedback knowledge graph are compared;
[0111] Calculate the association degree between the computing operation intention and each related entity in the set of related entities. Among them, the calculation process of the association degree includes performing word embedding on the operation intention and the related entities to obtain the intention word embedding sequence corresponding to the operation intention and the entity word embedding sequence corresponding to the related entities, and calculating the sequence similarity between the intention word embedding sequence and the entity word embedding sequence. Among them, the sequence similarity is the association degree;
[0112] Compare the association degree with a preset association degree threshold. If it is greater than the preset association degree threshold, add the related entities in the set of related entities corresponding to the association degree to the set of strongly related entities, and repeat the above operations until all related entities in the set of related entities are compared.
[0113] In an alternative embodiment, encode the fusion feature and the entity embedding matrix corresponding to the set of strongly related entities to obtain enhanced context features, including:
[0114] Encode the fusion feature according to the bidirectional LSTM to obtain context features. Among them, the calculation formula of the context features is as follows:
[0115]
[0116] Among them, represents the context feature at the t step, represents the fusion feature at the t step, and BiLSTM represents the bidirectional LSTM;
[0117] Concatenate the sentiment polarity label of the fusion feature with the entity embedding matrix corresponding to the set of strongly related entities to obtain a context conditional vector. Among them, the calculation formula of the context conditional vector is as follows:
[0118]
[0119] Among them, c cond represents the context conditional vector, represents the entity embedding matrix corresponding to the set of strongly related entities, and MeanPooling( ) represents the average pooling operation;
[0120] Enhance the context features according to the context conditional vector to obtain enhanced context features. Among them, the calculation formula of the enhanced context features is as follows:
[0121]
[0122] Among them, represents the enhanced context feature, and W c represents the weight matrix.
[0123] In an alternative embodiment, the enhanced context features are decoded to obtain an answer word sequence, and the answer text of the target customer is generated according to the answer word sequence, including:
[0124] Obtain the state at the last moment of the enhanced context features, use the state at the last moment of the enhanced context features as the initial decoding state, and obtain the decoder states at each step according to the initial decoding state;
[0125] Calculate the entity attention of the strongly relevant entity set through the decoder states at each step according to the attention mechanism, where the calculation formula of the entity attention is as follows:
[0126]
[0127] Where, represents the entity attention of the i-th relevant entity in the relevant entity set, represents the decoder state at step, W k represents the attention weight matrix of the entity;
[0128] Calculate the fusion attention of the fusion features through the decoder states at each step according to the attention mechanism, where the calculation formula of the fusion attention is as follows:
[0129]
[0130] Where, represents the fusion attention of the fusion features, W d represents the attention weight matrix of the fusion features.
[0131] It should be noted that the initial decoding state is the first state in the decoding process of the model. This state is usually based on the state at the last moment generated by the encoder and serves as the input to the decoder. The task of the decoder is usually to generate an output sequence, and the generation or prediction of the sequence starts from this initial state; the decoder state refers to the internal state of the decoder at each time step in the sequence generation or prediction task. The decoder gradually generates the output sequence, and the decoder state at each time step reflects the model's understanding of the current and historical input information. The state of the decoder depends on the context information of the input sequence and is passed through a recursive or cyclic structure.
[0132] In an alternative embodiment, the enhanced context features are decoded to obtain an answer word sequence, and the answer text of the target customer is generated according to the answer word sequence, further including:
[0133] Calculate the first weighted context vector of the fusion features and the second weighted context vector of the strongly relevant entity set according to the entity attention and the fusion attention;
[0134] Calculate the generation probability based on the decoder state, the first weighted context vector, and the second weighted context vector for each step size. The calculation formula for the generation probability is as follows:
[0135]
[0136] where P gen represents the generation probability, σ represents the activation function, and represent the first weighted context vector and the second weighted context vector;
[0137] Calculate the generation probability of the answer word based on the generation probability and the generation probability of the vocabulary. The calculation formula for the generation probability of the answer word is as follows:
[0138]
[0139] where P(w) represents the generation probability of the answer word w, P vocab (w) represents the generation probability of the answer word w in the vocabulary, represents the attention distribution of the first weighted context vector and the second weighted context vector for the answer word w.
[0140] It should be noted that the generation probability of the vocabulary refers to the probability of generating a certain word for a given input sequence. In the generation task, the vocabulary contains all possible generated words, and the model will calculate the generation probability of each word based on the input and context information.
[0141] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made without departing from the spirit of the present invention within the knowledge scope of those skilled in the art.
Claims
1. An intelligent customer feedback system based on natural language processing, characterized in that: include: A modal fusion unit (1), the modal fusion unit (1) is used to obtain multimodal feedback data of a target customer, encode the multimodal feedback data according to an encoder to obtain multimodal feedback features, and map the multimodal feedback features to a shared high-dimensional feature space according to a preset alignment loss function to obtain fusion features; An intention extraction unit (2), the intention extraction unit (2) being used to train the fusion feature according to a joint probability model to obtain a synchronized operation intention corresponding to the emotional polarity of the multimodal feedback data; A knowledge retrieval unit (3), the knowledge retrieval unit (3) is used to obtain a feedback knowledge graph, embed each entity in the feedback knowledge graph to obtain an entity embedding matrix corresponding to the feedback knowledge graph, and calculate the semantic similarity between the fusion feature and the entity embedding matrix; An entity screening unit (4), the entity screening unit (4) is used to screen out a set of related entities in the feedback knowledge graph according to the semantic similarity, and filter the set of related entities according to the operation intention to obtain a set of strongly related entities; A feature encoding unit (5), the feature encoding unit (5) is used to encode the fusion feature and the entity embedding matrix corresponding to the strongly related entity set to obtain an enhanced context feature; An answer generation unit (6) is used to decode the enhanced context feature to obtain an answer word sequence, and generate an answer text of the target customer according to the answer word sequence.
2. The intelligent customer feedback system based on natural language processing according to claim 1, characterized in that: The multimodal feedback data includes feedback text, feedback voice and feedback image, and the multimodal feedback features include feedback text features, feedback voice features and feedback image features.
3. The intelligent customer feedback system based on natural language processing according to claim 2, characterized in that: Encoding the multimodal feedback data according to the encoder to obtain multimodal feedback features, and mapping the multimodal feedback features to a shared high-dimensional feature space according to a preset alignment loss function, including: The feedback text is converted into a word embedding sequence after word segmentation, and a context-aware vector of the word embedding sequence is generated according to the bidirectional Transformer of the BERT-based model, that is, the feedback text feature is obtained; The feedback speech is segmented into a frame sequence according to a Wav2Vec framework, local features of the frame sequence are extracted according to a convolutional neural network, and long-term temporal dependencies of the local features of the frame sequence are captured according to a Transformer to obtain feedback speech features; The feedback image is divided into a plurality of pixel blocks, the plurality of pixel blocks are linearly projected into a pixel block sequence, and feedback image features of the pixel block sequence are generated according to the global attention mechanism of the Vision Transformer; The feedback text feature, the feedback speech feature, and the feedback image feature are mapped to a shared high-dimensional feature space according to a preset alignment loss function, wherein the alignment loss function is as follows: L align =∑ (x,y,z)∈D (λ(||E T (x)-E A (y)|| 2 +‖E T (x)-E I (z)|| 2 )+(1-λ)InfoNCE(E T (x),E A (y),E I (z))); Among them, L align represents the alignment loss function, λ represents a hyperparameter, and the hyperparameter is used to balance the alignment strength and modality specificity, E T (x), E A (y) and E I (z) represents feedback text features, feedback speech features and feedback image features, D represents a shared high-dimensional feature space, and InfoNCE() represents a calculation function for enhancing inter-modal discrimination through negative sample sampling.
4. The intelligent customer feedback system based on natural language processing according to claim 3, characterized in that: The fusion feature is trained according to a joint probability model to obtain the emotional polarity and synchronized operation intention of the multimodal feedback data, including: Performing sentiment extraction on the fused feature according to a sentiment analysis channel to obtain a sentiment hidden state of the fused feature, wherein the sentiment analysis channel adopts LSTM; Performing sentiment classification on the sentiment hidden state of the fused feature according to Softmax to obtain the sentiment polarity probability distribution of the fused feature, and determining the sentiment polarity label of the fused feature according to the sentiment polarity probability distribution of the fused feature; Performing label embedding on the sentiment polarity label to obtain a label embedding vector of the sentiment polarity label; Define an emotion-intention coupling matrix, wherein the emotion-intention coupling matrix is used to represent the association strength of the emotion polarity label to the intention, wherein the association strength is a learnable parameter; The emotion intention coupling matrix, the fusion feature and the corresponding label embedding vector are jointly modeled to obtain the emotion intention joint distribution, wherein the calculation formula of the emotion intention joint distribution is as follows: Among them, P(y,s|H fusion ) represents the joint distribution of sentiment intention, I s,y represents the emotion-intention coupling matrix, e s represents the label embedding vector, H fusion represents the fusion feature, W[] represents the projection matrix, S represents the emotion label set, and Y represents the intent set.
5. The intelligent customer feedback system based on natural language processing according to claim 4, characterized in that: The fusion feature is trained according to a joint probability model to obtain the emotional polarity and synchronized operation intention of the multimodal feedback data, and further includes: The annotator independently annotates the sentiment polarity label and the intention, and generates the fuzzy sentiment intention joint distribution between the sentiment polarity label and the intention through voting statistics, wherein the calculation formula of the fuzzy sentiment intention joint distribution is as follows: Among them, P true (s|X s ) represents the joint distribution of fuzzy sentiment intention, N(s) represents the number of votes for sentiment labels, N(y) represents the number of votes for intentions, X s Represents a multimodal feedback data sample; An overall loss function is defined according to the joint distribution of the emotional intention and the joint distribution of the fuzzy emotional intention, and the joint distribution of the emotional intention is made to approach the joint distribution of the fuzzy emotional intention according to the overall loss function, wherein the overall loss function is as follows: Among them, L total represents the overall loss function, X represents the multimodal feedback data sample training dataset, D KL Indicates P true (y,s|X s ) and P(y,s|H fusion ), λ1 and λ2 represent the hyperparameters of the regularization term, Represents a measure of the difference in intention distribution between different emotions; The operation intention corresponding to the emotion polarity label is obtained according to the improved joint distribution of the emotion intention.
6. The intelligent customer feedback system based on natural language processing according to claim 5, characterized in that: The feedback knowledge graph includes product entities, service entities and fault entities, and the calculation formula of the semantic similarity is as follows: Among them, Sim represents semantic similarity, E KG Represents the entity embedding matrix.
7. The intelligent customer feedback system based on natural language processing according to claim 6, characterized in that: A set of related entities in the feedback knowledge graph is screened out according to the semantic similarity, and the set of related entities is filtered according to the operation intention to obtain a set of strongly related entities, including: The semantic similarity is compared with a preset similarity threshold. If it is greater than the preset similarity threshold, the entity in the feedback knowledge graph corresponding to the semantic similarity is added to the related entity set, and the above operation is repeated until all entities in the feedback knowledge graph are compared; Calculating the degree of association between the operation intention and each related entity in the related entity set, wherein the calculation process of the degree of association includes word embedding of the operation intention and the related entity to obtain an intention word embedding sequence corresponding to the operation intention and an entity word embedding sequence corresponding to the related entity and calculating the sequence similarity between the intention word embedding sequence and the entity word embedding sequence, wherein the sequence similarity is the degree of association; The correlation degree is compared with a preset correlation degree threshold. If it is greater than the preset correlation degree threshold, the related entities in the related entity set corresponding to the correlation degree are added to the strongly related entity set. The above operation is repeated until all the related entities in the related entity set are compared.
8. The intelligent customer feedback system based on natural language processing according to claim 7, characterized in that: Encoding the fused features and the entity embedding matrix corresponding to the strongly related entity set to obtain enhanced context features, including: The fusion feature is encoded according to the bidirectional LSTM to obtain the context feature, wherein the calculation formula of the context feature is as follows: in, represents the context feature at step t, represents the fusion feature at step length t, and BiLSTM represents bidirectional LSTM; The sentiment polarity label of the fusion feature is concatenated with the entity embedding matrix corresponding to the strongly related entity set to obtain a contextual condition vector, wherein the calculation formula of the contextual condition vector is as follows: Among them, c cond represents the context condition vector, Represents the entity embedding matrix corresponding to the set of strongly related entities, and MeanPooling() represents the average pooling operation; The context feature is enhanced according to the context condition vector to obtain an enhanced context feature, wherein the calculation formula of the enhanced context feature is as follows: in, represents the enhanced context feature, W c represents the weight matrix.
9. The intelligent customer feedback system based on natural language processing according to claim 8, characterized in that: Decoding the enhanced context feature to obtain an answer word sequence, and generating an answer text of the target customer according to the answer word sequence, including: Acquire the last moment state of the enhanced context feature, use the last moment state of the enhanced context feature as the initial decoding state, and acquire the decoder state of each step length according to the initial decoding state; The entity attention of the strongly related entity set is calculated by the decoder state of each step according to the attention mechanism, wherein the calculation formula of the entity attention is as follows: in, represents the entity attention of the i-th related entity in the related entity set, represents the decoder state at the step length, and Wk represents the attention weight matrix of the entity; The fusion attention of the fusion feature is calculated by the decoder states of each step according to the attention mechanism, wherein the calculation formula of the fusion attention is as follows: in, represents the fusion attention of the fused features, W d Represents the attention weight matrix of the fused features.
10. The intelligent customer feedback system based on natural language processing according to claim 9, characterized in that: Decoding the enhanced context features to obtain an answer word sequence, and generating an answer text of the target customer according to the answer word sequence, further comprising: Calculating a first weighted context vector of the fused feature and a second weighted context vector of the strongly related entity set according to the entity attention and the fused attention; The generation probability is calculated according to the decoder states of the respective step sizes, the first weighted context vector, and the second weighted context vector, wherein the calculation formula of the generation probability is as follows: Among them, P gen represents the generation probability, σ represents the activation function, and represents a first weighted context vector and a second weighted context vector; The answer word generation probability is calculated according to the generation probability and the generation probability of the vocabulary, wherein the calculation formula of the answer word generation probability is as follows: Among them, P(w) represents the generation probability of answer word w, P(w) represents, P vocab (w) represents the generation probability of answer word w in the vocabulary, Represents the attention allocation of the first weighted context vector and the second weighted context vector to the answer word w.