Intelligent customer service dialogue generation method and system based on user intention recognition
By adopting a dialogue generation method based on user intent recognition in the intelligent customer service system, using the fusion intent feature matrix and intent migration analysis, combined with the improved Transformer network model, the shortcomings of the existing system in understanding user intent and maintaining dialogue coherence are solved, and more accurate and professional dialogue generation is achieved.
Patent Information
- Application Number
- CN202510076172.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
The existing intelligent customer service system is difficult to accurately understand user intentions, which leads to a large deviation from user needs. Especially in multiple rounds of dialogue scenarios, the dynamic migration rules of intentions cannot be effectively captured, affecting the consistency and consistency of reply.
Using an intelligent customer service dialogue generation method based on user intent recognition, a fused intent feature matrix is extracted from user text, and intention migration analysis is performed in combination with historical dialogue data, and a joint training is used for improved Transformer network model to generate a reply that meets the expected intent transfer rules.
It improves the accuracy and robustness of intention recognition, ensures that the generated content conforms to the expected intent transfer rules, enhances the professionalism and coherence of dialogue generation, and improves the quality and reliability of system responses.
Smart Images

Figure CN119990149A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of intelligent customer service technology, and in particular to a method and system for generating intelligent customer service dialogues based on user intent recognition. Background Art
[0002] With the rapid development of artificial intelligence technology, intelligent customer service systems have been widely used in various fields. Traditional customer service systems based on rule matching and retrieval often use keyword matching and template filling to generate responses, which makes it difficult to accurately understand the user's true intentions, resulting in a large deviation between the response content and user needs. Most existing methods treat intent recognition and dialogue generation as independent tasks, lacking an effective intent information interaction mechanism.
[0003] At present, the mainstream dialogue generation method mainly adopts sequence-to-sequence model. Although it can generate fluent reply text, due to the lack of in-depth understanding and modeling of user intentions, the generated replies often have problems such as irrelevant answers and semantic incoherence. Especially in multi-round dialogue scenarios, due to the failure to effectively capture the dynamic migration rules of intentions, the generated replies cannot accurately grasp the contextual semantics and it is difficult to maintain the coherence and consistency of the dialogue. In addition, the existing customer service dialogue system performs poorly when dealing with complex intention transfer scenarios and has difficulty in accurately predicting the dynamic changes of user intentions. This is mainly due to the lack of systematic modeling and constraint mechanism for intention transfer patterns, and the inability to effectively utilize the intention transfer rules contained in historical dialogue data. At the same time, existing methods often ignore the domain knowledge contained in the customer service speech templates and fail to effectively integrate them into the dialogue generation process, affecting the professionalism and accuracy of the replies. Summary of the invention
[0004] The present application provides a method and system for generating intelligent customer service dialogues based on user intent recognition. The present application improves the accuracy and robustness of intent recognition, and ensures that the generated content conforms to the expected rules of intent transfer while ensuring the fluency of responses.
[0005] In a first aspect, the present application provides a method for generating an intelligent customer service dialogue based on user intent recognition, and the method for generating an intelligent customer service dialogue based on user intent recognition includes:
[0006] Obtaining a first user text input from the intelligent customer service system, and performing feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix;
[0007] Based on the fusion intention feature matrix, intention migration analysis and dynamic Bayesian analysis are performed on the dialogue sequence data in the historical dialogue corpus to obtain an intention state transition probability matrix;
[0008] Inputting the fused intent feature matrix and the intent state transition probability matrix into an improved Transformer network model for joint training to obtain a dialogue generation model;
[0009] Based on the dialogue generation model, the second user text newly input into the intelligent customer service system is decoded and generated, and the intention state transition probability matrix is used as a constraint condition to obtain a reply text that meets the dialogue intention.
[0010] A second aspect of the present application provides an intelligent customer service dialogue generation system based on user intent recognition, and the intelligent customer service dialogue generation system based on user intent recognition includes:
[0011] An acquisition module, used to acquire a first user text input from the intelligent customer service system, and perform feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix;
[0012] An analysis module, configured to perform intention migration analysis and dynamic Bayesian analysis on the dialogue sequence data in the historical dialogue corpus based on the fused intention feature matrix to obtain an intention state transition probability matrix;
[0013] A training module, used for inputting the fusion intention feature matrix and the intention state transition probability matrix into an improved Transformer network model for joint training to obtain a dialogue generation model;
[0014] A generation module is used to decode and generate a newly input second user text in the intelligent customer service system based on the dialogue generation model, and use the intention state transition probability matrix as a constraint condition to obtain a reply text that meets the dialogue intention.
[0015] Compared with the prior art, the present application has the following beneficial effects: by designing a two-layer intent feature extraction network, the main intent features and the secondary intent features are extracted respectively, which can fully capture the intent information in the user text and improve the accuracy and robustness of intent recognition. By adopting an improved Transformer encoder structure, introducing a contextual intent understanding sublayer and a residual connection mechanism, the effective fusion of primary and secondary intent features is achieved, and the model's ability to understand complex intents is enhanced. The parameter learning method based on the intent migration graph and the dynamic Bayesian network effectively models the dynamic characteristics of intent transfer and provides precise intent constraints for dialogue generation. The intent recognition and dialogue generation tasks are jointly optimized, and the speech template matching mechanism is combined to ensure the professionalism and coherence of the generated replies. A decoding constraint mechanism based on the intent state transition probability matrix is designed to ensure that the generated content conforms to the expected intent transfer law while ensuring the fluency of the reply. A comprehensive scoring mechanism is used to screen candidate replies, comprehensively considering the intent relevance and template matching, and improving the quality and reliability of the system reply. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0017] The structures, proportions, sizes, etc. illustrated in the drawings of this specification are only used to match the contents disclosed in the specification so as to facilitate understanding and reading by persons familiar with this technology. They are not used to limit the conditions under which the present invention can be implemented, and therefore have no substantive technical significance. Any structural modification, change in proportion or adjustment of size, without affecting the effects and purposes that can be achieved by the present invention, should still fall within the scope of the technical contents disclosed by the present invention.
[0018] Figure 1 It is a flowchart of a method for generating an intelligent customer service dialogue based on user intent recognition provided by an embodiment of the present invention;
[0019] Figure 2 It is a schematic block diagram of the structure of an intelligent customer service dialogue generation system based on user intent recognition provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.
[0022] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0023] It should be further understood that the term "and / or" used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. Figure 1 In the embodiment of the present application, an embodiment of the intelligent customer service dialogue generation method based on user intent recognition includes:
[0024] Step 100: Obtain input first user text from the intelligent customer service system, and perform feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix;
[0025] It is understandable that the execution subject of the present application can be an intelligent customer service dialogue generation system based on user intention recognition, or a terminal or a server, which is not limited here. The present application embodiment is described by taking the server as the execution subject as an example.
[0026] Specifically, a first user text input by a user is obtained from the intelligent customer service system, and the text is a user question or request in natural language form. The first user text is segmented and converted into a word unit sequence composed of independent word units. The word segmentation tools in natural language processing, such as NLTK, Jieba and other word segmentation tools, are used to decompose the long text into basic semantic units. The word unit sequence is input into the Word2Vec word vector model to complete the vector mapping. The Word2Vec model can learn the semantic distribution of each word through the training corpus and convert the word unit sequence into the first word vector input sequence. The word vector is a dense numerical vector representation that can effectively capture the semantic association between words and provide compatible numerical input for subsequent deep learning models. The first word vector input sequence is input into the main intention feature extraction layer to extract the main features of the user's intention. The main intention feature extraction layer adopts a multi-head attention mechanism, which consists of 8 attention heads, and the dimension of each attention head is 64. The multi-head attention mechanism can focus on different parts of the user text at different semantic levels and capture multi-dimensional semantic relationships and features. Each attention head learns an attention pattern independently, and interacts the input word vector with the global context by calculating the weighted sum, and generates 8 sets of attention head feature vectors. The 8 sets of attention head feature vectors are concatenated, and the concatenated high-dimensional vectors are reduced in dimension through the linear transformation layer to reduce redundant features and retain the most informative features to obtain the main feature data of user intent. At the same time, in order to capture the auxiliary features of user intent, the first word vector input sequence is input into the secondary intent feature extraction layer for processing. This layer consists of a forward long short-term memory unit (LSTM) and a backward long short-term memory unit, and its hidden layer dimension is set to 256. The forward LSTM captures the semantic dependency of the text from front to back, while the backward LSTM processes in reverse and captures the semantic association from back to front. The feature sequences generated by the bidirectional LSTM are concatenated in the temporal dimension to obtain the auxiliary feature data of user intent, which supplements the detailed semantic information not captured in the main intent feature data. The main feature data of user intent and the auxiliary feature data of user intent are fused through the attention mechanism. The attention mechanism dynamically adjusts the weights according to the correlation between the primary and auxiliary feature data, highlights important features and suppresses secondary features, and obtains a fusion intent feature matrix.
[0027] The main feature data of user intention is input into the first self-attention sublayer for dot product attention calculation. The self-attention mechanism is a calculation method based on query, key and value. By performing inner product operation on the input features, the similarity distribution between each group of features and the remaining features is obtained. The main feature data of user intention is mapped into query matrix, key matrix and value matrix, and a main feature attention weight matrix is generated through matrix calculation. This matrix reflects the correlation weights between the internal parts of the main feature data and can highlight the important information. At the same time, the auxiliary feature data of user intention is input into the second self-attention sublayer for dot product attention calculation. The auxiliary feature data generates query matrix, key matrix and value matrix through mapping, and generates auxiliary feature attention weight matrix through inner product calculation, which reflects the semantic correlation structure in the auxiliary feature data. The main feature attention weight matrix and the auxiliary feature attention weight matrix are concatenated to form a combined feature weight matrix. In order to enhance the feature representation capability, the combined feature weight matrix is nonlinearly transformed, and the matrix is element-wise operated using activation function (such as ReLU) to generate the initial fused feature matrix. In order to make full use of the main feature data, the initial fusion feature matrix is residually connected with the main feature data of user intent. By adding the initial fusion feature matrix to the original main feature data, the model can retain the original information of the main feature and introduce the enhanced information of the fusion feature. The first layer of normalization units is used to normalize the result after residual connection to generate the main image fusion matrix. The normalization operation improves the stability of model training and prevents gradient disappearance or explosion by standardizing the range of eigenvalues. Similarly, the main image fusion matrix is residually connected with the auxiliary feature data of user intent, and the auxiliary feature information is injected into the fusion result in this way. After being processed by the second layer of normalization units, the auxiliary intent fusion matrix is generated. In order to improve the expressive power of the fusion feature, the main image fusion matrix and the auxiliary intent fusion matrix are input into the feedforward neural network for feature enhancement. The feedforward neural network consists of multiple fully connected layers and activation functions. By performing layer-by-layer nonlinear transformations on the input features, a high-dimensional enhanced feature matrix is generated. The enhanced feature matrix is input into the feature fusion layer for final processing. The feature fusion layer generates the final fusion intent feature matrix by weighting, normalizing and optimizing the enhanced features.
[0028] Step 200: Based on the fusion intention feature matrix, perform intention migration analysis and dynamic Bayesian analysis on the dialogue sequence data in the historical dialogue corpus to obtain an intention state transition probability matrix;
[0029] Specifically, the dialogue sequence data is extracted from the historical dialogue corpus and segmented according to the dialogue identifier. The dialogue turns under the same dialogue identifier are grouped into multiple dialogue subsequences by classifying the dialogue records in the original corpus. This segmentation method retains the logical integrity of the conversation and ensures that the conversion process of user intent and the coherence of the dialogue context can be accurately captured in subsequent analysis. The segmented multiple dialogue subsequences are analyzed one by one using the fusion intention feature matrix. Based on the semantic vector in the fusion intention feature matrix, feature matching is performed on each dialogue text. The feature matching method is implemented by calculating semantic similarity, for example, using cosine similarity, Euclidean distance, etc. to measure the matching degree between the fusion intention feature matrix and each dialogue text. Through feature matching, a most matching intent label is assigned to each dialogue text, and a corresponding dialogue intent label sequence is generated. The dialogue intent label sequence is statistically analyzed to calculate the transfer frequency between the intent labels of adjacent dialogue turns in the sequence. By constructing an intent transfer count matrix, each element in the matrix represents the transfer frequency from one intent category to another intent category. In order to make the data in the matrix more meaningful, each count value in the intention transfer count matrix is divided by the sum of the corresponding row to complete the normalization process and obtain the intention transfer probability distribution matrix. Each row in the matrix represents the transfer probability distribution of an intention category. Based on the intention transfer probability distribution matrix, the initial intention node graph is constructed. Each node of the initial intention node graph corresponds to an intention category, and the probability value in the matrix is used to represent the transfer relationship between intentions. According to the probability value, a directed connection is established between the intention nodes, and the probability value is used as the weight of the connection to form a weighted intention transfer graph structure, which reflects the transfer rules and strengths between different intentions. In order to optimize the structure of the intention transfer graph, the connection paths in the weighted intention transfer graph are screened, and only the connection paths with weights greater than the preset threshold are retained to obtain the optimized intention transfer graph structure. Each intention node in the optimized intention transfer graph structure is associated with the customer service speech template, and each intention node is directly mapped to a specific customer service speech set, so that the intention node can support the needs of actual dialogue generation, and the intention migration graph is obtained, which contains the transfer relationship between intentions and combines the speech template to enhance its practicality. The intention migration graph is input into the dynamic Bayesian network for parameter learning. The dynamic Bayesian network is a time series model that captures the dynamic characteristics of the time dimension by learning from historical data. In the parameter learning process, the dynamic Bayesian network combines the structure of the intention transition graph with the characteristics of the historical corpus, optimizes the transition probability estimation between intention states, and generates the final intention state transition probability matrix.
[0030] The nodes in the intent transition graph are connected in a directed manner according to the temporal dependency relationship to form a temporal dependency graph. In the intent transition graph, each node represents a user intent, and each edge represents the transfer relationship between intents. By analyzing the probability distribution and transfer direction in the intent transition graph, these nodes are connected in a directed manner according to the temporal dependency to generate a temporal dependency graph that can characterize the dynamic transfer of intents. Its directed edges reflect the causal relationship or conditional dependency between intent nodes. Based on the temporal dependency graph, each node in it is set as a random variable node. Each random variable node represents a possible state of an intent, and the directed connection between nodes represents the conditional dependency between these states. A conditional probability distribution is established between these random variable nodes to characterize how the state of each node depends on the state of its parent node, and a Bayesian network structure is obtained. The conditional probability distribution of the Bayesian network is a key part, reflecting the transfer pattern between intents and possible state dependencies. The conditional probability distribution in the Bayesian network structure is initialized to a uniform distribution. This initialization method ensures that all possible transfers are given equal probability in the initial stage of learning, avoiding bias towards certain specific transfer patterns too early. At the same time, probability is assigned to the state space of each node to generate an initial parameter matrix, which describes the initial probability distribution of each intent node under different states. In order to optimize the parameters of the Bayesian network, a variational lower bound function is constructed for the initial parameter matrix. The variational lower bound function is a lower bound based on the log-likelihood, and its function is to approximate the posterior probability distribution through optimization. When constructing the variational lower bound function, the dimension of the variational parameter is set equal to the number of intent nodes to ensure that each intent node has an independent parameter optimization target. Based on the lower bound function, a variational objective function is obtained, and its optimization result can reflect the dependency between states in the network. The variational objective function is input into the variational inference module, and the parameters of the Bayesian network are learned by optimization. In the variational inference process, iterative calculations are performed by minimizing the variational free energy, so as to gradually approach the posterior probability distribution of the target. Each iteration updates the parameters so that the network can better fit the transition law in the data. After the optimization is completed, the target posterior probability distribution obtained contains the precise probability relationship between the intent node and its state. The target posterior probability distribution is further processed, and the possible intention state sequence is simulated by parameter sampling. The number of state transitions between adjacent time series nodes is counted to generate a state transition frequency matrix. The state transition frequency matrix records the actual transition frequency between different intention state pairs, which can intuitively reflect the transition rules in the corpus. The maximum likelihood probability is calculated according to the state transition frequency matrix, and the transition probability of each state pair is normalized to obtain the state transition rule. The normalization process ensures that the sum of the transition probabilities is 1, forming a standardized probability distribution. The state transition rules are sorted according to the corresponding relationship of the intention nodes so as to connect with the semantic structure of the intelligent customer service system and generate the intention state transition probability matrix.
[0031] Step 300: Input the fusion intention feature matrix and the intention state transition probability matrix into the improved Transformer network model for joint training to obtain a dialogue generation model;
[0032] It should be noted that the fused intent feature matrix is input into the encoder part of the improved Transformer network model for feature encoding. The encoder of the improved Transformer network model is composed of 6 stacked encoding layers, each of which includes three main sublayers, namely, the multi-head self-attention sublayer, the first feedforward neural network sublayer, and the contextual intent understanding sublayer. In the multi-head self-attention sublayer, the fused intent feature matrix generates global semantic relevance weights through the self-attention mechanism, so that the model can capture the dependencies between different parts of the input feature matrix. The first feedforward neural network sublayer extracts and enhances these features locally through nonlinear transformation. The contextual intent understanding sublayer introduces additional contextual information to ensure that the processing of the fused intent feature matrix can fully consider the semantic fluidity of the conversation context. After the 6-layer processing of the encoder, the encoded feature sequence is obtained, which is a high-dimensional semantic representation of the input features, including global intent features and their contextual relationships. The encoded feature sequence is input into the decoder part of the improved Transformer network model for sequence decoding. The decoder also consists of 6 decoding layers, each of which includes a masked self-attention sublayer, a cross attention sublayer, and a second feedforward neural network sublayer. In the masked self-attention sublayer, the decoder introduces a masking mechanism to limit the model to only focus on the previously generated sequence part, thereby ensuring the causality of the decoding process. The cross-attention sublayer generates a correlation representation between the decoding state and the input feature by interacting the current state of the decoder with the encoded feature sequence output by the encoder. The second feedforward neural network sublayer further performs a nonlinear transformation on the interacted features to obtain the output of the decoding layer. After 6 layers of decoding, the initial dialogue generation probability distribution is generated, which represents the probability of each candidate word that may be generated under the current dialogue state. At the same time, the intention state transition probability matrix is input into the intention constraint network of the improved Transformer network model for processing. The intention constraint network consists of three fully connected layers, whose function is to map the intention state transition probability matrix to a low-dimensional vector space and generate an intention constraint vector. The intention constraint vector can provide a compact representation of the intention state transition rules and provide additional guidance information for the model to ensure that the generated dialogue conforms to the expected intention transition pattern. The KL divergence calculation is performed on the initial dialogue generation probability distribution and the intention constraint vector to obtain the intention Figure 1The KL divergence is used to measure the difference between the initial generated probability distribution and the intent constraint vector. The optimization goal is to minimize this difference so that the generated dialogue is consistent with the expected intent. The initial dialogue generation probability distribution is input into the template matching network of the improved Transformer network model for template matching. The template matching network consists of a bidirectional long short-term memory layer and a dot product attention layer. It uses a bidirectional LSTM to capture the global sequence information in the generated probability distribution, and then calculates the similarity score between the generated distribution and the predefined dialogue template through dot product attention. Based on the template matching similarity score, the template matching cross entropy loss is calculated to measure the degree of match between the generated dialogue and the predefined template. Figure 1 The consistency loss value and the template matching cross entropy loss are weightedly combined to form a multi-task joint loss function. Figure 1 The optimization goals of consistency and template matching accuracy provide multi-dimensional training constraints for the model. During the back propagation process, the parameters of the entire improved Transformer network model are optimized by minimizing the multi-task joint loss function. After the optimization is completed, the updated network parameters are loaded into the model to generate the final dialogue generation model.
[0033] Step 400: Decode and generate the second user text newly input into the intelligent customer service system based on the dialogue generation model, and use the intention state transition probability matrix as a constraint condition to obtain a reply text that meets the dialogue intention.
[0034] Specifically, the newly input second user text in the intelligent customer service system is segmented and decomposed into word-unit sequences for subsequent vectorization operations. Through word segmentation, natural language is converted into a structured data form that is easy for computer processing. The segmentation result is input into a word vector model (such as Word2Vec or GloVe) to complete the vector mapping and obtain the second word vector input sequence. The second word vector input sequence is subjected to primary and secondary intention feature extraction and feature fusion. The primary intention feature extraction is implemented through a multi-head attention mechanism to capture the most relevant part of the input text with global semantics, while the secondary intention feature extraction is extracted through a bidirectional long short-term memory network (BiLSTM) to extract the context dependency information of the input text. The primary and secondary intention features are combined, and a fused input feature representation is generated through a feature fusion mechanism. The fused input feature representation is input into the encoder module in the dialogue generation model for feature encoding. The multi-layer structure and self-attention mechanism of the encoder can extract the semantic information in the fused input features and generate a high-dimensional feature encoding result. At the same time, based on the intention state transition probability matrix, the feature encoding result is predicted for state transition. The core of state transition prediction is to combine the current user text and historical context, and calculate the target intent vector of the next round of dialogue through dynamic Bayesian analysis or other prediction mechanisms. The target intent vector represents the system's prediction result of the user's intent, providing constraints for the subsequent generation stage. In the decoding process, the decoding process is constrained based on the target intent vector so that the generated candidate replies can better meet the expected intent. By combining the target intent vector with the generation mechanism of the decoder, the model generates multiple candidate reply sequences that cover the possibilities of different semantic expressions and meet the requirements of the target intent within a certain range. In order to evaluate the quality of these candidate reply sequences, they are input into the intent matching unit for evaluation. The intent matching unit calculates the cosine similarity between each candidate reply sequence and the target intent vector to obtain the intent relevance score, which reflects the degree of match between the candidate reply and the target intent. The higher the score, the more the candidate reply meets the user's intent. Template matching calculation is performed on the candidate reply sequence. The template matching score is calculated by comparing the candidate reply with the predefined dialogue template. The higher the score, the higher the fit between the candidate reply and the template. In order to comprehensively evaluate the candidate reply sequence, the intent relevance score and the template matching score are weighted summed to obtain the comprehensive score of each candidate reply sequence. The comprehensive score combines the evaluation results of intent matching and template matching, which can more comprehensively measure the quality of candidate responses. Based on the comprehensive score, the candidate response sequences are sorted, and the candidate response sequence with the highest score is selected as the final response text.
[0035] In the embodiment of the present application, by designing a two-layer intent feature extraction network, the main intent features and the secondary intent features are extracted respectively, which can fully capture the intent information in the user text and improve the accuracy and robustness of intent recognition. The improved Transformer encoder structure is adopted, and the contextual intent understanding sublayer and residual connection mechanism are introduced to realize the effective fusion of the main and secondary intent features, and enhance the model's ability to understand complex intents. The parameter learning method based on the intent migration graph and the dynamic Bayesian network effectively models the dynamic characteristics of intent transfer and provides accurate intent constraints for dialogue generation. The intent recognition and dialogue generation tasks are jointly optimized, and the speech template matching mechanism is combined to ensure the professionalism and coherence of the generated replies. A decoding constraint mechanism based on the intent state transition probability matrix is designed to ensure that the generated content conforms to the expected intent transfer law while ensuring the fluency of the reply. A comprehensive scoring mechanism is used to screen candidate replies, comprehensively considering the intent relevance and template matching, and improving the quality and reliability of the system reply.
[0036] In a specific embodiment, the process of executing step 100 may specifically include the following steps:
[0037] Obtain a first user text input from the intelligent customer service system, perform word segmentation on the first user text to obtain a word unit sequence, input the word unit sequence into a Word2Vec word vector model for vector mapping, and obtain a first word vector input sequence;
[0038] The first word vector input sequence is input into the main intent feature extraction layer for processing. The main intent feature extraction layer includes 8 attention heads, each with a dimension of 64, and 8 groups of attention head feature vectors are obtained;
[0039] The 8 groups of attention head feature vectors are concatenated and dimensionality reduction is performed through a linear transformation layer to obtain the main feature data of user intention;
[0040] The first word vector input sequence is input into the secondary intent feature extraction layer for processing. The secondary intent feature extraction layer includes a forward long short-term memory unit and a backward long short-term memory unit. The hidden layer dimension is 256, and a bidirectional feature sequence is obtained.
[0041] Perform time-series splicing on the bidirectional feature sequences to obtain auxiliary feature data of user intention;
[0042] Attention fusion is performed on the main feature data of user intention and the auxiliary feature data of user intention to obtain a fused intention feature matrix.
[0043] Specifically, the text input by the user is obtained from the intelligent customer service system, and the text is segmented and decomposed into a sequence of word units. The word segmentation tool is used to decompose the text into a sequence of word units, each of which represents a basic semantic unit. The word unit sequence after word segmentation is input into the Word2Vec word vector model for vector mapping. Assuming that Word2Vec has been pre-trained on a large-scale corpus, each word unit is mapped to a vector of a fixed dimension by looking up the word vector table. For the entire word unit sequence, a word vector input sequence is obtained:
[0044] V=[v1,v2,…,v n ];
[0045] Among them, v i is the word vector of the i-th word unit, and n is the number of word units. The first word vector input sequence V is input into the main intent feature extraction layer for processing. The main intent feature extraction layer adopts a multi-head attention mechanism, including 8 attention heads, and the dimension of each attention head is 64. Each attention head calculates attention through the following formula:
[0046]
[0047] Where Q = VW Q , K = VW K , V=VW V is the query, key and value matrix, W Q ,W K ,W V is a trainable weight matrix, is a scaling factor used to prevent the inner product value from being too large and causing the softmax function to be oversaturated. For each attention head, the output feature dimension generated is 64, and the 8 attention heads independently calculate and output 8 sets of attention feature vectors A1, A2, …, A8. These feature vectors are concatenated into a matrix:
[0048] A=[A1;A2;…;A8];
[0049] Then, the concatenated matrix is reduced in dimension through the linear transformation layer to generate the main feature data F of the user’s intention. main :
[0050] F min =AW min ;
[0051] Among them, W main is the weight matrix of dimensionality reduction. At the same time, the first word vector is input into the sequence V and input into the secondary intent feature extraction layer for processing. The secondary intent feature extraction layer includes forward and backward long short-term memory networks (LSTM). The forward LSTM calculates the time series features H from left to right. f, the backward LSTM calculates the time series feature H from right to left b Assuming the hidden layer dimension is 256, the output dimension of both the forward and backward LSTM is 256. The bidirectional feature sequence is obtained by concatenating the forward and backward feature representations:
[0052] H=[H f ;H b ];
[0053] Each line of H represents the bidirectional semantics of the current word in the context. The bidirectional feature sequence H is spliced in time sequence to generate the user intention auxiliary feature data F sub , which integrates the context information of each word and retains the temporal dependency. main and auxiliary feature data F sub Perform attention fusion to generate the fusion intention feature matrix F fusion The process of attention fusion is calculated by the following formula:
[0054]
[0055] Among them, W1, W2 are trainable weight matrices, and d is the scaling factor of the feature dimension.
[0056] In a specific embodiment, the execution step performs attention fusion on the user intention main feature data and the user intention auxiliary feature data to obtain a fused intention feature matrix, which may specifically include the following steps:
[0057] Input the main feature data of user intention into the first self-attention sublayer for dot product attention calculation to obtain the main feature attention weight matrix, and input the auxiliary feature data of user intention into the second self-attention sublayer for dot product attention calculation to obtain the auxiliary feature attention weight matrix;
[0058] The main feature attention weight matrix and the auxiliary feature attention weight matrix are concatenated to obtain a combined feature weight matrix, and the combined feature weight matrix is nonlinearly transformed to obtain an initial fused feature matrix;
[0059] Perform residual connection on the initial fusion feature matrix and the main feature data of user intention, and normalize them through the first layer normalization unit to obtain the main image fusion matrix;
[0060] The main intention fusion matrix and the user intention auxiliary feature data are residually connected and normalized through the second layer normalization unit to obtain the auxiliary intention fusion matrix;
[0061] The main intention fusion matrix and the auxiliary intention fusion matrix are input into the feedforward neural network for feature enhancement to obtain an enhanced feature matrix, and the enhanced feature matrix is input into the feature fusion layer for processing to obtain a fused intention feature matrix.
[0062] Specifically, the main feature data of user intention is input into the first self-attention sub-layer, and dot product attention calculation is performed to generate the main feature attention weight matrix. Assume that the main feature data of user intention is represented by matrix F main ∈R n×d Represented as n, where n is the number of features and d is the dimension of each feature. The dot product attention is calculated by the following formula:
[0063] Q min =F min W Q ,K min =F min W K ,V min =F min W V ;
[0064] Among them, W Q ,W K ,W V ∈R d×dk is the trainable weight matrix, Q main ,K main ,V main are query, key, and value matrices, respectively, and d k is the scaling factor. Then calculate the attention weight:
[0065]
[0066] Among them, A main is the main feature attention weight matrix, which represents the correlation between the main features. Similarly, the user intention auxiliary feature data F sub ∈R n×d Enter the second self-attention sub-layer and use the same formula as the main feature to calculate the dot product attention:
[0067] Q sub =F sub W Q ,K sub =F sub W K ,V sub =F sub W V ;
[0068]
[0069] Finally, the auxiliary feature attention weight matrix A is obtained sub. The main feature attention weight matrix A main And the auxiliary feature attention weight matrix A sub Perform concatenation operations to generate a combined feature weight matrix A comb :
[0070] A comb =[A main ; A sub ];
[0071] Among them, [;] represents the vertical concatenation of matrices. The combined feature weight matrix is transformed nonlinearly. The nonlinear transformation is achieved through activation functions (such as ReLU):
[0072] F init =ReLU(A comb W comb );
[0073] Among them, F init is the initial fusion feature matrix, W comb ∈R 2d×d is the dimension reduction weight matrix. In order to retain the main feature information, the initial fusion feature matrix F init and user intention main feature data F main Make a residual connection:
[0074] F res-main =F init +F main ;
[0075] Then, the first layer normalization unit F res-main Perform normalization to generate the main image fusion matrix:
[0076] F main-fuse =LayerNorm(F res-main );
[0077] Similarly, the main image fusion matrix F main-fuse and user intention auxiliary feature data F sub Make a residual connection:
[0078] F res-sub =F main-fuse +F sub ;
[0079] Through the second layer normalization unit F res-sub After normalization, the auxiliary intention fusion matrix is obtained:
[0080] F sub-fuse =LayerNorm(F res-sub );
[0081] The main image is fused into the matrix Fmain-fuse and auxiliary intention fusion matrix F sub-fuse Input feedforward neural network (FFN) for feature enhancement. The feedforward neural network consists of two linear transformation layers and an activation function:
[0082] F enh =ReLU(F main-fuse W1+b1)W2+b2;
[0083] Among them, W1, W2 are weight matrices of linear transformation, b1, b2 are bias terms, and F enh is the enhanced feature matrix. enh Input feature fusion layer for processing to generate the final fusion intent feature matrix F fusion The feature fusion layer combines different features through a weighted mechanism:
[0084] F fusion =F enh W fusion ;
[0085] Among them, W fusion is the fusion weight matrix, F fusion Contains semantic information of main features and auxiliary features.
[0086] In a specific embodiment, the process of executing step 200 may specifically include the following steps:
[0087] According to the conversation identifier, the conversation sequence data in the historical conversation corpus is segmented to obtain multiple groups of conversation subsequences;
[0088] Based on the fusion intention feature matrix, feature matching is performed on each dialogue text in multiple groups of dialogue subsequences to obtain a dialogue intention label sequence;
[0089] Perform frequency statistics on the intent labels of adjacent dialogue turns in the dialogue intent label sequence to obtain the intent transfer count matrix, and divide the count values in the intent transfer count matrix by the sum of the corresponding rows for normalization to obtain the intent transfer probability distribution matrix;
[0090] Construct an initial intent node graph based on the intent transfer probability distribution matrix. Each node in the initial intent node graph corresponds to an intent category.
[0091] Establish directed connections between nodes according to the probability values in the intention transfer probability distribution matrix, use the probability values as the weights of the connections, and obtain a weighted intention transfer graph structure;
[0092] The connection paths with weights greater than a preset threshold in the weighted intention transfer graph structure are retained to obtain an optimized intention transfer graph structure, and each intention node in the optimized intention transfer graph structure is associated with a customer service speech template to obtain an intention migration graph;
[0093] The intention transition graph is input into the dynamic Bayesian network for parameter learning to obtain the intention state transition probability matrix.
[0094] Specifically, the conversation data is extracted from the historical conversation corpus and segmented according to the conversation identifier to generate multiple conversation subsequences. Assume that the historical corpus contains multiple complete conversation records, each of which contains multiple rounds of conversation data. For example, a conversation record is a multiple round of interaction between a user and a customer service representative. After segmenting these data according to the conversation identifier, they are converted into {S1, S2, …, S m}, where S i represents the ith conversation subsequence, and m is the total number of conversations. The conversation turns in each subsequence are represented by {T1, T2, …, T n}, where n is the number of dialogue turns in the subsequence. The fusion intention feature matrix is used to perform feature matching on the dialogue text in each subsequence to generate a dialogue intention label sequence. Assume that the fusion intention feature matrix is F intent ∈R k×d , where k is the number of intent categories and d is the feature dimension. Each conversation text generates its feature vector v through feature extraction text ∈R d By calculating v text With F intent The cosine similarity of each row in determines the intent label of the current text:
[0095]
[0096] Among them, f i Yes F intent The i-th row represents the feature vector of the i-th type of intent. The intent label of the current text is determined based on the maximum similarity, for example, it is labeled as Intent j By matching the features of each conversation text, the corresponding intent tag sequence {L1, L2, …, L n}, where L i Represents the intent label of the i-th round of dialogue. After generating the intent label sequence, the frequency of the intent labels of adjacent rounds in the sequence is counted to construct the intent transfer count matrix C∈R k×k . The element C in the matrix ij Intent i Transfer to Intent jThe frequency of the intention transfer in all subsequences is counted to fill the matrix completely. Each row is normalized to convert the transfer frequency into the transfer probability:
[0097]
[0098] Among them, P ij Intent i Transfer to Intent j The probability of intention transfer is obtained, and the intention transfer probability distribution matrix P is obtained. The initial intention node graph is constructed based on P. In the initial intention node graph, each node represents an intention category, and the directed connection between nodes represents the intention transfer relationship. The weight of the connection is determined by the probability value in P. For example, if P 12 =0.6, a directed edge with a weight of 0.6 is established between nodes Intent1 and Intent2. In order to optimize the graph structure, the connection paths are screened according to the preset threshold τ, and only paths with weights greater than τ are retained. For example, if τ = 0.5, only paths with weights greater than 0.5 are retained, thereby obtaining a simplified intent transfer graph structure. Associate the optimized intent transfer graph structure with the customer service speech template. Each intent node is associated with one or more predefined dialogue templates so that the corresponding template can be directly called when generating a dialogue. The optimized intent transition graph is input into the dynamic Bayesian network for parameter learning to generate an intent state transition probability matrix. The parameter learning process of the dynamic Bayesian network is based on the intent transfer data and optimizes the conditional probability distribution in the network through maximum likelihood estimation:
[0099]
[0100] Among them, Θ is the parameter set of the dynamic Bayesian network, P(Intent t |Intent t-1 ,Θ) indicates that the intent in the previous round is Intent t-1 The current round intent is Intent t Through the above steps, the intention state transition probability matrix is generated, which integrates the intention transfer rules in historical data.
[0101] In a specific embodiment, the execution step of inputting the intention transition graph into the dynamic Bayesian network for parameter learning to obtain the intention state transition probability matrix may specifically include the following steps:
[0102] Directed connections are constructed for the nodes in the intention migration graph according to the temporal dependency relationship to obtain a temporal dependency graph;
[0103] Each node in the time series dependency graph is set as a random variable node, and conditional probability distribution is established between random variable nodes to obtain a Bayesian network structure;
[0104] Initialize the conditional probability distribution in the Bayesian network structure to a uniform distribution, perform probability distribution on the state space of each node, and obtain an initial parameter matrix;
[0105] Construct a variational lower bound function for the initial parameter matrix, and set the dimension of the variational parameters to be equal to the number of intention nodes to obtain the variational objective function;
[0106] The variational objective function is input into the variational inference module for optimization. The target posterior probability distribution is obtained by iterative calculation by minimizing the variational free energy.
[0107] Perform parameter sampling on the target posterior probability distribution and count the number of state transitions between adjacent time series nodes to obtain the state transition frequency matrix;
[0108] The maximum likelihood probability is calculated according to the state transition frequency matrix, the transition probability of each state pair is normalized to obtain the state transition rules, and the state transition rules are sorted according to the corresponding relationship of the intention nodes to obtain the intention state transition probability matrix.
[0109] Specifically, according to the nodes and their associations in the intent migration graph, directed connections are constructed according to the timing dependency logic to generate a timing dependency graph. Assume that the node set in the intent migration graph is {I1,I2,…,I n}, where I i Represents the i-th intention node. By analyzing the probability distribution matrix P∈R n×n , determine the transfer relationship between nodes. If P ij Indicates intention i To Intention I j The transition probability is , then add a line from node I to the timing dependency graph i To Node I j The directed edge with weight P ij For example, if P 12 =0.8 and P 23 =0.6, then nodes I1→I2→I3 form a timing path. In the constructed timing dependency graph, each node is regarded as a random variable node, and the state space of the random variable node represents the possible state of the intention node. For example, for node I1, its state space is defined as {s1, s2, …, s k}, where k is the number of states. A conditional probability distribution is established between each random variable node and its parent node to obtain a Bayesian network structure. Each conditional probability distribution of the Bayesian network is expressed as:
[0110] P(I i |Pa(I i ))=f(I i ,Pa(I i ),Θ);
[0111] Among them, Pa(I i ) represents node I i The parent node set of I is Θ, the parameter set of the Bayesian network, and f is the function form of the conditional probability. The conditional probability distribution in the Bayesian network structure is initialized to a uniform distribution. The purpose of the uniform distribution is to assign equal probabilities to all possible states in the initial stage, thereby avoiding any prior bias. For each node I i Status j , whose initial probability is:
[0112]
[0113] This constructs the initial parameter matrix Θ0, where each element Θ ij Represents the probability of going from a state of the parent node to the current node state. In order to optimize the parameters of the Bayesian network, a variational lower bound function is constructed for the initial parameter matrix Θ0. The variational lower bound function is used to approximate the posterior distribution of the Bayesian network and is defined as:
[0114]
[0115] Where q(Θ) is a variational distribution used to approximate the posterior distribution P(Θ|I). The dimension of the variational distribution is set equal to the number of intent nodes n to ensure that each node has an independent optimization variable. Construct the variational objective function. Input the variational objective function into the variational inference module and iteratively update the variational parameters through optimization. Progressively approach the posterior probability distribution of the target:
[0116]
[0117] Each iteration adjusts the parameter matrix Θ to improve the Bayesian network's ability to fit the data. After the optimization is completed, the target posterior probability distribution is sampled. The purpose of sampling is to generate multiple possible state transition sequences, and through these sequences, the number of state transitions between adjacent time series nodes is counted to obtain the state transition frequency matrix C∈R k×k Among them, C ij Indicates state s i To state s j The transition frequency. Calculate the maximum likelihood probability based on the state transition frequency matrix and normalize the transition frequency to a probability distribution:
[0118]
[0119] After normalization, the state transition rules are generated and organized into the intention state transition probability matrix P according to the corresponding relationship of the intention nodes. final The element P in this matrix ij Intent I i Transfer to Intention I j The final probability of .
[0120] In a specific embodiment, the process of executing step 300 may specifically include the following steps:
[0121] The fusion intent feature matrix is input into the encoder of the improved Transformer network model for feature encoding. The encoder contains 6 encoding layers, each of which includes a multi-head self-attention sublayer, a first feedforward neural network sublayer, and a contextual intent understanding sublayer to obtain an encoded feature sequence.
[0122] The encoded feature sequence is input into the decoder of the improved Transformer network model for sequence decoding. The decoder contains 6 decoding layers, each of which includes a masked self-attention sublayer, a cross-attention sublayer, and a second feedforward neural network sublayer to obtain the initial dialogue generation probability distribution;
[0123] The intention state transition probability matrix is input into the intention constraint network of the improved Transformer network model for processing. The intention constraint network includes three fully connected layers to obtain the intention constraint vector.
[0124] The KL divergence calculation is performed on the probability distribution of the initial dialogue generation and the intention constraint vector to obtain the intention Figure 1 Causative loss value;
[0125] The initial dialogue generation probability distribution is input into the template matching network of the improved Transformer network model. The template matching network includes a bidirectional long short-term memory layer and a dot product attention layer to obtain a template matching similarity score.
[0126] The template matching cross entropy loss is calculated based on the template matching similarity score, and the template matching cross entropy loss is compared with the Figure 1 The consistency loss values are weighted combined to obtain the multi-task joint loss function;
[0127] Back propagation calculation is performed based on the multi-task joint loss function to obtain the optimized network parameters, which are then loaded into the improved Transformer network model for model parameter optimization to obtain the dialogue generation model.
[0128] Specifically, the fusion intent feature matrix is input into the encoder part of the improved Transformer network model for feature encoding. Assume that the fusion intent feature matrix is represented as F∈R n×d , where n is the sequence length of the feature matrix and d is the dimension of each feature. The encoder consists of 6 stacked encoding layers, each of which contains a multi-head self-attention sublayer, a first feedforward neural network sublayer, and a contextual intent understanding sublayer. The multi-head self-attention sublayer calculates the correlation between each feature and the rest of the features using the following formula.
[0129]
[0130] Where, Q = FW Q , K=FW K 、V=FW V , respectively, the query, key, and value matrices, is the trainable weight matrix, d k is a scaling factor, usually taken as the square root of d / head. The results of multiple attention heads are concatenated and linearly transformed to obtain the multi-head self-attention output:
[0131] MultiHead(Q,K,V)=Concat(head1,…,head h )W O ;
[0132] in is the linear transformation matrix. The feedforward neural network sublayer performs a point-by-point nonlinear transformation on each feature vector:
[0133] FFN(x)=max(0,xW1+b1)W2+b2;
[0134] in, is the weight matrix, b1, b2 are bias terms, d ff is the dimension of the hidden layer. The contextual intent understanding sublayer further introduces global context information to enhance the semantic expression ability of the feature matrix. After 6 layers of stacking, the output encoding feature sequence E∈R n×d . The encoded feature sequence E is input into the decoder of the improved Transformer network model for sequence decoding. The decoder also consists of 6 decoding layers, each of which contains a masked self-attention sublayer, a cross-attention sublayer, and a second feedforward neural network sublayer. The masked self-attention sublayer ensures the causality of the decoding process by masking the data of future time steps during the decoding process. The formula is the same as the self-attention calculation in the encoder. The cross-attention sublayer associates the state of the decoder with the encoded feature sequence through the following formula:
[0135] CrossAttention(Q,K,V)=Attention(Q d ,K e ,V e );
[0136] Among them, Q d ,K e ,V e are the query matrix of the current state of the decoder and the key and value matrices of the encoded feature sequence. After decoding is completed, the initial dialogue generation probability distribution P is generated. g , each row represents the generation probability of each word at the current time step. At the same time, the intention state transition probability matrix P t ∈R c×c Input into the intention constraint network, P t The element P ij represents the transition probability from intention i to intention j. The intention constraint network consists of three fully connected layers, which generates the intention constraint vector v through linear transformation. t ∈R d :
[0137] v t =ReLU(P t W1+b1)W2+b2;
[0138] Where W1, W2 are weight matrices, b1, b2 are biases. Generate probability distribution P for the initial dialogue g and the intention constraint vector v t The KL divergence is calculated to measure the deviation between the generated distribution and the constraints:
[0139]
[0140] The result of KL divergence is used as Figure 1 The consistency loss value. g The template matching network is input for evaluation. The template matching network consists of a bidirectional LSTM and a dot product attention layer. The bidirectional LSTM is used to extract the context information of the generated sequence, and then the dot product calculation is performed with the predefined template to obtain the matching similarity score s. template . Calculate the template matching cross entropy loss based on the matching score:
[0141]
[0142] Among them, y i is the true label of the template. Figure 1 Consistency loss value Cross entropy loss with template matching Weighted combination, we get the multi-task joint loss function:
[0143]
[0144] Among them, α and β are weighting coefficients that control the relative importance of the two losses. The multi-task joint loss function is minimized through the back-propagation algorithm to optimize the network parameters. The updated parameters are loaded into the improved Transformer network model to complete the training of the dialogue generation model. After the above steps, the model can comprehensively consider the intention state transfer and template matching in the generation process and generate high-quality dialogue responses that conform to the context semantics.
[0145] In this embodiment, back propagation calculation is performed based on a multi-task joint loss function to obtain optimized network parameters, and the optimized network parameters are loaded into an improved Transformer network model for model parameter optimization to obtain a dialogue generation model, including: dividing the improved Transformer network model into an encoder network and a decoder network, and initializing 100 search particles in the parameter space of the encoder network and the decoder network respectively, each particle including a position vector and a velocity vector, to obtain an initial particle swarm; calculating the fitness value of each particle in the initial particle swarm based on the multi-task joint loss function, setting the particle position with the best fitness value as the global optimal position, setting the optimal position experienced by each particle as the individual optimal position, to obtain an initial optimal position set; fixing the parameters of the decoder network to the initial parameter values, performing parameter search on the particle swarm of the encoder network, wherein the particle velocity update weight during the parameter search is 0.7, the individual cognitive factor is 1.5, and the group cognitive factor is 1.5, to obtain an encoder Parameter search results; Based on the encoder parameter search results, the encoder network parameters are configured, and gradient backpropagation and parameter update are performed through the multi-task joint loss function to obtain the encoder optimization parameters; The parameters of the encoder network are fixed as the encoder optimization parameters, and the particle swarm of the decoder network is searched for parameters. During the parameter search process, the particle speed update weight is 0.7, the individual cognitive factor is 1.5, and the group cognitive factor is 1.5 to obtain the decoder parameter search results; Based on the decoder parameter search results, the decoder network parameters are configured, and gradient backpropagation and parameter update are performed through the multi-task joint loss function to obtain the decoder optimization parameters; The encoder optimization parameters and the decoder optimization parameters are integrated and input into the improved Transformer network model for global parameter optimization. After 50 rounds of iterative training, the learning rate is set to 0.001 to obtain the global optimization parameters; The global optimization parameters are loaded into the improved Transformer network model, the model parameter configuration is completed, and the dialogue generation model is obtained.
[0146] In a specific embodiment, the process of executing step 400 may specifically include the following steps:
[0147] Perform word segmentation and vector mapping on the second user text newly input into the intelligent customer service system to obtain a second word vector input sequence, and perform primary and secondary intention feature extraction and feature fusion on the second word vector input sequence to obtain a fused input feature representation;
[0148] The fused input feature representation is input to the encoder in the dialogue generation model to perform feature encoding to obtain the feature encoding result, and the state transition prediction of the feature encoding result is performed based on the intention state transition probability matrix to obtain the target intention vector for the next round of dialogue;
[0149] Constrain the decoding process of the dialogue generation model based on the target intention vector of the next round of dialogue to obtain multiple candidate response sequences;
[0150] Input multiple candidate response sequences into the intent matching unit, calculate the cosine similarity between each candidate response sequence and the target intent vector to obtain an intent relevance score, and perform template matching calculation on the multiple candidate response sequences to obtain a template matching score;
[0151] The intention relevance score and the template matching score are weightedly summed to obtain the comprehensive score of each candidate response sequence. Multiple candidate response sequences are sorted based on the comprehensive score, and the candidate response sequence with the highest score is selected as the response text that meets the conversation intent.
[0152] Specifically, the second user text input by the user is obtained from the intelligent customer service system, and the text is segmented and converted into a word unit sequence. The word unit sequence after segmentation is vectorized, and each word unit is mapped to a high-dimensional vector using a pre-trained word vector model (such as Word2Vec or GloVe). Assuming that the dimension of the word vector is d, each word unit is mapped to v i ∈R d , where v i is the vector representation of the i-th word. The vector representation of the entire input sequence is V = [v1, v2, …, v n ], where n is the number of tokens. The second word vector is input into the sequence V for primary and secondary intent feature extraction and feature fusion. The primary intent feature extraction captures the key semantic relationship in the sequence through a multi-head attention mechanism, and the calculation formula is as follows:
[0153]
[0154] Where Q = VW Q , K = VW K , V=VW V , respectively, the query, key, and value matrices, is a trainable weight matrix, d k=d / head is the scaling factor. The outputs of multiple attention heads are concatenated and transformed linearly to obtain the main image feature matrix F main At the same time, the sub-intention feature extraction captures the contextual dependency of the sequence through bidirectional LSTM. The forward LSTM calculates the time sequence features from left to right, and the backward LSTM calculates the time sequence features from right to left. The two are concatenated to obtain the sub-intention feature matrix F sub ∈R n×2h , where h is the dimension of the hidden layer. main and sub-intention feature F sub The final fused input feature representation F is generated by weighted fusion fusion :
[0155] F fusion =αF main +βF sub ;
[0156] Among them, α and β are adjustable parameters used to balance the weights of primary and secondary features. The fused input feature is represented by F fusion Input to the encoder of the dialogue generation model for feature encoding, generating the encoding feature result E∈R n×d On this basis, combined with the intention state transition probability matrix P intent ∈R c×c Perform state transition prediction on the encoded features to generate the target intention vector v for the next round of dialogue target .P intent The element P ij represents the probability of transferring from intention i to intention j. The prediction process is implemented by the following formula:
[0157] v target =softmax(EW t +P intent );
[0158] Among them, W t is a trainable weight matrix. In the decoding stage, the target intention vector v target Used to constrain the generation process of the decoder. The decoder generates multiple candidate reply sequences {S1, S2, …, S k}, where each candidate sequence contains possible generated text. These candidate response sequences are input into the intent matching unit and similarity is calculated with the target intent vector. The similarity is measured using cosine similarity, which is calculated as follows:
[0159]
[0160] Among them, S iis the feature representation of the candidate reply sequence, and ∥·∥ is the norm of the vector. The obtained intent relevance score reflects the degree of match between the reply and the target intent. At the same time, template matching is performed on the candidate reply sequence to calculate the matching score with the predefined template. Template matching is completed through the dot product attention layer to generate the matching score score template The comprehensive score of each candidate reply sequence is calculated by combining the intent relevance score and the template matching score:
[0161] score combined =γ·similarity(S i ,v target )+δ·score template ;
[0162] Among them, γ and δ are weighting coefficients. The candidate response sequences are sorted according to the comprehensive scores, and the sequence with the highest score is selected as the final response.
[0163] The above describes the intelligent customer service dialogue generation method based on user intent recognition in the embodiment of the present application. The following describes the intelligent customer service dialogue generation system 10 based on user intent recognition in the embodiment of the present application. Figure 2 In the embodiment of the present application, an embodiment of an intelligent customer service dialogue generation system 10 based on user intent recognition includes:
[0164] An acquisition module 11 is used to acquire a first user text input from the intelligent customer service system, and perform feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix;
[0165] An analysis module 12 is used to perform intention migration analysis and dynamic Bayesian analysis on the dialogue sequence data in the historical dialogue corpus based on the fusion intention feature matrix to obtain an intention state transition probability matrix;
[0166] A training module 13 is used to input the fusion intention feature matrix and the intention state transition probability matrix into the improved Transformer network model for joint training to obtain a dialogue generation model;
[0167] The generation module 14 is used to decode and generate the second user text newly input into the intelligent customer service system based on the dialogue generation model, and use the intention state transition probability matrix as a constraint condition to obtain a reply text that meets the dialogue intention.
[0168] Through the collaboration of the above components, a two-layer intent feature extraction network is designed to extract the main intent features and the secondary intent features respectively, which can fully capture the intent information in the user text and improve the accuracy and robustness of intent recognition. The improved Transformer encoder structure is adopted, and the contextual intent understanding sublayer and residual connection mechanism are introduced to achieve the effective fusion of the main and secondary intent features, and enhance the model's ability to understand complex intents. The parameter learning method based on the intent transfer graph and the dynamic Bayesian network effectively models the dynamic characteristics of intent transfer and provides accurate intent constraints for dialogue generation. The intent recognition and dialogue generation tasks are jointly optimized, and the speech template matching mechanism is combined to ensure the professionalism and coherence of the generated replies. A decoding constraint mechanism based on the intent state transition probability matrix is designed to ensure that the generated content conforms to the expected intent transfer law while ensuring the fluency of the reply. The candidate replies are screened by a comprehensive scoring mechanism, which comprehensively considers the intent relevance and template matching degree, and improves the quality and reliability of the system replies.
[0169] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0170] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0171] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating intelligent customer service dialogue based on user intention recognition, characterized in that: The method comprises: Obtaining a first user text input from the intelligent customer service system, and performing feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix; Based on the fusion intention feature matrix, intention migration analysis and dynamic Bayesian analysis are performed on the dialogue sequence data in the historical dialogue corpus to obtain an intention state transition probability matrix; Inputting the fused intent feature matrix and the intent state transition probability matrix into an improved Transformer network model for joint training to obtain a dialogue generation model; Based on the dialogue generation model, the second user text newly input into the intelligent customer service system is decoded and generated, and the intention state transition probability matrix is used as a constraint condition to obtain a reply text that meets the dialogue intention.
2. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 1, characterized in that: The step of obtaining the first user text input from the intelligent customer service system, and performing feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix includes: Obtaining a first user text input from the intelligent customer service system, performing word segmentation processing on the first user text to obtain a word unit sequence, and inputting the word unit sequence into a Word2Vec word vector model for vector mapping to obtain a first word vector input sequence; Input the first word vector input sequence into the main intention feature extraction layer for processing, wherein the main intention feature extraction layer includes 8 attention heads, each of which has a dimension of 64, and obtains 8 groups of attention head feature vectors; The eight groups of attention head feature vectors are concatenated and dimensionality reduction is performed through a linear transformation layer to obtain the main feature data of the user's intention; Input the first word vector input sequence into the secondary intent feature extraction layer for processing, wherein the secondary intent feature extraction layer includes a forward long short-term memory unit and a backward long short-term memory unit, and the hidden layer dimension is 256, to obtain a bidirectional feature sequence; Performing time-sequential splicing on the bidirectional feature sequence to obtain auxiliary feature data of user intention; Attention fusion is performed on the main feature data of the user intention and the auxiliary feature data of the user intention to obtain a fused intention feature matrix.
3. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 2, characterized in that: The attention fusion of the user intention main feature data and the user intention auxiliary feature data to obtain a fused intention feature matrix includes: Input the user intention main feature data into the first self-attention sublayer to perform dot product attention calculation to obtain a main feature attention weight matrix, and input the user intention auxiliary feature data into the second self-attention sublayer to perform dot product attention calculation to obtain an auxiliary feature attention weight matrix; Performing matrix concatenation on the main feature attention weight matrix and the auxiliary feature attention weight matrix to obtain a combined feature weight matrix, and performing a nonlinear transformation on the combined feature weight matrix to obtain an initial fused feature matrix; Performing residual connection on the initial fusion feature matrix and the main feature data of the user's intention, and performing normalization processing through a first-layer normalization unit to obtain a main image fusion matrix; Performing a residual connection on the main intention fusion matrix and the auxiliary feature data of the user intention, and performing normalization processing through a second-layer normalization unit to obtain an auxiliary intention fusion matrix; The main intention fusion matrix and the auxiliary intention fusion matrix are input into a feedforward neural network for feature enhancement to obtain an enhanced feature matrix, and the enhanced feature matrix is input into a feature fusion layer for processing to obtain a fused intention feature matrix.
4. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 1, characterized in that: Based on the fusion intention feature matrix, the intention migration analysis and dynamic Bayesian analysis are performed on the dialogue sequence data in the historical dialogue corpus to obtain the intention state transition probability matrix, including: Segmenting the dialogue sequence data in the historical dialogue corpus according to the dialogue identifier to obtain a plurality of dialogue subsequences; Based on the fusion intention feature matrix, feature matching is performed on each dialogue text in the multiple groups of dialogue subsequences to obtain a dialogue intention label sequence; Performing frequency statistics on the intention labels of adjacent dialogue turns in the dialogue intention label sequence to obtain an intention transfer count matrix, and dividing the count values in the intention transfer count matrix by the sum of the corresponding rows for normalization to obtain an intention transfer probability distribution matrix; Constructing an initial intention node graph based on the intention transfer probability distribution matrix, wherein each node in the initial intention node graph corresponds to an intention category; Establishing directed connections between nodes according to the probability values in the intention transfer probability distribution matrix, using the probability values as connection weights, and obtaining a weighted intention transfer graph structure; Retain the connection paths whose weights are greater than a preset threshold in the weighted intention transfer graph structure to obtain an optimized intention transfer graph structure, and associate each intention node in the optimized intention transfer graph structure with a customer service speech template to obtain an intention migration graph; The intention transition graph is input into a dynamic Bayesian network for parameter learning to obtain an intention state transition probability matrix.
5. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 4, characterized in that: The intention transition graph is input into a dynamic Bayesian network for parameter learning to obtain an intention state transition probability matrix, including: Directed connections are constructed for the nodes in the intention migration graph according to the timing dependency relationship to obtain a timing dependency graph; Setting each node in the timing dependency graph as a random variable node, establishing a conditional probability distribution between the random variable nodes, and obtaining a Bayesian network structure; Initializing the conditional probability distribution in the Bayesian network structure to a uniform distribution, performing probability distribution on the state space of each node, and obtaining an initial parameter matrix; Constructing a variational lower bound function for the initial parameter matrix, and setting the dimension of the variational parameter to be equal to the number of intention nodes, to obtain a variational objective function; The variational objective function is input into the variational inference module for optimization and solution, and the target posterior probability distribution is obtained by iterative calculation in a manner of minimizing variational free energy; Parameter sampling is performed on the target posterior probability distribution, and the number of state transitions between adjacent time series nodes is counted to obtain a state transition frequency matrix; The maximum likelihood probability is calculated according to the state transition frequency matrix, the transition probability of each state pair is normalized to obtain the state transition rule, and the state transition rule is sorted according to the corresponding relationship of the intention node to obtain the intention state transition probability matrix.
6. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 1, characterized in that: The step of inputting the fusion intention feature matrix and the intention state transition probability matrix into an improved Transformer network model for joint training to obtain a dialogue generation model includes: Inputting the fusion intent feature matrix into the encoder of the improved Transformer network model for feature encoding, the encoder comprises 6 encoding layers, each encoding layer comprises a multi-head self-attention sublayer, a first feedforward neural network sublayer and a contextual intent understanding sublayer, to obtain an encoded feature sequence; Inputting the encoded feature sequence into the decoder of the improved Transformer network model for sequence decoding, wherein the decoder comprises 6 decoding layers, each decoding layer comprises a masked self-attention sublayer, a cross-attention sublayer and a second feedforward neural network sublayer, and obtaining an initial dialogue generation probability distribution; Inputting the intention state transition probability matrix into the intention constraint network of the improved Transformer network model for processing, wherein the intention constraint network includes three fully connected layers, to obtain an intention constraint vector; Perform KL divergence calculation on the initial dialogue generation probability distribution and the intention constraint vector to obtain an intention consistency loss value; Inputting the initial dialogue generation probability distribution into the template matching network of the improved Transformer network model, wherein the template matching network includes a bidirectional long short-term memory layer and a dot product attention layer, to obtain a template matching similarity score; Calculating a template matching cross entropy loss based on the template matching similarity score, and weightedly combining the template matching cross entropy loss with the intent consistency loss value to obtain a multi-task joint loss function; Back propagation calculation is performed based on the multi-task joint loss function to obtain optimized network parameters, and the optimized network parameters are loaded into the improved Transformer network model to optimize the model parameters to obtain a dialogue generation model.
7. The method for generating intelligent customer service dialogue based on user intention recognition according to claim 1, characterized in that: The method of decoding and generating a second user text newly input in the intelligent customer service system based on the dialogue generation model, and taking the intention state transition probability matrix as a constraint condition to obtain a reply text that meets the dialogue intention, includes: Performing word segmentation and vector mapping on the second user text newly input into the intelligent customer service system to obtain a second word vector input sequence, and performing primary and secondary intention feature extraction and feature fusion on the second word vector input sequence to obtain a fused input feature representation; Inputting the fused input feature representation into the encoder in the dialogue generation model for feature encoding to obtain a feature encoding result, and performing state transition prediction on the feature encoding result based on the intention state transition probability matrix to obtain a target intention vector for the next round of dialogue; Constraining the decoding process of the dialogue generation model based on the target intention vector of the next round of dialogue to obtain multiple candidate response sequences; Input the multiple candidate response sequences into the intention matching unit, calculate the cosine similarity between each candidate response sequence and the target intention vector to obtain an intention relevance score, and perform template matching calculation on the multiple candidate response sequences to obtain a template matching score; The intention relevance score and the template matching score are weightedly summed to obtain a comprehensive score for each candidate reply sequence, and the multiple candidate reply sequences are sorted based on the comprehensive score, and the candidate reply sequence with the highest score is selected as the reply text that meets the conversation intent.
8. An intelligent customer service dialogue generation system based on user intention recognition, characterized in that: The system is used to execute the intelligent customer service dialogue generation method based on user intent recognition according to any one of claims 1 to 7, comprising: An acquisition module, used to acquire a first user text input from the intelligent customer service system, and perform feature decomposition and attention fusion on the first user text to obtain a fusion intention feature matrix; An analysis module, configured to perform intention migration analysis and dynamic Bayesian analysis on the dialogue sequence data in the historical dialogue corpus based on the fused intention feature matrix to obtain an intention state transition probability matrix; A training module, used for inputting the fusion intention feature matrix and the intention state transition probability matrix into an improved Transformer network model for joint training to obtain a dialogue generation model; A generation module is used to decode and generate a newly input second user text in the intelligent customer service system based on the dialogue generation model, and use the intention state transition probability matrix as a constraint condition to obtain a reply text that meets the dialogue intention.
Citation Information
Cited By
AI multi-technology fusion intelligent dialogue model construction method and system for old people
CN120317381A
AI multi-technology fusion intelligent dialogue model construction method and system for elderly care
CN120317381B
Man-machine multi-round interaction method and device facing visual image
CN120353959A
Multi-round dialogue model training method and device and related equipment
CN120687567A
Natural language interaction intelligent customer service system
CN120893583A