Method and system for context semantic extraction and intention recognition in voice conversations

By combining sparse hybrid expert Transformer network and ordered graph neural network, dynamic update of dialogue context cache and deeply fusion dialogue semantic graphs, the limitations of existing speech dialogue systems in context understanding and intent recognition are solved, and more efficient and accurate speech dialogue processing is achieved.

CN119862891BActive Publication Date: 2025-05-30ZHEJIANG XINBA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510345684.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-30
Estimated Expiration
2045-03-24

AI Technical Summary

Technical Problem

The existing speech dialogue system has limitations in context semantic extraction and intent recognition, and cannot fully understand the diversity and uncertainty of speech input, lacks the utilization of in-depth dialogue historical information, and consumes a lot of computing and storage resources.

Method used

The sparse hybrid expert Transformer network and ordered graph neural network are used to dynamically update the dialogue context cache, deeply fusion dialogue semantic graphs, and adjust the intent recognition results through the expert subnet to improve the system's context understanding and accuracy of intent recognition.

Benefits of technology

It improves the processing capability and dialogue fluency of the voice dialogue system, enhances the robustness and generalization ability of complex dialogue environments, reduces information loss and misjudgment problems, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119862891B_ABST
    Figure CN119862891B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for context semantic extraction and intention recognition in voice conversations, including the following steps: S1, obtaining a voice input signal, converting it into text data, and performing preprocessing; S2, constructing a dialogue context cache and updating the cache content; S3, constructing a sparse mixture of experts Transformer network and performing semantic encoding; S4, calculating the semantic dependency relationship between dialogue statements using an ordered graph neural network; S5, calculating the intention probability distribution and adjusting the weights of the sparse mixture of experts Transformer network and the ordered graph neural network; S6, updating the dialogue state and generating a system response. The present invention combines a sparse mixture of experts Transformer network and an ordered graph neural network to achieve context semantic extraction and intention recognition in voice conversations, and has the advantages of accurate semantic understanding, high computational efficiency, and strong context awareness ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic extraction and recognition, and particularly to a method and system for context semantic extraction and recognition of intentions in voice conversations. Background Art

[0002] With the rapid development of artificial intelligence technology, voice dialogue systems have been widely used in various intelligent devices, especially in the fields of intelligent assistants, intelligent customer service, and speech recognition. The core goal of a voice dialogue system is to enable a machine to understand and generate speech through natural language processing technology, and then interact with users efficiently and smoothly. Existing voice dialogue systems usually rely on modules such as speech recognition, semantic understanding, and intention recognition. The speech recognition module converts the user's speech signal into text, and the semantic understanding module extracts the user's intention based on the converted text. Through this process, the system can respond to the user's needs and thus achieve various intelligent operations. However, in practical applications, despite the significant progress made by these technologies, existing voice dialogue systems still face many challenges, especially in context understanding and intention recognition.

[0003] In the process of context semantic extraction and intention recognition of traditional voice dialogue systems, they often rely on simple lexical matching and fixed pattern matching rules, and this method has great limitations. First, the diversity and uncertainty of speech input signals make it possible for the system to encounter understanding obstacles when processing various speech inputs. Especially in the case of synonyms, polysemous words, accent differences, and noise interference, the accuracy of traditional speech recognition and text analysis technologies is greatly reduced. Second, existing systems usually lack a deep understanding of the context and cannot effectively remember and utilize dialogue history information. Even in some advanced systems, the function of using dialogue history to enhance semantic understanding still has certain limitations. The dialogue content is usually segmented into single, isolated fragments, and the system cannot capture the deep connections between sentences, resulting in the loss of context information and thus affecting the accuracy of intention recognition.

[0004] In addition, traditional intention recognition methods mainly rely on predefined rules or simple classification algorithms, and these methods often perform poorly when dealing with complex and dynamically changing dialogue scenarios. Many existing systems use machine learning methods based on statistical models, which can learn certain semantic rules through a large amount of training data. However, these methods usually cannot fully utilize the implicit information in the dialogue context, especially in the process of multi-round conversations, it is difficult to maintain the coherence of the dialogue and the consistency of semantics.

[0005] In recent years, deep learning models based on Transformer have achieved remarkable results, especially in natural language processing tasks. Through the self-attention mechanism, the Transformer model effectively captures the global dependencies of the input data, significantly improving the text processing ability. Although the Transformer model has achieved good results in some fields, in speech dialogue systems, especially in complex context understanding and intent recognition tasks, it still faces certain challenges. First, traditional Transformer models usually cannot fully process the unstructured information of speech input, especially in aspects such as multi-turn interactions, context dependencies, and dynamic changes in intent involved in conversations, and the performance is still not ideal. Second, the computational and storage overheads of the standard Transformer model are relatively large, especially when dealing with long sequence data, and the problems of computational efficiency and resource consumption are particularly prominent.

[0006] Therefore, there are still many deficiencies in the existing technology during the process of context semantic extraction and intent recognition. First, traditional methods cannot fully consider the deep associations of the context and the semantic dependencies of multi-turn conversations. Second, although deep learning technologies have solved the problem of semantic recognition to a certain extent, how to achieve deep integration and real-time feedback mechanisms while ensuring high efficiency and accuracy is still a difficult point in the current technology.

[0007] Therefore, how to provide a method and system for context semantic extraction and intent recognition in speech dialogue is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0008] An object of the present invention is to propose a method for context semantic extraction and intent recognition in speech dialogue. The present invention combines a sparse mixture-of-experts Transformer network and an ordered graph neural network, and improves the accuracy of context understanding and intent recognition of the speech dialogue system, as well as the processing ability of the system and the fluency of the dialogue by dynamically updating the dialogue context cache, deeply integrating the dialogue semantic graph, and adjusting the intent recognition result through the expert sub-network.

[0009] A method for context semantic extraction and intent recognition in speech dialogue according to an embodiment of the present invention includes the following steps:

[0010] S1. Obtain a speech input signal, convert the speech input signal into text data, and perform preprocessing to extract basic semantic features and generate a standardized text representation;

[0011] S2. Use the standardized text representation to construct a dialogue context cache, store the current dialogue history information, and dynamically update the cache content according to the latest input to form a context-enhanced representation;

[0012] S3. Construct a sparse mixture-of-experts Transformer network to semantically encode the context-enhanced representation using the sparse mixture-of-experts Transformer network, generating a context-encoded representation;

[0013] S4. Based on the context-encoded representation, construct a dialogue semantic graph, calculate the semantic dependency relationship between dialogue utterances using an ordered graph neural network, and adjust the current semantic expression to obtain a context-related representation;

[0014] S5. Calculate the intent probability distribution according to the context-related representation, output the intent recognition result, and feedback the intent recognition result to the sparse mixture-of-experts Transformer network and the ordered graph neural network to adjust the selection weights of the expert sub-networks and the edge weights of the dialogue semantic graph, and output the final intent recognition result;

[0015] S6. Update the dialogue state based on the final intent recognition result and input it into the dialogue manager to generate a system response in combination with the current dialogue information.

[0016] Optionally, the preprocessing includes word segmentation, part-of-speech tagging, named entity recognition, syntactic parsing, and redundancy removal.

[0017] Optionally, S3 specifically includes:

[0018] S31. Construct a sparse mixture-of-experts Transformer network, which includes a gating network, expert sub-networks, a multi-head attention mechanism, and a normalization and residual connection network. The gating network is used to determine which expert sub-networks should process the current input and activates the optimal expert sub-networks using the Top K selection strategy. The expert sub-networks are a feed-forward neural network;

[0019] S32. Obtain the context-enhanced representation , and perform word vector encoding on the context-enhanced representation using Word2Vec to convert the text input into a vector representation:

[0020] ;

[0021] where X represents the vector representation of , represents the embedding vector of the M-th word, M represents the length of the input text, and d represents the embedding dimension;

[0022] S33. Calculate the output of the expert network according to the vector representation of ;

[0023] S34. Take the output Mapped to three different subspaces:

[0024] ;

[0025] where Q, K, and V represent the query, key, and value matrices respectively, , and represent the trained projection matrices used to map the query, key, and value respectively, represents the output of the expert subnetwork;

[0026] S35. Calculate the attention scores based on the query matrix and the key matrix:

[0027] ;

[0028] where A represents the attention scores, T represents the transpose operation, represents the dimension of the key, represents the normalization factor used to control numerical stability and avoid gradient vanishing or explosion;

[0029] S36. Normalize the attention scores A through the Softmax function and apply them to the value matrix to calculate the attention-weighted output:

[0030] ;

[0031] where Attention(Q, K, V) represents the attention-weighted output, Softmax represents the normalization function, and V represents the value matrix;

[0032] S37. Adopt the multi-head attention mechanism to calculate the multi-head attention-weighted output :

[0033] ;

[0034] ;

[0035] where MHA(Q, K, V) represents the multi-head attention-weighted output, , and represent the training parameters of different attention heads, represents the attention-weighted output of the h-th head, represents the linear transformation matrix, and Concat represents the concatenation operation;

[0036] S38. Pass the multi-head attention-weighted output and the output of the expert subnetwork through residual connection and layer normalization, and further map them through a feed-forward neural network to obtain the context encoding representation :

[0037] ;

[0038] ;

[0039] Among them, represents the context encoding representation, FeedForward represents a two-layer feedforward neural network, and LayerNorm represents layer normalization, which is used to standardize data, represents the output of the expert sub-network, represents the output after residual connection and layer normalization.

[0040] Optionally, the S33 specifically includes:

[0041] S331. Calculate the gating weight according to the vector representation X of :

[0042] ;

[0043] Among them, G(X) represents the gating weight, which is a probability distribution used to dynamically activate different expert sub-networks, and Softmax represents normalization, and represent training parameters;

[0044] S332. Adopt the Top K selection strategy to select K optimal expert sub-networks from N expert sub-networks ;

[0045] ;

[0046] Among them, represents the selected K optimal expert sub-networks, and N represents the number of expert sub-networks;

[0047] S333. Calculate the weighted output of the expert sub-network:

[0048] ;

[0049] ;

[0050] Among them, H represents the weighted output of the expert sub-network, represents the weight of the selected i-th expert sub-network, represents the output of the i-th expert sub-network, , , and represent the training parameters of the expert sub-network, and ReLU represents the activation function;

[0051] S334. Use the expert attention mechanism to perform weighted fusion on the results of different expert sub-networks to generate the output of the expert network :

[0052] ;

[0053] ;

[0054] Among them, represents the output of the expert network, represents the weighted output of the i-th expert sub-network, MLP represents the multi-layer perceptron, represents calculating the matching score between expert sub-networks through the multi-layer perceptron, exp represents the natural exponential function, represents the normalized weight, which is used to ensure the rationality of the fusion of expert sub-networks.

[0055] Optionally, the specific steps of S4 include:

[0056] S41. Based on the context encoding representation , construct the initial embedding of each node of the dialogue semantic graph :

[0057] ;

[0058] Among them, represents the initial representation of all dialogue units of the dialogue semantic graph, represents the vector representation of node f, f represents the number of dialogue units, represents the ReLU activation function, represents the training parameter matrix, represents the bias term;

[0059] S42. Calculate the semantic relevance between nodes according to the initial embedding , and calculate the edge weights in the dialogue semantic graph :

[0060] ;

[0061] Among them, represents the edge weight between node i and node j, sim represents the cosine similarity, , and respectively represent the vector representations of node i, node j, and node k;

[0062] S43. Use the calculated edge weights to construct the adjacency matrix B of the dialogue semantic graph:

[0063] ;

[0064] Among them, represents the path deviation value, represents the k-th node of the actual path, represents the k-th node of the planned path, and L represents the total number of path nodes;

[0065] S44. Based on the initial embedding and the adjacency matrix B, use an ordered graph neural network to update the node representation of each layer. The ordered graph neural network is a deep learning network that combines a graph neural network and sequential information;

[0066] S45. Update the node representation V through multiple rounds of iteration until the preset number of iterations is reached or the node representation converges, and obtain the final node representation , and through the final node representation perform a pooling operation to obtain the context-related representation.

[0067] Optionally, the S44 specifically includes:

[0068] S441. Determine the neighbor node set N(i) from the adjacency matrix B, and transmit the information of each node i through the neighbor node j ∈ N(i) to calculate the aggregated representation of the neighbor nodes:

[0069] ;

[0070] Among them, represents the message received by node i from the neighbor nodes in the t-th layer, W represents the training weight matrix for message passing, N(i) represents the neighbor node set of node i, represents the vector representation of node j in the t-th layer, represents the adaptive attention weight:

[0071] ;

[0072] Among them, represents the relative time difference information between node i and node j, represents the relative time difference information between node i and node k, represents the scoring function, using a multi-layer perceptron:

[0073] ;

[0074] Among them, ReLU represents the non-linear activation function, represents the vector concatenation operation, , , and represent the training parameters, The vector representation of node i, The vector representation of node j;

[0075] S442. After receiving the aggregated representation of neighbor nodes, the ordered graph neural network calculates the state update of the current node representation through a gating network:

[0076] ;

[0077] Among them, represents the updated node representation, which is the vector representation of node i in the (t + 1)-th layer, represents the updated path planning scheme, represents the training matrix for node update, represents the bias term, and ReLU represents the activation function;

[0078] S443. Calculate the attention for all nodes in each layer to ensure that information flows in chronological order:

[0079] ;

[0080] ;

[0081] Among them, represents the global ordered attention matrix, Softmax represents normalization, represents the time encoding matrix, which represents the sequential information between different time steps, and respectively represent the query vector and key vector after the vector representation undergoes projection transformation, represents the scaling factor of the embedding dimension, represents the trainable matrix, represents the time step difference between node i and node j, and PositionalEmbedding represents relative position encoding;

[0082] S444. Calculate the node representation after attention:

[0083] .

[0084] Optionally, the specific content of S5 includes:

[0085] S51. Use a fully connected layer to perform feature transformation on the context-related representation to generate a feature representation for intent classification;

[0086] S52. Based on generating a feature representation for intent classification, use a Softmax classification layer to calculate the intent probability distribution P(I) of the current speech input, and take the intent probability distribution P(I) as the intent recognition result;

[0087] S53. Feed the intent recognition result back to the sparse mixture-of-experts Transformer network and the ordered graph neural network to update the parameters:

[0088] ;

[0089] ;

[0090] where, represents the selection weight of the updated expert sub-network, G(x) represents the selection weight of the original expert sub-network, P(I) represents the probability distribution of the intent category, represents the feedback coefficient, which is used to adjust the learning rate of the expert sub-network weights, represents the initial edge weight, represents the updated edge weight, β represents the edge weight adjustment factor, represents the cosine similarity between the recognized intents;

[0091] S54. Output the final intent recognition result:

[0092] ;

[0093] where, represents the final intent recognition result, which is the intent category with the highest probability.

[0094] A system for context semantic extraction and intent recognition in a voice dialogue according to an embodiment of the present invention includes:

[0095] A voice input module, configured to receive a user's voice input and convert it into text data;

[0096] A text preprocessing module, configured to preprocess the text data, extract basic semantic features and generate a standardized text representation;

[0097] A context caching module, configured to store the dialogue history and dynamically update the cache to form a context enhanced representation;

[0098] A sparse mixture-of-experts Transformer network module, configured to perform semantic encoding on the context enhanced representation to generate a context encoding representation;

[0099] A semantic graph construction and ordered graph neural network module, configured to construct a dialogue semantic graph and calculate the semantic dependency relationship between statements to generate a context association representation;

[0100] An intention recognition module, which is used to calculate and output an intention recognition result based on the context - associated representation;

[0101] A feedback module, which is used to feedback the intention recognition result to the sparse mixture - of - experts Transformer network and the semantic graph construction and ordered graph neural network module, adjust the network weights, and output the final intention recognition result;

[0102] A dialogue management module, which is used to update the dialogue state and generate a system response according to the final intention recognition result.

[0103] The beneficial effects of the present invention are as follows:

[0104] First of all, by constructing a dialogue context cache, the system of the present invention can store and dynamically update historical dialogue information, ensuring the consistency of multi - turn conversations, enabling the system to maintain the coherence of dialogue logic in continuous dialogue scenarios, and reducing recognition errors caused by semantic drift. Using the sparse mixture - of - experts Transformer network for semantic encoding and adopting a gating network to activate the optimal expert sub - network, the sparse mixture - of - experts Transformer network can achieve a balance between computational efficiency and expressive ability, improving the processing ability of large - scale corpora and reducing the consumption of computing resources, thus enhancing the real - time performance and scalability of the system.

[0105] Secondly, the present invention uses an ordered graph neural network to model the dialogue semantic graph, ensuring the correct propagation of semantic information between different dialogue turns. Traditional methods usually regard dialogue turns as independent units, ignoring the influence of historical conversations on the current semantics. However, by constructing a dialogue semantic graph, the present invention uses a graph neural network to calculate the semantic dependency relationship between statements and combines an attention mechanism to adjust semantic expressions, enabling the system to better understand the semantic hierarchical relationship and reducing information loss and misjudgment problems in multi - turn conversations of traditional speech dialogue systems. At the same time, the dynamic edge - weight adjustment mechanism of the present invention adaptively optimizes the structure of the dialogue semantic graph according to the intention recognition result, enhancing the robustness and generalization ability of the system in complex dialogue environments.

[0106] In addition, in terms of intention recognition, compared with traditional static intention recognition methods, the present invention feeds the intention recognition result back to the sparse mixture - of - experts Transformer network and the ordered graph neural network, dynamically adjusting the selection weights of expert sub - networks and the edge weights of the semantic graph, thereby enhancing the adaptive ability and context - awareness ability of the dialogue model. The system can improve its adaptability to complex contexts by continuously optimizing the activation strategy of expert sub - networks, making intention recognition more accurate. Especially in long - dialogue, fuzzy intention expression, and dialogue scenarios with strong cross - turn semantic dependencies, the present invention shows obvious advantages.

[0107] Finally, the present invention also performs excellently in terms of the consistency and interpretability of multi-round conversations. Through the construction and dynamic update of the dialogue semantic graph, the system can clearly track and model the evolution of the user's intention, making the reasoning process of the dialogue system more interpretable. Traditional voice dialogue systems often have difficulty providing stable outputs when facing contexts with strong ambiguity and fuzziness. However, the semantic graph modeling and the dynamic adjustment mechanism of the expert network in the present invention enable the system to provide more reasonable intention recognition results when facing complex problems, reducing misrecognition and off-topic responses, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used in conjunction with the embodiments of the present invention to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0109] Figure 1 is a flowchart of a method for context semantic extraction and intention recognition in voice conversations proposed by the present invention;

[0110] Figure 2 is a structural diagram of a sparse mixture-of-experts Transformer network of a method for context semantic extraction and intention recognition in voice conversations proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0111] Now, the present invention will be further described in detail with reference to the accompanying drawings. These drawings are all simplified schematic diagrams, only illustrating the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0112] Referring to Figure 1 and Figure 2 , a method for context semantic extraction and intention recognition in voice conversations includes the following steps:

[0113] S1. Obtain a voice input signal, convert the voice input signal into text data, and perform preprocessing to extract basic semantic features and generate a standardized text representation;

[0114] S2. Use the standardized text representation to construct a dialogue context cache, store the current dialogue history information, and dynamically update the cache content according to the latest input to form a context-enhanced representation;

[0115] S3. Construct a sparse mixture-of-experts Transformer network, and perform semantic encoding on the context-enhanced representation using the sparse mixture-of-experts Transformer network to generate a context encoding representation;

[0116] S4. Construct a dialogue semantic graph based on the context encoding representation, use an ordered graph neural network to calculate the semantic dependency relationship between dialogue statements, and adjust the current semantic expression to obtain a context-related representation;

[0117] S5. Calculate the intention probability distribution according to the context-related representation, output the intention recognition result, and feedback the intention recognition result to the sparse mixture of experts Transformer network and the ordered graph neural network to adjust the selection weights of the expert sub-networks and the edge weights of the dialogue semantic graph, and output the final intention recognition result;

[0118] S6. Update the dialogue state based on the final intention recognition result and input it into the dialogue manager, and generate a system response in combination with the current dialogue information.

[0119] In this embodiment, the preprocessing includes word segmentation, part-of-speech tagging, named entity recognition, syntactic parsing, and redundancy removal.

[0120] In this embodiment, the specific content of S3 is as follows:

[0121] S31. Construct a sparse mixture of experts Transformer network, which includes a gating network, expert sub-networks, a multi-head attention mechanism, and a normalization and residual connection network. The gating network is used to determine which expert sub-networks should process the current input and activates the optimal expert sub-networks using the Top K selection strategy. The expert sub-network is a feed-forward neural network;

[0122] S32. Obtain a context-enhanced representation , and perform word vector encoding on the context-enhanced representation using Word2Vec to convert the text input into a vector representation:

[0123] ;

[0124] where X represents the vector representation of denotes the embedding vector of the M-th word, M represents the length of the input text, and d represents the embedding dimension;

[0125] S33. Calculate the output of the expert network according to the vector representation of ;

[0126] S34. Map the output of the expert network to three different subspaces:

[0127] ;

[0128] Among them, Q, K, and V respectively represent the query, key, and value matrices, , and respectively represent the trained projection matrices, which are respectively used to map the query, key, and value, represents the output of the expert sub-network;

[0129] S35. Calculate the attention score based on the query matrix and the key matrix:

[0130] ;

[0131] Among them, A represents the attention score, T represents the transpose operation, represents the dimension of the key, represents the normalization factor, which is used to control numerical stability and avoid gradient vanishing or explosion;

[0132] S36. Normalize the attention score A through the Softmax function and apply it to the value matrix to calculate the attention-weighted output:

[0133] ;

[0134] Among them, Attention(Q, K, V) represents the attention-weighted output, Softmax represents the normalization function, and V represents the value matrix;

[0135] S37. Adopt the multi-head attention mechanism to calculate the multi-head attention-weighted output :

[0136] ;

[0137] ;

[0138] Among them, MHA(Q, K, V) represents the multi-head attention-weighted output, , and represent the training parameters of different attention heads, represents the attention-weighted output of the h-th head, represents the linear transformation matrix, and Concat represents the concatenation operation;

[0139] S38. Pass the multi-head attention-weighted output and the output of the expert sub-network through residual connection and layer normalization, and further map through the feed-forward neural network to obtain the context encoding representation :

[0140] ;

[0141] ;

[0142] Among them, represents the context encoding representation, FeedForward represents a two-layer feedforward neural network, and LayerNorm represents layer normalization, which is used to standardize data. represents the output of the expert sub-network. represents the output after residual connection and layer normalization.

[0143] In this embodiment, the S33 specifically includes:

[0144] S331. Calculate the gating weight according to the vector representation X of:

[0145] ;

[0146] Among them, G(X) represents the gating weight, which is a probability distribution and is used to dynamically activate different expert sub-networks. Softmax represents normalization. and represent the training parameters;

[0147] S332. Adopt the Top K selection strategy to select K optimal expert sub-networks from N expert sub-networks ;

[0148] ;

[0149] Among them, represents the selected K optimal expert sub-networks, and N represents the number of expert sub-networks;

[0150] S333. Calculate the weighted output of the expert sub-network:

[0151] ;

[0152] ;

[0153] Among them, H represents the weighted output of the expert sub-network. represents the weight of the selected i-th expert sub-network. represents the output of the i-th expert sub-network. , , and represent the training parameters of the expert sub-network, and ReLU represents the activation function;

[0154] S334. Adopt the expert attention mechanism to perform weighted fusion on the results of different expert sub-networks to generate the output of the expert network :

[0155] ;

[0156] ;

[0157] Among them, represents the output of the expert network, represents the weighted output of the i-th expert sub-network, and MLP represents a multi-layer perceptron. represents calculating the matching score between expert sub-networks through a multi-layer perceptron, and exp represents the natural exponential function. represents the normalized weight, which is used to ensure the rationality of the fusion of expert sub-networks.

[0158] In this embodiment, the S4 specifically includes:

[0159] S41. Based on the context encoding representation , construct the initial embedding of each node of the dialogue semantic graph :

[0160] ;

[0161] Among them, represents the initial representation of all dialogue units of the dialogue semantic graph, represents the vector representation of node f, and f represents the number of dialogue units. represents the ReLU activation function, represents the training parameter matrix, represents the bias term;

[0162] S42. Calculate the semantic correlation between nodes according to the initial embedding , and calculate the edge weights in the dialogue semantic graph :

[0163] ;

[0164] Among them, represents the edge weight between node i and node j, and sim represents the cosine similarity. , and respectively represent the vector representations of node i, node j, and node k;

[0165] S43. Use the calculated edge weights to construct the adjacency matrix B of the dialogue semantic graph:

[0166] ;

[0167] Among them, represents the path deviation value, Represents the k-th node of the actual path, Represents the k-th node of the planned path, and L represents the total number of path nodes;

[0168] S44. Based on the initial embedding and the adjacency matrix B, use an ordered graph neural network to update the node representations of each layer. The ordered graph neural network is a deep learning network that combines a graph neural network and sequential information;

[0169] S45. Update the node representation V through multiple rounds of iteration until the preset number of iterations is reached or the node representation converges to obtain the final node representation , and through the final node representation perform a pooling operation to obtain the context-related representation.

[0170] In this embodiment, S44 specifically includes:

[0171] S441. Determine the neighbor node set N(i) from the adjacency matrix B, and transmit the information of each node i through the neighbor node j ∈ N(i) to calculate the aggregated representation of the neighbor nodes:

[0172] ;

[0173] Among them, represents the message received by node i from the neighbor nodes in the t-th layer, W represents the training weight matrix for message passing, N(i) represents the neighbor node set of node i, represents the vector representation of node j in the t-th layer, represents the adaptive attention weight:

[0174] ;

[0175] Among them, represents the relative time difference information between node i and node j, represents the relative time difference information between node i and node k, represents the scoring function, using a multi-layer perceptron:

[0176] ;

[0177] Among them, ReLU represents the non-linear activation function, represents the vector concatenation operation, , , and represent the training parameters, represents the vector representation of node i, represents the vector representation of node j;

[0178] S442. After receiving the aggregated representation of neighbor nodes, the ordered graph neural network calculates the state update of the current node representation through a gating network:

[0179] ;

[0180] Among them, represents the updated node representation, which is the vector representation of node i in the (t + 1)-th layer, represents the updated path planning scheme, represents the training matrix for node update, represents the bias term, and ReLU represents the activation function;

[0181] S443. Calculate the attention for all nodes in each layer to ensure that information flows in chronological order:

[0182] ;

[0183] ;

[0184] Among them, represents the global ordered attention matrix, and Softmax represents normalization, represents the time encoding matrix, which represents the sequential information between different time steps, and respectively represent the query vector and key vector after the vector representation undergoes projection transformation, represents the scaling factor of the embedding dimension, represents the trainable matrix, represents the time step difference between node i and node j, and PositionalEmbedding represents relative position encoding;

[0185] S444. Calculate the node representation after attention:

[0186] .

[0187] In this embodiment, the specific content of S5 includes:

[0188] S51. Use a fully connected layer to perform feature transformation on the context-related representation to generate a feature representation for intent classification;

[0189] S52. Based on the generated feature representation for intent classification, use a Softmax classification layer to calculate the intent probability distribution P(I) of the current speech input, and use the intent probability distribution P(I) as the intent recognition result;

[0190] S53. Feed the intention recognition result back to the sparse mixture-of-experts Transformer network and the ordered graph neural network to update the parameters:

[0191] ;

[0192] ;

[0193] where, represents the selection weight of the updated expert sub-network, G(x) represents the selection weight of the original expert sub-network, P(I) represents the probability distribution of the intention category, represents the feedback coefficient, which is used to adjust the learning rate of the expert sub-network weights, represents the initial edge weight, represents the updated edge weight, β represents the edge weight adjustment factor, represents the cosine similarity between the recognized intentions;

[0194] S54. Output the final intention recognition result:

[0195] ;

[0196] where, represents the final intention recognition result, which is the intention category with the highest probability.

[0197] A system for context semantic extraction and intention recognition in voice conversations, comprising:

[0198] A voice input module, configured to receive the user's voice input and convert it into text data;

[0199] A text preprocessing module, configured to preprocess the text data, extract basic semantic features, and generate a standardized text representation;

[0200] A context caching module, configured to store the conversation history and dynamically update the cache to form a context-enhanced representation;

[0201] A sparse mixture-of-experts Transformer network module, configured to perform semantic encoding on the context-enhanced representation to generate a context encoding representation;

[0202] A semantic graph construction and ordered graph neural network module, configured to construct a conversation semantic graph and calculate the semantic dependency relationship between statements to generate a context association representation;

[0203] An intention recognition module, configured to calculate and output an intention recognition result based on the context association representation;

[0204] A feedback module, which is used to feedback the intent recognition result to the sparse mixture-of-experts Transformer network and the semantic graph construction and ordered graph neural network module, adjust the network weights, and output the final intent recognition result;

[0205] A dialogue management module, which is used to update the dialogue state according to the final intent recognition result and generate a system response.

[0206] Embodiment 1:

[0207] To verify the feasibility of the present invention in implementation, the present invention is applied to the customer service system of a certain intelligent e-commerce platform. This platform receives thousands of user requests every day, covering multiple fields such as product query, order processing, and after-sales service. In order to improve customer satisfaction and reduce the burden on human customer service, the platform decides to introduce a voice-based intelligent customer service system, and this system can handle multi-turn conversations, and accurately understand the intentions and needs of users in the conversations, especially in the aspect of cross-turn conversation context semantic understanding, and can effectively improve the accuracy of intent recognition.

[0208] In the actual application of this e-commerce platform, users interact with the customer service system through voice. For example, when a user asks "I want to check the order situation last month", the system needs to convert this request into text data through voice recognition and extract the semantic information of "check" and "last month's order". Next, the system will rely on its dialogue context cache module to identify whether the user has had relevant order query behaviors before, and then generate an enhanced context representation, and perform semantic encoding on it through the sparse mixture-of-experts Transformer network module.

[0209] This system adopts a sparse mixture-of-experts Transformer network, which can dynamically select relevant expert sub-networks for processing according to the current conversation content, and while improving the computing efficiency, ensure that the context information of the conversation will not be lost in semantic encoding. For example, if a user mentions "order" or "delivery" in a multi-turn conversation, the system can identify these key information and maintain the consistency of the context. This process can effectively avoid the semantic deviation that may occur in traditional systems when processing cross-turn conversations.

[0210] To comprehensively evaluate the effect of the method of the present invention, we conducted an experiment for two weeks in the intelligent customer service system of an e-commerce platform. During the test period, more than 100,000 rounds of user conversations were carried out. The following are the detailed data of the experimental results.

[0211] Table 1 Comparison table of experimental data

[0212] ;

[0213] From the comparison of the intent recognition accuracy, the accuracy of the system of the present invention in each round of conversation is significantly higher than that of the traditional speech recognition system. For example, in the first round of experiments, the accuracy of the traditional system was 78.5%, while the system of the present invention reached 91.2%, with a promotion ratio of 16.7%. This gap was continuously reflected in other conversation rounds, and the average accuracy increased from 80.4% of the traditional system to 92.6% of the system of the present invention, with a promotion amplitude of approximately 15.8%. This result indicates that by introducing the sparse mixture of experts Transformer network and the ordered graph neural network, the system can more accurately understand the conversation context, thereby effectively improving the accuracy of intent recognition.

[0214] In terms of response time, the system of the present invention shows significant advantages compared with the traditional speech recognition system. The average response time of the traditional system is 5.3 seconds, while the average response time of the system of the present invention is reduced to 3.8 seconds, with a promotion ratio of 28.2%. This shortening of the response time means that users can obtain the feedback of the system in a shorter time, improving the interaction efficiency of the system. Especially in high-concurrency conversation scenarios, the optimization of the response time is of great significance for improving the user experience.

[0215] The change in user satisfaction further verifies the effectiveness of the present invention in improving the user experience. Although in some rounds, the user satisfaction of the traditional system is relatively high, the overall performance of the system of the present invention is relatively stable, with an average user satisfaction of 100%. This reflects that by optimizing the intent recognition accuracy and response time, the overall performance of the system of the present invention better meets the needs of users, thereby greatly improving the user experience. Especially in scenarios with frequent conversations and complex content, accurate and rapid feedback can effectively reduce the waiting anxiety of users and improve the trust and satisfaction of users in the system.

[0216] Generally speaking, the present invention not only improves the accuracy of intent recognition, but also effectively reduces the system response time, thereby improving the overall user experience. These improvements make the present invention more competitive in practical applications and can meet the higher requirements of intelligent voice interaction.

[0217] As described above, it is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. A method for extracting contextual semantics and identifying intent in voice conversations, characterized in that: The steps include: S1, obtaining a speech input signal, converting the speech input signal into text data, and performing preprocessing to extract basic semantic features and generate a standardized text representation; S2. Use standardized text representation to build a conversation context cache, store the current conversation history information, and dynamically update the cache content based on the latest input to form a context-enhanced representation; S3, construct a sparse hybrid expert Transformer network, use the sparse hybrid expert Transformer network to semantically encode the context-enhanced representation, and generate a context-encoded representation; S4. Construct a conversation semantic graph based on the context encoding representation, use an ordered graph neural network to calculate the semantic dependency between conversation sentences, and adjust the current semantic expression to obtain the context association representation; S5. Calculate the intent probability distribution based on the context association representation, output the intent recognition result, and feed the intent recognition result back to the sparse hybrid expert Transformer network and the ordered graph neural network, adjust the selection weight of the expert subnetwork and the edge weight of the conversation semantic graph, and output the final intent recognition result; S6. Update the dialog state based on the final intent recognition result, input it into the dialog manager, and generate a system response based on the current dialog information.

2. A method for extracting contextual semantics and identifying intentions in voice conversations according to claim 1, characterized in that: The preprocessing includes word segmentation, part-of-speech tagging, named entity recognition, syntactic parsing and redundancy removal.

3. The method for extracting contextual semantics and identifying intentions in voice conversations according to claim 1, characterized in that: The S3 specifically includes: S31, constructing a sparse hybrid expert Transformer network, wherein the sparse hybrid expert Transformer network includes a gating network, an expert sub-network, a multi-head attention mechanism, and a normalization and residual connection network, wherein the gating network is used to determine which expert sub-networks should process the current input, and activate the optimal expert sub-network using a Top K selection strategy, and wherein the expert sub-network is a feed-forward neural network; S32. Obtaining context-enhanced representation , to enhance the context representation Word2Vec is used for word vector encoding to convert text input into vector representation: ; Among them, X represents The vector representation of represents the embedding vector of the Mth word, M represents the length of the input text, and d represents the embedding dimension; S33, according to The vector represents the output of the expert network ; S34. Output of the expert network Mapped into three different subspaces: ; Where Q, K, and V represent query, key, and value matrices, respectively. , and Represent the trained projection matrices, used to map queries, keys, and values, respectively. represents the output of the expert sub-network; S35. Calculate the attention score based on the query matrix and the key matrix: ; Among them, A represents the attention score, T represents the transposition operation, represents the dimension of the key, Represents the normalization factor, which is used to control numerical stability and avoid gradient disappearance or explosion; S36. Normalize the attention score A through the Softmax function and apply it to the value matrix to calculate the attention weighted output: ; Among them, Attention(Q,K,V) represents the attention weighted output, Softmax represents the normalization function, and V represents the value matrix; S37. Use multi-head attention mechanism to calculate multi-head attention weighted output : ; ; Among them, MHA(Q,K,V) represents the multi-head attention weighted output, , and represents the training parameters of different attention heads, represents the attention weighted output of the h-th head, Represents a linear transformation matrix, and Concat represents a concatenation operation; S38. Output multi-head attention weighted and the output of the expert sub-network After residual connection and layer normalization, and further mapping through a feedforward neural network, the context encoding representation is obtained : ; ; in, represents context encoding representation, FeedForward represents a two-layer feedforward neural network, and LayerNorm represents layer normalization, which is used to standardize data. represents the output of the expert sub-network, Represents the output after residual connection and layer normalization.

4. A method for extracting contextual semantics and identifying intentions in voice conversations according to claim 3, characterized in that: The S33 specifically includes: S331, according to The vector representation of X calculates the gating weights: ; Among them, G(X) represents the gate weight, which is a probability distribution used to dynamically activate different expert sub-networks. Softmax represents normalization. and represents the training parameters; S332, using the Top K selection strategy, select K best expert sub-networks from N expert sub-networks ; ; in, represents the selected K best expert sub-networks, and N represents the number of expert sub-networks; S333. Calculate the weighted output of the expert subnetwork: ; ; Among them, H represents the weighted output of the expert sub-network, represents the weight of the selected i-th expert sub-network, represents the output of the i-th expert sub-network, , , and represents the training parameters of the expert sub-network, and ReLU represents the activation function; S334. Use the expert attention mechanism to perform weighted fusion on the results of different expert sub-networks to generate the output of the expert network : ; ; in, represents the output of the expert network, represents the weighted output of the i-th expert sub-network, MLP represents a multi-layer perceptron, indicates that the matching scores between expert sub-networks are calculated by multi-layer perceptron, exp indicates the natural exponential function, Represents the normalized weight, which is used to ensure the rationality of the expert sub-network fusion.

5. The method for extracting contextual semantics and identifying intentions in voice conversations according to claim 1, characterized in that: The S4 specifically includes: S41. Context-based encoding representation , construct the initial embedding of each node in the conversation semantic graph : ; in, represents the initial representation of all dialogue units in the dialogue semantic graph, represents the vector representation of node f, where f represents the number of dialogue units, represents the ReLU activation function, represents the training parameter matrix, represents the bias term; S42. Calculate the semantic correlation between nodes based on the initial embedding , and calculate the edge weights in the conversation semantic graph : ; in, represents the edge weight between node i and node j, sim represents the cosine similarity, , and Represent the vector representation of node i, node j and node k respectively; S43. Using the calculated edge weights Construct the adjacency matrix B of the conversation semantic graph: ; in, Represents the path deviation value, represents the kth node of the actual path, represents the kth node of the planned path, and L represents the total number of path nodes; S44, based on initial embedding and the adjacency matrix B, and updating the node representation of each layer using an ordered graph neural network, wherein the ordered graph neural network is a deep learning network that combines a graph neural network and sequential information; S45, updating the node representation V through multiple rounds of iterations until a preset number of iterations is reached or the node representation converges, and obtaining the final node representation , represented by the final node Perform pooling operation to obtain contextual association representation.

6. A method for extracting contextual semantics and identifying intentions in voice conversations according to claim 5, characterized in that: The S44 specifically includes: S441. Determine the neighbor node set N(i) by the adjacency matrix B, transmit the information of each node i through the neighbor node j∈N(i), and calculate the aggregate representation of the neighbor nodes: ; in, represents the message received by node i from its neighbor nodes in the tth layer, W represents the training weight matrix of message transmission, N(i) represents the set of neighbor nodes of node i, represents the vector representation of node j in the tth layer, Represents adaptive attention weights: ; in, Represents the relative time difference between node i and node j, Represents the relative time difference between node i and node k, Represents the scoring function, using a multi-layer perceptron: ; Among them, ReLU represents the nonlinear activation function, represents the vector concatenation operation, , , and represents the training parameters, represents the vector representation of node i, represents the vector representation of node j; S442. After receiving the aggregate representation of the neighboring nodes, the ordered graph neural network calculates the state update of the current node representation through the gating network: ; in, represents the updated node representation, which is the vector representation of node i in the t+1th layer. represents the updated path planning solution, represents the training matrix for node updates, represents the bias term, ReLU represents the activation function; S443. Perform attention calculation on all nodes at each layer to ensure that information flows in chronological order: ; ; in, represents the global ordered attention matrix, Softmax represents normalization, represents the time encoding matrix, which represents the sequential information between different time steps, and Respectively represent vector representation The query vector and key vector after projection transformation, represents the scaling factor of the embedding dimension, represents the trainable matrix, Represents the time step difference between node i and node j, and PositionalEmbedding represents the relative position encoding; S444, node representation after calculating attention: 。 7. The method for extracting contextual semantics and identifying intentions in voice conversations according to claim 1, characterized in that: The S5 specifically includes: S51, using a fully connected layer to perform feature transformation on the context association representation to generate a feature representation for intent classification; S52, based on the feature representation generated for intent classification, the Softmax classification layer is used to calculate the intent probability distribution P(I) of the current speech input, and the intent probability distribution P(I) is used as the intent recognition result; S53, feed back the intent recognition results to the sparse hybrid expert Transformer network and the ordered graph neural network, and update the parameters: ; ; in, represents the selection weight of the updated expert sub-network, G(x) represents the selection weight of the original expert sub-network, P(I) represents the probability distribution of the intent category, represents the feedback coefficient, which is used to adjust the learning rate of the expert sub-network weights. represents the initial edge weight, represents the updated edge weight, β represents the edge weight adjustment factor, represents the cosine similarity between the identified intents; S54, output the final intention recognition result: ; in, Represents the final intent recognition result, which is the intent category with the highest probability.

8. A system for extracting contextual semantics and identifying intentions in voice conversations, executing the method for extracting contextual semantics and identifying intentions in voice conversations as claimed in any one of claims 1 to 7, characterized in that: include: A voice input module, used to receive the user's voice input and convert it into text data; The text preprocessing module is used to preprocess text data, extract basic semantic features and generate standardized text representation; A context cache module, which is used to store conversation history and dynamically update the cache to form a context-enhanced representation; A sparse hybrid expert Transformer network module is used to semantically encode the context-enhanced representation and generate a context-encoded representation; Semantic graph construction and ordered graph neural network module, which is used to construct the conversation semantic graph and calculate the semantic dependency between sentences to generate contextual association representation; An intent recognition module, used to calculate and output intent recognition results based on context association representation; Feedback module, used to feed back the intent recognition results to the sparse hybrid expert Transformer network and semantic graph construction and ordered graph neural network modules, adjust the network weights, and output the final intent recognition results; The dialogue management module is used to update the dialogue state and generate system responses based on the final intent recognition results.

Citation Information

Patent Citations

  • Multi-round dialogue reply generation method based on dual-channel semantic enhancement and terminal equipment

    CN115495552A

  • Natural language understanding method and device fusing dialogue context information

    CN116542256A