End-to-End Task-Oriented Dialogue System in the Field of Architecture Based on Hierarchical Memory Network
By adopting a hierarchical memory network in a task-oriented dialogue system, the dialogue history and knowledge base information are separated and stored, and knowledge inference is used using memory network pointers and replication enhancement decoders for knowledge inference, the problem of difficulty in knowledge in existing systems is solved, and a more efficient and low-cost dialogue system is achieved.
Patent Information
- Application Number
- CN202210540052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Existing task-oriented dialogue systems are difficult to perform knowledge inference when storing dialogue history and knowledge base information in the same memory network, and traditional pipeline methods are costly.
The end-to-end building field task-based dialogue system is adopted based on hierarchical memory network, and the dialogue history is encoded through a context encoder. The hierarchical memory network is divided into dialogue memory network and knowledge base memory network, and knowledge network pointer and replication enhancement decoder are used for knowledge reasoning and answer generation.
It realizes effective separation and reasoning of dialogue history and knowledge base information, reduces the cost and complexity of the system, and improves the task completion ability of the dialogue system.
Smart Images

Figure CN114969331B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dialogue systems, and particularly to an end-to-end task-oriented dialogue system in the construction field based on a hierarchical memory network. Background Art
[0002] For a long time, having a virtual voice assistant or chat companion system with sufficient intelligence has seemed illusory and only existed in science fiction movies. Recently, human-computer dialogue has attracted much attention due to its huge potential and attractive commercial value. With the development of big data and deep learning technologies, the goal of creating an automatic human-computer dialogue system as our personal assistant or chat companion is no longer a fantasy. On the one hand, we can easily obtain the "big data" of conversations on the Internet, so that we can learn how to reply to any input. This enables us to build a data-driven human-computer dialogue system. On the other hand, deep learning technologies have been proven to be effective in identifying complex patterns in big data. Currently, dialogue systems are mainly divided into two categories, one is the task-oriented dialogue system, and the other is the non-task-oriented dialogue system, that is, the open-domain dialogue system. The existing task-oriented dialogue systems mainly have the pipeline method and the end-to-end method.
[0003] The task-oriented dialogue system aims to help users complete tasks in a specific field, such as navigation systems, restaurant recommendations, and accommodation reservations. Different from the open-domain dialogue system, the task-oriented dialogue system usually involves information from an external knowledge base.
[0004] In the traditional pipeline method, each module needs to be designed and trained separately, and the cost is very high; in the existing end-to-end solutions, the dialogue history and knowledge base information are stored in the same memory network, and the knowledge base information is represented in the form of knowledge base triples, making it difficult for the memory network to perform knowledge reasoning. Summary of the Invention
[0005] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes an end-to-end task-oriented dialogue system in the construction field based on a hierarchical memory network.
[0006] To achieve the above object of the present invention, the present invention provides an end-to-end task-oriented dialogue system in the construction field based on a hierarchical memory network, including:
[0007] A context encoder: used to encode the dialogue history into a vector The encoded vector representing the nth word in the dialogue;
[0008] A hierarchical memory network: used to store the dialogue history and knowledge base information;
[0009] A memory network pointer: As a condition, query relevant rows S in the knowledge base kb , relevant rows S in the dialogue memory network d and content o in the relevant knowledge base kb ;
[0010] Copy enhancement decoder: used to obtain specific answer information;
[0011] The data output end of the context encoder is connected to the data input end of the memory network pointer and the data input end of the copy enhancement decoder. The data end of the memory network pointer is connected to the data end of the hierarchical memory network. The data output end of the hierarchical memory network is connected to the data input end of the copy enhancement decoder. The data output end of the copy enhancement decoder is connected to the data input end of the hierarchical memory network.
[0012] Furthermore, the context encoder includes:
[0013] First, use the embedding function φ emb to encode each word into a vector, then obtain the context representation of each word in the dialogue history through a single-layer gated recurrent unit, add the context representation to the corresponding position in the dialogue memory network, and use the encoded vector of the nth word in the dialogue as the hidden representation of the dialogue state.
[0014] Furthermore, the hierarchical memory network includes:
[0015] Store the dialogue history in the dialogue memory network, store the external knowledge base in the knowledge base memory network, and represent the knowledge base information by knowledge base rows.
[0016] Furthermore, the hierarchical memory network also includes:
[0017] Each row of the memory network contains historical dialogue information or knowledge base information. For the dialogue memory network, each row contains a word and its corresponding speaker, time, and location information. For the knowledge memory network, each row stores a topic in the knowledge base and its attribute values;
[0018] Each column of the memory network is a specific word that the decoder needs to receive. Store each attribute value of the topic in the relevant column of the knowledge base memory network, and add a special symbol at the end as the end position marker.
[0019] Furthermore, the hierarchical memory network also includes using an end-to-end memory network to model the memory reading process, including the following steps:
[0020] S-1, first, a series of trainable embedding matrices C = {C 1 , C 2,...,C K+1} as Then each word v i,j is encoded into a vector;
[0021] where C 1 represents the first trainable embedding matrix, C 2 represents the second trainable embedding matrix, C K+1 represents the (K + 1)-th trainable embedding matrix, represents the k-th trainable embedding matrix of the word v i,j and v i,j represents the word in the i-th row and j-th column of the memory network;
[0022] S - 2, and then each row of the memory network is represented as where represents the i-th row of the embedded memory matrix R k , represents the k-th embedding vector of the i-th row and j-th column, and |C| represents the number of columns of the matrix C;
[0023] Read the memory through the following formula:
[0024] p k = Softmax(q k R k ) (1)
[0025] where p k represents the probability distribution of the rows of the memory network;
[0026] q k represents the query data for the k-th hop;
[0027] R k represents the embedded memory matrix;
[0028] Softmax(·) represents the Softmax function;
[0029] Read out the content o of the k-th hop of the memory network through the following formula k :
[0030]
[0031] where represents the probability of each row for the k-th inference;
[0032] is the i-th row of the embedded memory matrix R k ;
[0033] The query data for the (k + 1)-th hop can be obtained through the following formula:
[0034] qk+1 = q k + o k (3)
[0035] Where q k represents the query data for the k-th hop;
[0036] o k represents the content of the k-th hop of the memory network.
[0037] Furthermore, the memory network pointer includes:
[0038] S-11, using the encoding vector of the n-th word in the dialogue as the query of the memory network to obtain the probability of each row in the two memory networks:
[0039]
[0040] Where q K represents the query data for the K-th hop;
[0041] represents the transpose of the matrix;
[0042] represents the i-th row of the embedded memory matrix R k ;
[0043] S-12, defining the pointer label of the dialogue memory network by checking whether the word appears in the system answer represents the pointer label of the first row of the dialogue history memory network, represents the pointer label of the second row of the dialogue history memory network, represents the pointer label of the |R|-th row of the dialogue history memory network d ;
[0044] And by checking the rows in the knowledge base memory network that contain the true answer, selecting the row with the largest number of entities to define the pointer label of the knowledge base memory network, and defining the label of the knowledge base memory network as Where represents the pointer label of the first row of the knowledge base memory network, represents the pointer label of the first row of the knowledge base memory network, represents the pointer label of the |R|kb-th row of the knowledge base memory network;
[0045] If the checked word appears in the system answer and the row in the knowledge base memory network contains the true answer, then If not satisfied represents whether each row is selected, 1 for selected, otherwise 0;
[0046] S-13, use the binary cross-entropy loss function Loss between the S label and the S l label to train the memory network pointer. s The loss function of the dialogue memory network is as follows:
[0047]
[0048]
[0049]
[0050]
[0051] s d,i represents the probability of the i-th row of the dialogue history memory network;
[0051] |R| d represents the |R|-th d row of the dialogue history memory network;
[0052] The loss function of the knowledge base memory network is as follows:
[0053]
[0054]
[0055] s kb,i represents the probability of the i-th row of the knowledge base memory network;
[0056] |R| kb represents the |R|-th kb row of the knowledge base memory network;
[0057] For the loss function Loss of the memory network pointer s is defined as follows:
[0058] Loss s = Loss kbs + Loss ds (7)
[0059] where Loss ds represents the loss function of the dialogue memory network;
[0060] Loss kbs represents the loss function of the knowledge base memory network;
[0061] S-14, use the probability of the memory network pointer to adjust each row in the two memory networks:
[0062]
[0063] where is the embedded memory matrix Rk The i-th row of;
[0064] Represents the probability distribution of the i-th row of the memory network.
[0065] Furthermore, the copy-enhanced decoder includes:
[0066] First, generate a rough answer through a rough recurrent neural network model, and then use an entity decoder to query a knowledge base memory E kb and the dialogue history P d to decode the rough answer by querying entities in it to obtain specific answer information.
[0067] Furthermore, the rough recurrent neural network model includes:
[0068] Read from the memory network pointer and the knowledge base content as input:
[0069]
[0070] Where Represents the matrix W 1 And The composed matrix Are concatenated;
[0071] Represents the encoded vector of the n-th word in the dialogue;
[0072] Represents the content of the k-th hop in the knowledge base memory network;
[0073] The meaning of is the decoding state of the 0-th time of the dialogue state;
[0074] W 1 Is a learnable matrix parameter;
[0075] Then replace the entities in the answer with a rough label, and the decoding process is as follows:
[0076]
[0077] Where, Represents the state vector of the t-th decoding;
[0078] Represents the state vector of the (t - 1)-th decoding;
[0079] Represents the word The embedding matrix parameter of;
[0080] GRU(·) is a single-layer gated recurrent unit;
[0081] C 1 is the embedding matrix parameter used by the hierarchical memory network;
[0082] Then, the probability distribution of the word in the current step t can be obtained through the following formula
[0083]
[0084] where W 2 is a learnable matrix parameter;
[0085] Finally, the word distribution probability and the ground truth of the rough answer are used to train the rough recurrent neural network model through the cross-entropy loss function:
[0086]
[0087] where represents the rough answer at the t-th time;
[0088] m represents the total number of steps, i.e., the total number of loops;
[0089] represents the word distribution probability of.
[0090] Furthermore, the entity decoder includes:
[0091] Using an entity from the hierarchical memory network to replace a rough label generated by the rough recurrent neural network model, including the following steps:
[0092] S-111, using the memory reader with the encoded vector of the n-th word in the conversation as the input to obtain the probability of each row and at the current step t, represents the word distribution probability of the knowledge base network at the current step t, represents the word distribution probability of the conversation memory network at the current step t;
[0093] S-112, secondly, the probability of each column in the knowledge base memory network can be obtained through the following formula:
[0094]
[0095] where φ emb (·) is the embedding function;
[0096] n jis the name of the entity type in the j-th column of the knowledge memory network;
[0097] represents the state vector of the t-th encoding;
[0098] S - 113. Finally, the following formula is used to calculate the probability of each entity in the knowledge base memory network:
[0099]
[0100] where, represents the probability distribution of the entity in the knowledge base memory network at the current step t;
[0101] represents the probability of the knowledge base at the t-th step;
[0102] score t represents the score of the attribute at the current step t;
[0103] By looking up the entity v that satisfies v i,j = y t in the i-th row and j-th column, and the corresponding i,j in the memory network pointer that satisfies satisfies to define the label of the entity decoder in the knowledge base memory network:
[0104]
[0105] where max(·) represents taking the maximum value;
[0106] i represents the i-th row of the knowledge base memory network;
[0107] |C| kb represents the number of columns of the knowledge base;
[0108] j represents the j-th column of the knowledge base memory network;
[0109] v i,j represents the entity in the i-th row and j-th column;
[0110] y t represents the true answer;
[0111] represents the pointer label of the i-th row of the knowledge base memory network;
[0112] represents the number of rows in the knowledge base memory network;
[0113] Define the label of the entity decoder of the dialogue memory network as shown in the following formula:
[0114]
[0115] wherein represents the number of rows in the dialogue memory network;
[0116] max(·) represents taking the maximum value;
[0117] v i,0 represents the entity in the i-th row and the 0-th column;
[0118] y t represents the true answer;
[0119] Then the loss function of the entity selector is defined as follows:
[0120]
[0121] wherein represents the probability distribution of the entity in the knowledge base memory network at the current step t;
[0122] represents the label of the entity decoder in the knowledge base memory network;
[0123] represents the probability distribution of the entity in the dialogue memory network at the current step t;
[0124] represents the label of the entity decoder of the dialogue memory network.
[0125] Furthermore, it also includes the loss function Loss:
[0126] Loss = αLoss s + βLoss v + γLoss ent (18)
[0127] where α, β, and γ are all hyperparameters;
[0128] Loss s represents the loss function of the memory network pointer;
[0129] Loss v represents the cross-entropy loss function;
[0130] Loss ent represents the loss function of the entity selector.
[0131] In summary, due to the adoption of the above technical solutions, the present invention can divide the dialogue history and knowledge base information into two memory networks and represent knowledge in the form of knowledge rows, thereby enabling the memory network to smoothly perform knowledge reasoning.
[0132] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0133] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of embodiments in conjunction with the accompanying drawings, in which:
[0134] Figure 1 is a schematic diagram of the hierarchical memory network of the present invention.
[0135] Figure 2 is an overall architecture diagram of the hierarchical memory model of the task-oriented dialogue system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0136] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0137] The overall architecture diagram of the end-to-end task-oriented dialogue system model based on the hierarchical memory network is as Figure 2 shown.
[0138] The dialogue history X = (x 1 , x 2 ,..., x n ) and the knowledge base are used as the model input, and the system answer Y = (y 1 , y 2 ,..., y m ) is used as the model output.
[0139] This model is based on the encoder-decoder architecture and consists of four parts: a context encoder, a hierarchical memory network, a memory network pointer, and a copy-augmented decoder.
[0140] First, the context encoder encodes the dialogue history into a vector representing the dialogue state, that is, the encoded vector of the last word in a dialogue. Then, the memory network pointer queries the relevant rows S in the knowledge base, the relevant rows S kb in the dialogue memory network, and the content o d in the relevant knowledge base with the dialogue state kb as a condition. Finally, the copy-augmented decoder first generates a rough answer through a rough recurrent neural network Sketch RNN model, and then uses an entity decoder to query a knowledge base memory E kband the dialogue history P d Decode the rough answer from the entities in
[0141] 1. Context Encoder
[0142] Use a single-layer gated recurrent unit (GRU) to encode the dialogue history into a hidden state where represents the encoded vector of the first word in the dialogue, represents the encoded vector of the first word in the dialogue, represents the encoded vector of the nth word in the dialogue. First, use the embedding function φ emb to encode each word into a vector, then obtain the context representation of each word in the dialogue history through a single-layer gated recurrent unit, and add the context representation to the corresponding position in the dialogue memory network. And use the last hidden state as the hidden representation of the dialogue state. The dialogue state is the hidden state, and here only the last state is read.
[0143] 2. Hierarchical Memory Network
[0144] Currently, most methods store the dialogue history and knowledge base information in the same memory network and represent the knowledge base information as knowledge base triples. These defects make it difficult for the memory network pointer to perform knowledge reasoning on the memory network. Therefore, a hierarchical memory network is proposed to store the dialogue history and knowledge base information, maintaining the hierarchy of reasoning. As Figure 1 shown in the hierarchical memory network framework, the dialogue history and the external knowledge base are stored in two memory networks (dialogue memory network and knowledge base memory network), and the knowledge base information is represented by knowledge base rows. In the first level, each row of each memory network contains historical dialogue information or knowledge base information. For the dialogue memory network, each row contains a word and its corresponding speaker, time, and location information. For the knowledge memory network, each row stores a topic in the knowledge base and its related attribute values. In the second level, each column is a specific word that the decoder needs to receive. Each attribute value of the topic is stored in the relevant column of the knowledge base memory network, and the attributes of the topic in each domain (dataset) are shown in Table 2. A special word symbol $ is added at the end of both memory networks as the end position marker.
[0145] Table 1 Partial knowledge information of triples in the SMD dataset
[0146]
[0147]
[0148] Table 2 Attributes of Different Datasets
[0149]
[0150] 2.1. Memory Reading
[0151] An end-to-end memory network is adopted to model the memory reading process. Here, the working principle of the K-hop memory network is briefly described. Each word v in the memory network i,j , first, a series of trainable embedding matrices C = {C 1 , C 2 ,..., C K+1} are used as to encode each word into a vector (C K+1 represents the (K + 1)-th trainable embedding matrix C k ∈ C = {C 1 , C 2 ,..., C K+1}), C k (v i,j ) represents the embedding vector encoded by the embedding matrix C; v i,j represents the word at the i-th row and j-th column in the memory network, represents the k-th trainable embedding matrix. Then, each row of the memory network is represented by |R| rows and |C| columns as where is the k-th embedding vector at the i-th row and j-th column, and |·| represents the determinant of the matrix. For a k-hop (k-th training) memory network, it consists of trainable parameter matrices R = {R 1 , R 2 ,..., R K+1}, is the |R|-th row of the embedding memory matrix R k . The memory can be read through the following steps as Figure 1 shown:
[0152] p k = Softmax(q k R k ) (1)
[0153] where q k is the query data for the k-th hop, and p k represents the probability distribution of the rows of the memory network. Equation 1 represents the probability distribution obtained by multiplying the query data for the k-th hop and the embedding matrix R k and passing it through the Softmax function. The content o of the k-th hop of the memory network can be read out through the formula
[0154] 2k :
[0155]
[0156] Among them, represents the probability of each line of the k-th inference, is the i-th row of the embedded memory matrix R k The query data q for the (k + 1)-th hop is updated by the query data q of the k-th hop and the memory network o k of the k-th hop, that is, Equation 3:
[0157] q k+1 = q k + o k (3)
[0158] 3. Memory network pointer
[0159] The function of the memory network pointer is to point to the lines related to the conversation. Using the dialogue state h n enc as the query of the memory network to obtain the probability of each line in the two memory networks:
[0160]
[0161] Define the pointer label of the dialogue memory network by checking whether the corresponding word appears in the system answer where |R| d is the number of rows of the dialogue history memory network. And define the pointer label of the knowledge base memory network by checking which rows in the knowledge base memory network contain the largest number of entities in the true answer. Whether the corresponding knowledge base row in the system answer has the largest number of entities defines the label of the knowledge base memory network as |R| kb is the number of rows of the knowledge memory network. Assuming the above requirements are met, then correspondingly If not met indicates whether each line is selected, 1 for selected, otherwise 0. Then use the binary cross-entropy loss function Loss between the S label and the S l label to train the memory network pointer. Equation 5 is the loss function of the dialogue memory network and Equation 6 is the loss function of the knowledge base memory network, s Then, for the loss function Loss of the memory network pointer
[0162]
[0163]
[0164] Then, for the loss function Loss of the memory network pointer s is defined as shown in Equation (7):
[0165] Loss s = Loss kbs + Loss ds (7)
[0166] Adjust each row in the two memory networks using the probability of the memory network pointer:
[0167]
[0168] where is the i-th row of the embedded memory matrix R k ;
[0169] represents the probability distribution of the i-th row of the memory network;
[0170] 4. Copy-Augmented Decoder
[0171] In the decoding phase, first use the rough recurrent neural network Sketch RNN model to generate a rough answer. Then use the entity decoder to copy entities from the hierarchical memory network in two steps to decode the rough answer.
[0172] 4.1. Rough Recurrent Neural Network Sketch RNN Model
[0173] Use the gated recurrent unit GRU to model the rough recurrent neural network Sketch RNN model. The last dialogue state read from the memory network pointer and the knowledge base content are used as inputs:
[0174]
[0175] In Equation 9, represents the matrix composed of matrix W 1 and concatenated, represents the n-th encoded state of the dialogue state, represents the content of the k-th hop in the knowledge base memory network, means the 0-th decoded state of the dialogue state, W 1 is a learnable matrix parameter, where dec represents decoding, the encoded hidden state; enc represents encoding, the encoded hidden state.
[0176] Replace the entity in the answer with a coarse label. For example, the coarse Recurrent Neural Network Sketch RNN model will generate "@poi is @distance away" instead of "chef_chu_s is 5_miles away". This decoding process is shown in Equation 10 as follows:
[0177]
[0178] where represents the state vector of the t-th decoding, represents the state vector of the (t - 1)-th decoding, represents the embedding matrix parameter of the word , GRU(·) is a single-layer gated recurrent unit, and C 1 is the embedding matrix parameter used by the hierarchical memory network, and is the word generated by the coarse Recurrent Neural Network Sketch RNN model in the previous step. Then, the probability distribution of the word in the current step t can be obtained through Equation 11
[0179]
[0180] W 2 is a learnable matrix parameter.
[0181] Use the word distribution probability and the ground truth of the coarse answer to train the coarse Recurrent Neural Network Sketch RNN model through the cross-entropy loss function:
[0182]
[0183] where represents the t-th coarse answer;
[0184] m represents the total number of steps, i.e., the total number of loops;
[0185] represents the word distribution probability;
[0186] 4.2 Entity Decoder
[0187] When the coarse Recurrent Neural Network Sketch RNN model generates a coarse label, it needs to be replaced with an entity from the hierarchical memory network. For example, when the coarse Recurrent Neural Network Sketch RNN model generates a coarse label @address, it should be replaced with an entity in the memory network, and the type of this entity is the address type. First, the memory reader is also used with the hidden state Obtain the probability for each row in the current step t as the input and the probability of indicating the word distribution probability in the knowledge base network at the current step t indicating the word distribution probability in the dialogue memory network at the current step t. Secondly, the probability for each column in the knowledge base memory network can be obtained through Equation 13:
[0188]
[0189] where φ emb (·) is the embedding function, n j is the entity type name of the j-th column in the knowledge memory network represents the hidden state at the current step t, dec represents decoding, the encoded hidden state; that is represents the state vector of the t-th encoding;
[0190] and represents the score of the attribute of the j-th column at the current step t. Finally, Equation 14 is used to calculate the probability of each entity in the knowledge base memory network:
[0191]
[0192] where represents the probability distribution of the entity in the knowledge base memory network at the current step t represents the probability of the t-th step of the knowledge base, score t represents the score of the attribute at the current step t. By looking up which entity (word) v i,j in the i-th row and j-th column satisfies v i,j = y t , that is, the solved answer is to be compared with the true answer y t ; and the corresponding in the memory network pointer satisfies to define the label of the entity decoder in the knowledge base memory network:
[0193]
[0194] where represents the number of rows in the knowledge base memory network, |C| kb represents the number of columns of the knowledge base. indicates that the word does not appear in the memory network. Similar to the knowledge base memory network, the label of the entity decoder of the dialogue memory network is defined as shown in Equation 16:
[0195]
[0196] Then the loss function of the entity selector is defined as shown in Equation 17:
[0197]
[0198] Where represents the probability distribution of the entity in the knowledge base memory network at the current step t;
[0199] Try to optimize the loss function Loss in Equation 18 to train the entire model, where α, β, γ are hyperparameters:
[0200] Loss = αLoss s + βLoss v + γLoss ent (18)
[0201] In addition, two recording matrices are used to avoid duplicating the same entity. Initially, the matrix elements are all 1. If a word in the memory network has been duplicated, the corresponding position in the matrix will become 0. During decoding, the entity corresponding to the 0 element in the matrix in the memory network will be masked. The entity decoder always selects the word with the highest copying probability from E t kb and P t d in the highest copying probability.
[0202] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.
Claims
1. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network, characterized in that, it includes: Context Encoder: Used to encode the conversation history into a vector Denote the encoding vector of the n-th word in the conversation; Hierarchical memory network: used to store dialogue history and knowledge base information; Memory network pointer: Using as a condition to query the relevant row S in the knowledge base kb , the relevant row S in the dialogue memory network d and the content o in the relevant knowledge base kb ; Copy-enhanced decoder: used to obtain specific answer information; The data output end of the context encoder is connected to the data input end of the memory network pointer and the data input end of the copy-enhanced decoder. The data end of the memory network pointer is connected to the data end of the hierarchical memory network. The data output end of the hierarchical memory network is connected to the data input end of the copy-enhanced decoder. The data output end of the copy-enhanced decoder is connected to the data input end of the hierarchical memory network; The replication-enhanced decoder includes: first generating a rough answer through a rough recurrent neural network model, and then using an entity decoder to query the entities in a knowledge base memory E kb and the dialogue history P d to decode the rough answer to obtain specific answer information; The entity decoder includes: Replacing a rough label generated by a rough recurrent neural network model with an entity from the hierarchical memory network, including the following steps: S-111, use the memory reader to obtain the encoding vector of the nth word in the conversation as the input to get the probability of each row and in the current step t, representing the word distribution probability of the knowledge base network in the current step t, representing the word distribution probability of the dialogue memory network in the current step t; S-112. Secondly, the probability of each column in the knowledge base memory network can be obtained through the following formula: where, φ emb (·) is an embedding function; n j is the entity type name of the j-th column in the knowledge memory network; Denote the state vector of the t-th encoding; S-113. Finally, the following formula is used to calculate the probability of each entity in the knowledge base memory network: Among them, represents the probability distribution of entities in the knowledge base memory network at the current step t; P t kb represents the probability at the t-th step of the knowledge base; score t Represents the score of the attribute in the current step t; By finding the i-th row and j-th column that satisfies v i,j =y t Entity v i,j , and the corresponding satisfy To define the labels of the entity decoder in the knowledge base memory network: where max(·) represents taking the maximum value; i represents the i-th row of the knowledge base memory network; |C| kb Indicates the number of columns of the knowledge base; j represents the j-th column of the knowledge base memory network; v i,j represents the entity at the i-th row and j-th column; y t represents the true answer; Pointer label indicating the i-th row of the knowledge base memory network; Indicates the number of rows in the knowledge base memory network; Define the label of the entity decoder of the dialogue memory network as shown in the following formula: wherein represents the number of rows in the dialogue memory network; max(·) represents taking the maximum value; v i,0 represents the entity in the i-th row and the 0-th column; y t Represents the true answer.
2. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that, the context encoder includes: First, use the embedding function φ emb to encode each word into a vector, then obtain the context representation of each word in the dialogue history through a single-layer gated recurrent unit, add the context representation to the corresponding position in the dialogue memory network, and use the encoded vector of the nth word in the dialogue as the hidden representation of the dialogue state.
3. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that, the hierarchical memory network includes: Store the dialogue history in the dialogue memory network, store the external knowledge base in the knowledge base memory network, and represent the knowledge base information by knowledge base rows.
4. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that, the hierarchical memory network further includes: Each row of the memory network contains historical dialogue information or knowledge base information. For the dialogue memory network, each row contains a word and its corresponding speaker, time, and location information. For the knowledge memory network, each row stores a topic in the knowledge base and its attribute values; Each column of the memory network is a specific word that the decoder needs to receive. Each attribute value of the topic is stored in the relevant column of the knowledge base memory network, and a special symbol is added at the end as an end position marker.
5. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that, the hierarchical memory network further includes using an end-to-end memory network to model the memory reading process, including the following steps: S-1. First, a series of trainable embedding matrices C = {C 1 , C 2 ,..., C K+1} are used as Then, each word v i,j is encoded into a vector; Among them, C 1 represents the first trainable embedding matrix, C 2 represents the second trainable embedding matrix, C K+1 represents the (K + 1)-th trainable embedding matrix, represents the k-th trainable embedding matrix of the word v i,j where v i,j represents the word at the i-th row and j-th column in the memory network; S - 2, and then represent each row of the memory network as where r i k represents the i-th row of the embedded memory matrix R k , represents the k-th embedded vector of the i-th row and j-th column, and |C| represents the number of columns of matrix C; Read the memory through the following formula: p k = Softmax(q k R k )(1) where p k represents the probability distribution of the memory network rows; q k Represents the query data for the k-th hop; R k represents an embedded memory matrix; Softmax(·) represents the Softmax function; Read the content \(o\) of the \(k\)-th hop of the memory network through the following formula k : Among them, represents the probability of each line for the k-th inference; r i k is the i-th row of the embedded memory matrix R k ; The query data for the (k + 1)-th hop can be obtained through the following formula: q k+1 = q k + o k (3) where q k represents the query data for the k-th hop; o k Represents the content of the k-th hop of the memory network.
6. An end-to-end task-oriented dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that, the memory network pointer includes: S-11, use the encoding vector of the nth word in the dialogue as the query of the memory network to obtain the probabilities of each row in the two memory networks: where q K represents the query data for the K-th hop; · T Indicates the transpose of a matrix; r i K represents the i-th row of the embedded memory matrix R k ; S-12, defining the pointer label of the dialogue memory network by checking whether the word appears in the system answer The pointer label indicating the first row of the dialogue history memory network, The pointer label indicating the second row of the dialogue history memory network, Indicating that it is the |R|th d row of the pointer label; And by checking the rows in the knowledge base memory network that contain the true answers, select the row with the largest number of entities to define the pointer label of the knowledge base memory network, and define the label of the knowledge base memory network as where represents the pointer label of the first row of the knowledge base memory network, represents the pointer label of the first row of the knowledge base memory network, represents the pointer label of the |R| kb th row of the knowledge base memory network; If the check word appears in the system answer and the knowledge base memory network contains the line with the true answer, then If not satisfied, then Indicates whether each line is selected, 1 for selected and 0 otherwise; S-13, using the binary cross-entropy loss function Loss between the S label and the S l label to train the memory network pointer s The loss function of the dialogue memory network is as follows: Indicates the pointer label of the i-th row of the dialogue history memory network; s d,i Represents the probability of the i-th row of the dialogue history memory network; |R| d Indicates the |R|-th d row of the dialogue history memory network; The loss function of the knowledge base memory network is as follows: Pointer label representing the i-th row of the knowledge base memory network; s kb,i Represents the probability of the i-th row of the knowledge base memory network; |R| kb Indicates the |R|-th kb row of the knowledge base memory network; The loss function Loss for the memory network pointer s is defined as follows: Loss s = Loss kbs + Loss ds (7) Among them, Loss ds represents the loss function of the dialogue memory network; Loss kbs Indicates the loss function of the knowledge base memory network; S-14. Adjust each row in the two memory networks using the probability of the memory network pointer: where r i k is the i-th row of the embedded memory matrix R k ; Represents the probability distribution of the i-th row of the memory network.
7. An end-to-end task-based dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that the rough recurrent neural network model includes: Read from the memory network pointer and the knowledge base content as input: Among them represents matrix W 1 and the matrix formed by are concatenated; The encoding vector representing the n-th word in the dialogue; Represents the content of the k-th hop in the knowledge base memory network; means the decoding state at the 0th time of the dialogue state; W 1 is a learnable matrix parameter; Then replace the entity in the answer with a rough label, and the decoding process is as follows: Among them, represents the state vector of the t-th decoding; Denote the state vector of the (t - 1)-th decoding; C 1 (y t-1 ) represents the embedding matrix parameter of the word y t-1 ; GRU(·) is a single-layer gated recurrent unit; C 1 is the embedding matrix parameter used by the hierarchical memory network; Then the probability distribution P of the word in the current step t can be obtained by the following formula t v : where W 2 is a learnable matrix parameter; Finally, the word distribution probability P is used t v and the true value of the rough answer are used to train the rough recurrent neural network model through the cross-entropy loss function: Among them represents the t-th rough answer; m Indicates the total number of steps, i.e., the total number of loops; Indicates The word distribution probability of 8. An end-to-end task-based dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that it also includes the loss function of the entity selector defined as follows: where represents the probability distribution of the entity in the knowledge base memory network at the current step t; Label representing the entity decoder in the knowledge base memory network; Represents the probability distribution of the entity in the dialogue memory network at the current step t; Label for the entity decoder of the dialogue memory network.
9. An end-to-end task-based dialogue system in the field of architecture based on a hierarchical memory network according to claim 1, characterized in that it also includes the loss function Loss: Loss=αLoss s +βLoss v +γLoss ent (18) where α, β, and γ are all hyperparameters; Loss s The loss function representing the memory network pointer; Loss v represents the cross-entropy loss function; Loss ent The loss function representing the entity selector.
Citation Information
Patent Citations
Emotion attribute determination method and device and electronic equipment
CN112241453A
Intelligent dialogue system fusing multiple attention mechanisms
CN113505208A