Multi-turn dialogue chatbot system based on knowledge graph
Patent Information
- Application Number
- KR1020220186942
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-28
- Publication Date
- 2026-09-21
- Estimated Expiration
- 2042-12-28
Smart Images

Figure R1020220186942_ABST
Abstract
Description
Technology Field
[0001] The following description concerns chatbot creation technology. Background Technology
[0003] A chatbot is a type of chat system that generates real-time natural language feedback based on a user's natural language input. Due to their rich application scenarios and potential commercial value, chatbots capable of communicating seamlessly and naturally with people have always emerged as a key focus in the field of artificial intelligence. However, chatbots generally suffer from a lack of knowledge and the inability to fully understand the logical relationships between objects, which leads to a deficiency in knowledge reasoning capabilities. A knowledge graph is a type of graph data structure used to record knowledge and elucidate the logical relationships between objects. By utilizing effective algorithms, more complex logical reasoning tasks can be achieved, which is an effective method for improving chatbot performance.
[0004] Meanwhile, Korean Patent Publication No. 10-2021-0109208 (publication date September 6, 2021) discloses a technology that provides a question-response service by deriving a plurality of example query sentences related to a user query sentence based on an intention-based probability model, selecting at least one example query sentence among the derived plurality of example query sentences based on a sentence-based probability model, and transmitting an example response sentence corresponding to the selected example query sentence to a user terminal as a response sentence corresponding to the user query sentence. The problem to be solved
[0006] A method and system can be provided for generating a response to an input query by utilizing graph information of a subgraph retrieved from a knowledge graph constructed in a joint augmentation model based on an input query, in a joint augmentation model configured by combining a knowledge fusion mechanism with a natural language processing model composed of an encoder and a decoder. means of solving the problem
[0008] A chatbot system may include a step of receiving a query into a joint augmentation model configured by combining a knowledge fusion mechanism with a natural language processing model composed of an encoder and a decoder; and a response providing unit that generates a response to the received query using graph information of a subgraph retrieved from a knowledge graph built in the joint augmentation model based on the received query.
[0009] The above joint augmentation model may be composed of a KE-Seq2Seq-based encoder and a KD-Seq2Seq-based decoder to enable knowledge inference and multi-line conversation capabilities.
[0010] The above-mentioned co-enhancement model may use an attention mechanism through a recurrent neural network in the encoder to apply attention to a knowledge subgraph retrieved from an input sequence, connect a vector of the knowledge subgraph with a word vector, and then input it into a non-linear recurrent neural network.
[0011] The above-mentioned joint augmentation model may search for a subgraph related to the current input word in a knowledge graph and obtain a word vector and a knowledge graph vector from a list of words trained for the current input word.
[0012] The above-mentioned co-augmentation model obtains a history knowledge semantic vector by calculating a knowledge subgraph retrieved from a history conversation using a dual attention mechanism through a recurrent neural network in the decoder, and the history knowledge semantic vector may be added to a non-linear operation of decoding to obtain the output of the current moment.
[0013] The above joint augmentation model can calculate attention weights for each position state of the encoder using an attention mechanism based on the previous moment state, obtain a sub-semantic vector of the input sequence through a weighted sum of the calculated attention weights, calculate attention weights for the history knowledge subgraph using an attention mechanism based on the previous moment state, and obtain a history subgraph vector by weighted summing the vectors of each triple in the history subgraph.
[0014] The above joint augmentation model can generate a word to obtain a word vector from a trained word list using the obtained sub-semantic vector and the obtained historical sub-graph vector, generate a current word through an LSTM for non-linear operations to obtain a current moment state, and generate an output sequence for the generated current word.
[0015] A method for generating a question and answer performed by a chatbot system may include: a step of receiving a query into a joint augmentation model configured by combining a knowledge fusion mechanism with a natural language processing model composed of an encoder and a decoder; and a step of generating a response to the received query using graph information of a subgraph retrieved from a knowledge graph built in the joint augmentation model based on the received query. Effects of the invention
[0017] It can improve the knowledge reasoning ability and multi-round conversation ability of the chatbot system. Brief explanation of the drawing
[0019] FIG. 1 is a block diagram illustrating the configuration of a chatbot system in one embodiment. FIG. 2 is a flowchart illustrating a method for generating a conversation based on a knowledge graph in a chatbot system in one embodiment. FIG. 3 is an example illustrating the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment. FIG. 4 is another example for explaining the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment. FIG. 5 is another example illustrating the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment. Specific details for implementing the invention
[0020] Hereinafter, embodiments will be described in detail with reference to the attached drawings.
[0022] FIG. 1 is a block diagram illustrating the configuration of a chatbot system in one embodiment, and FIG. 2 is a flowchart illustrating a method for generating a conversation based on a knowledge graph in a chatbot system in one embodiment.
[0023] The processor of the chatbot system (100) may include a query input unit (110) and a conversation generation unit (120). These components of the processor may be representations of different functions performed by the processor according to control commands provided by program code stored in the chatbot system. The processor and the components of the processor may control the chatbot system to perform steps (210 to 220) included in a method for generating a conversation based on the knowledge graph of FIG. 2. In this case, the processor and the components of the processor may be implemented to execute instructions according to the code of an operating system included in memory and the code of at least one program.
[0024] The processor can load program code stored in a file of a program for a method of generating a conversation based on a knowledge graph into memory. For example, when a program is executed in a chatbot system, the processor can control the chatbot system to load program code from a file of a program into memory under the control of the operating system. At this time, the query input unit (110) and the conversation generation unit (120) may each be different functional representations of the processor for executing commands of corresponding parts of the program code loaded into memory to execute subsequent steps (210 to 220).
[0025] In step (210), the query input unit (110) can receive a query into a joint augmentation model composed of a natural language processing model composed of an encoder and a decoder, combined with a knowledge fusion mechanism. The joint augmentation model may be composed of a KE-Seq2Seq-based encoder and a KD-Seq2Seq-based decoder to enable knowledge inference and multi-line conversation capabilities.
[0026] In step (220), the conversation generation unit (120) can provide a response to the input query by using graph information of a subgraph retrieved from a knowledge graph built in a joint augmentation model based on the input query.
[0027] FIG. 3 is an example illustrating the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment.
[0028] We will now discuss the case where the natural language processing model is a KE-Seq2Seq model. Since traditional conversation generation models are trained based solely on large amounts of Q&A pair data, the encoder-decoder model can learn syntactic-semantic information of the sentences in the Q&A pairs; however, KE-Seq2Seq models tend to select words that occur more frequently after training based on empirical risk minimization. Consequently, chatbots fail to generate correct responses from semantically similar words rather than the correct words they encounter in knowledge-based questions. To address this deficiency, KE-Seq2Seq models utilize a knowledge graph to search for subgraphs within the knowledge graph based on input questions and integrate subgraph information into the encoder based on an attention mechanism. This provides the model with prior knowledge and improves its understanding of external knowledge.
[0029] We will now explain the detailed operation of the KE-Seq2Seq model. In the KE-Seq2Seq model, the encoder and decoder are in the LSTM neural network model Receives input and expects output as response It follows an encoder-decoder architecture that derives. Also, knowledge graph information It must be integrated into the following probability distribution.
[0030] Mathematical Formula 2-1:
[0031]
[0032] First, when a series of sentences is input, the subgraphs corresponding to the input sentences can be searched in the knowledge graph based on each word of the sentences. Subgraph It contains multiple triples, and since "of" cannot find an entity in the knowledge graph, a predefined subgraph is returned. For location i, the searched subgraph Sub-graph vectors based on information is generated, and the sub-graph vector is the corresponding word vector as input information for encoder position i, as shown in Equation 2-2. It continues with.
[0033] Mathematical formula 2-2:
[0034]
[0035] Here is the hidden state of the previous LSTM. And the subgraph vector is derived from the proposed encoder knowledge fusion mechanism, which basically uses an attention mechanism for subgraphs Current input in Calculate the connections between pairs and each triad, and can be expressed as follows.
[0036] Mathematical formula 2-3:
[0037]
[0038] Mathematical formula 2-4:
[0039]
[0040] Mathematical formula 2-5:
[0041]
[0042] Here , , are the mapping transformation weight matrices for the word vector, head entity, relationship, and tail entity, respectively. As shown in Equation 2-5, to obtain subgraph information, , , current word vector , triad head entity vector , tail entity vector Input it into the hyperbolic tangent tanh activation function, perform matrix transformation, sum them, and then the relational vector rafter Word vectors from the current triad by performing a dot product with the matrix result denormalized attention weights Calculate, and all triads are calculated after mathematical formula 2-4, and the current word Triad Calculate the normalized attention weight a up to. Finally, according to Equation 2-3, the subgraph The triads and tail entity vectors of all triangles are summed by the attention weight a to form the subgraph vector It is continued and normalized to obtain.
[0043] FIG. 4 is another example for explaining the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment.
[0044] We will now explain the case where the natural language processing model is a KD-Seq2Seq model. Similar to KE-Seq2Seq, the KD-Seq2Seq model searches for relevant subgraphs in the knowledge graph based on historical conversations; however, since historical conversation input is not required, it cannot be used in the same way as the KE-Seq2Seq model, such as concatenating subgraph vectors with corresponding word vectors in the encoder. Therefore, the KD-Seq2Seq model integrates subgraph information into the decoder, which is another module of the Seq2Seq architecture. The difference is that the KESeq2Seq encoder only needs to compute one piece of subgraph information at each position, whereas the KDSeq2Seq decoder must compute multiple different subgraphs at each moment. Thus, the proposed decoder historical knowledge fusion mechanism is a dual attention mechanism. First, attention will be calculated for all triples of the multi-subgraph to obtain the corresponding multi-subgraph vectors, and then attention can be calculated for these multi-subgraph vectors to obtain all historical information. This can then be added to the current decoding to complete the fusion of historical knowledge.
[0045] More specifically, in the case of the KD-Seq2Seq encoder, as described in Fig. 3, the input sequence It can be encoded following an LSTM model. For the decoder, at each moment t Mathematical formulas 3-1 and 3-2 are used to generate.
[0046] Mathematical Formula 3-1:
[0047]
[0048] Mathematical Formula 3-2:
[0049]
[0050] Here, w is the output weight matrix is a moment Decoder state at, is a moment Decoder state in, ct is the input sequence semantic vector, is a historical knowledge semantic vector, is a moment Decoder-generated words in It is a word vector of.
[0051] Input sequence semantic vector It is generated by calculating attention weights for each position in the input sequence through an attention mechanism and then assigning weights to the sum. Historical knowledge semantic vector In the case of, each history subgraph as follows First, attention is calculated at the triad level, and attention is generated by calculating it twice.
[0052] Mathematical formula 3-3:
[0053]
[0054] Mathematical formula 3-4:
[0055]
[0056] Mathematical formula 3-5:
[0057]
[0058] Here , , , are the mapping transformation weight matrices of the state of the previous moment, the head entity, the relationship, and the tail entity, respectively, and the decoder state of the previous moment is is. Triplet head entity vector , tail entity vector For each matrix transformation , , Multiply by and then sum. And the denormalized attention weights is the relationship after matrix transformation of the hyperbolic tangent function It is obtained by multiplying the transpose of the vector. is Triad and It can be viewed as a correlation of states, followed by normalized attention weights after the Softmax function can be obtained. j is the history subgraph Triad It can be viewed as the probability of selecting. Finally, the attention weight of each triad is calculated by weighting the head entity and tail entity splicing vectors of each triad and summing them to obtain the historical (past) subgraph vector Obtains. Historical knowledge subgraph set A set of historical knowledge subgraph vectors can be obtained by performing the above operations for each subgraph of.
[0059] FIG. 5 is another example illustrating the operation of generating a conversation based on a knowledge graph in a chatbot system in one embodiment.
[0060] We will now describe the case where the natural language processing model is a KED-Seq2Seq model. The KED-Seq2Seq model uses the encoder of the KE-Seq2Seq model; when the recurrent neural network encodes the input sequence, it also pays attention to the knowledge subgraph retrieved from the input sequence through an attention mechanism, concatenates the knowledge subgraph vectors with the corresponding word vectors, and then performs a non-linear recurrent neural network. At the same time, the KED-Seq2Seq model uses the decoder of the KD-Seq2Seq model. When the decoder's recurrent neural network decodes at each step, it computes the knowledge subgraph retrieved from the historical conversation using a dual attention mechanism to obtain the historical knowledge semantic vector. The knowledge semantic vector is added to the non-linear operation of the decoding to obtain the output of the current moment.
[0061] In the encoder section, each position is a related subgraph of the current input word x in the knowledge graph. Searching for, here is. And then, the current input word x is the corresponding word vector from the trained GloVe word list. Obtaining, similarly, the corresponding knowledge graph vector from the trained TransE word list It can be seen that not only the relationships within the knowledge subgraph that can be obtained, but also each triad of the entity is acquired. Here is. Knowledge subgraph vector based on word vectors and the corresponding knowledge graph vectors of each triad It is calculated by the formula and finally connected to the corresponding knowledge subgraph vector. Word vector is the input to the current position of the LSTM. This completes the encoding of knowledge graph information associated with the input sequence.
[0062] In the decoder section, the hidden state of the LSTM neural network at each moment depends on the previous moment state. First, the previous moment state The encoder's position state by the attention mechanism Calculate the attention weights for, and here is. Next, the semantic vector of the input sequence through the weighted sum. Find . Second, state The attention weights for the history knowledge subgraph by the attention mechanism Calculate, where j ∈ It is. There are two specific stages here. First, the previous moment state The attention weight of each triple in the history subgraph is calculated through the attention mechanism, and the history subgraph vector can be obtained by weighting and summing the vectors of each triple in the history subgraph. Second, the state It calculates the attention weights of each history subgraph vector through the attention mechanism, assigns weights to the vectors of each history subgraph, and sums them to form the history knowledge semantic vector. Obtain the corresponding word vector from the trained GloVe word list at the previous decoding moment. Finally, the word vector from the previous decoding moment. To get the word It can generate. Then, the current moment state vector According to the LSTM unit formula for the non-linear operation to obtain, and the current word According to the formula decoding for generating, cycle decoding is the output sequence It can generate.
[0063] To explain the technical effects of the method proposed in the examples, experiments can be performed on three datasets: KdConv, DuConv, and DuRecDial. To explain the effects of KED-Seq2Seq, it can be compared with KE-Seq2Seq, which has enhanced knowledge inference capabilities, and KD-Seq2Seq, which has enhanced multi-round conversational capabilities. Similarly, for the experiments, four automated test metrics—BLEU, Knowledge Precision, Knowledge Recall, and Knowledge—and one manual test metric—Fitness—can be selected. Additionally, recurrent neural networks can be stacked in multiple layers, and generally, recurrent neural networks possess stronger feature extraction and non-linear computation capabilities than single-layer recurrent neural networks. Therefore, when input data is relatively large or complex, multi-layer recurrent neural network stacking is typically used as an encoder and decoder; thus, the experiments explore the impact of the number of network layers on the model's performance, in addition to the size of the hidden state vector, word vector, and knowledge graph vector. The hyperparameters of LSTM, such as mini-batch size, training, hidden state vector, word vector, knowledge graph vector size, mini-batch size, training rate size, training optimizer, and training strategy settings, are consistent with KE-Seq2Seq and KD-Seq2Seq.
[0064] The device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and components described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include a plurality of processing elements and / or a plurality of types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.
[0065] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or instruct the processing unit independently or collectively. Software and / or data may be embodied in any type of machine, component, physical device, virtual equipment, computer storage medium, or device so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.
[0066] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc.
[0067] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.
[0068] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.
Claims
Claim 1 In a chatbot system, a step of receiving a query into a joint augmentation model configured by combining a knowledge fusion mechanism with a natural language processing model composed of an encoder and a decoder; and includes a response providing unit that generates a response to the received query using graph information of a subgraph retrieved from a knowledge graph constructed in the joint augmentation model based on the received query, wherein the joint augmentation model incorporates, in the encoder, an attention mechanism through a recurrent neural network to reflect attention on a knowledge subgraph retrieved from an input sequence to connect a vector of the knowledge subgraph with a word vector, and then inputs it into a non-linear recurrent neural network, retrieves a subgraph related to the current input word in the knowledge graph, and obtains a word vector and a knowledge graph vector from a list of words trained for the current input word, and in the decoder, calculates a knowledge subgraph retrieved from a historical conversation using a dual attention mechanism through a recurrent neural network to obtain a historical knowledge semantic vector, calculates an attention weight for each position state of the encoder by an attention mechanism using a previous moment state, obtains a sub-semantic vector of the input sequence through a weighted sum of the calculated attention weights, and uses the obtained sub-semantic vector and the obtained historical subgraph vector from a list of words trained A chatbot system comprising generating a word to obtain a word vector, generating a current word through an LSTM for non-linear operations to obtain a current moment state, generating an output sequence for the generated current word, calculating attention weights for a history knowledge subgraph by an attention mechanism using the previous moment state, and obtaining a history subgraph vector by weighted summing the vectors of each triple in the history subgraph. Claim 2 A chatbot system according to claim 1, wherein the joint augmentation model is composed of a KE-Seq2Seq-based encoder and a KD-Seq2Seq-based decoder to enable knowledge inference and multi-line conversation capabilities. Claim 3 delete Claim 4 delete Claim 5 A chatbot system according to claim 1, characterized in that the above-mentioned history knowledge semantic vector is added to a non-linear operation of decoding to obtain the output of the current moment. Claim 6 delete Claim 7 delete Claim 8 A method for generating a conversation based on a knowledge graph performed by a chatbot system, comprising the step of receiving a query into a joint augmentation model configured by combining a knowledge fusion mechanism with a natural language processing model composed of an encoder and a decoder; The method includes the step of generating a response to the received query using graph information of a subgraph retrieved from a knowledge graph constructed in the joint augmentation model based on the received query, wherein the joint augmentation model incorporates: in the encoder, using an attention mechanism through a recurrent neural network, attention is applied to the knowledge subgraph retrieved from the input sequence to connect the vector of the knowledge subgraph with a word vector, and then inputs it into a non-linear recurrent neural network; retrieves a subgraph related to the current input word in the knowledge graph; obtains a word vector and a knowledge graph vector from a list of words trained for the current input word; in the decoder, calculates the knowledge subgraph retrieved from the historical conversation using a dual attention mechanism through a recurrent neural network to obtain a historical knowledge semantic vector; calculates attention weights for each position state of the encoder by an attention mechanism using the previous moment state; obtains a sub-semantic vector of the input sequence through a weighted sum of the calculated attention weights; and uses the obtained sub-semantic vector and the obtained historical subgraph vector to obtain a word from a list of words trained for the current input word A method for generating a conversation based on a knowledge graph, comprising generating a word to obtain a vector, generating a current word through an LSTM for non-linear operations to obtain a current moment state, generating an output sequence for the generated current word, calculating attention weights for a history knowledge subgraph by an attention mechanism using the previous moment state, and obtaining a history subgraph vector by weighted summing the vectors of each triple in the history subgraph.
Citation Information
Patent Citations
System for question answering knowledge graphs using graph neural network
KR1020220019461A