Dialogue model training method and apparatus

JP7927510B2Active Publication Date: 2026-10-01HYPERCONNECT LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2022133106
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-22
Filing Date
2022-08-24
Publication Date
2026-10-01
Estimated Expiration
2042-08-24

AI Technical Summary

Benefits of technology

【0013】 本開示によると、検索基盤対話モデルが、大規模な言語モデルの豊富な知識に基づいて流暢な回答を生成する生成基盤対話モデルの回答クオリティーに対応する回答を返すことができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007927510000018
    Figure 0007927510000018
  • Figure 0007927510000019
    Figure 0007927510000019
  • Figure 0007927510000020
    Figure 0007927510000020
Patent Text Reader

Abstract

To provide a method for training a dialogue model of a user which reduces a high response latency of a generative-based dialogue model or compensates for a relatively low response quality of a retrieval-based dialogue model, and a device therefor.SOLUTION: A method includes: a step S101 of selecting a first context from a first dialogue data set including at least one pair of a context and a response corresponding to the context; a step S102 of generating a first response corresponding to the first context through a first dialogue model; a step S103 of generating an augmented dialogue data set by including a pair of the first context and the first response corresponding to the first context in the first dialogue data set; and a step S104 of training a second dialogue model based on the augmented dialogue data set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[[Technical Field]]

[0001] The present disclosure relates to a method for training a user dialogue model and an apparatus therefor. [[Background Art]]

[0002] With the development of artificial intelligence technology, people can now interact with chatbots as virtual characters rather than real persons. Such a chatbot can retrieve and output a predetermined response according to a predetermined dialogue topic, or can generate and output an appropriate response for a free dialogue topic. A dialogue in which a specific dialogue topic is not predetermined can be referred to as open domain conversation.

[0003] In order to derive a response in an open domain conversation, two main types of dialogue models are used: generation-based dialogue models and retrieval-based dialogue models. A generation-based dialogue model is a model that generates an appropriate answer based on an input dialogue context and returns the answer as a response. A retrieval-based dialogue model is a model that predefines a response set that can be used as answers, then retrieves the most appropriate answer for the input dialogue context from the response set and returns the answer as a response.

[0004] When such a generation-based dialogue model is used in conjunction with a large-scale language model, it can generate an answer suitable for a given dialogue context based on the rich knowledge of the language model. However, the generation-based dialogue model has high latency in answer generation because the decoder of the sequence-to-sequence structure spends a lot of time on the autoregressive decoding process. In an actual dialogue situation, the chatbot must return an answer to the user in real time, so such a heavy and slow characteristic of the generation-based dialogue model is actually difficult to apply to open domain conversations.

[0005] On the other hand, search-based dialogue models, when used with high-performance search libraries, can return appropriate answers to a given context much faster than generative dialogue models. However, because search-based dialogue models can only return answers that exist in a predefined set of responses, they may give irrelevant and unrelated answers if the appropriate response for the input dialogue context is not included in the response set. Also, because search-based dialogue models are highly dependent on the predefined set of responses, they may return less fluent answers compared to generative dialogue models. [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] This disclosure aims to solve the high response latency problem of generative dialogue models by generating responses to a dialogue context through a generative dialogue model and constructing a set of responses for a search-based dialogue model based on the generated responses.

[0007] This disclosure aims to address the relatively low response quality of search-based dialogue models by generating responses to a dialogue context through a generative dialogue model and training a search-based dialogue model based on the generated responses.

[0008] The technical problems that this disclosure aims to solve are not limited to those described above, and other technical problems may be inferred from the following embodiments. [Means for solving the problem]

[0009] A dialogue model training method performed in an electronic device for solving the aforementioned problems may include the steps of: selecting a first context from a first dialogue dataset containing one or more context and corresponding response pairs; generating a first response corresponding to the first context through a first dialogue model; generating an augmented dialogue dataset by adding the first context and the corresponding first response pairs to the first dialogue dataset; and training a second dialogue model based on the augmented dialogue dataset.

[0010] Furthermore, other dialogue model training methods performed in electronic devices for solving the aforementioned problems may include the steps of: acquiring a response set from a first dialogue dataset, including a first response subset corresponding to a first context and an arbitrarily selected second response subset; calculating a first score for the responses included in the response set for the first context based on the first dialogue model; calculating a second score for the responses included in the response set for the first context based on the second dialogue model; and training the second dialogue model based on the first and second scores.

[0011] Furthermore, an electronic device for training a dialogue model to solve the aforementioned problems includes a storage device and a controller, the controller which, through the storage device, obtains a response set from a first dialogue dataset, including a first response subset corresponding to a first context and an arbitrarily selected second response subset; calculates a first score for the responses included in the response set for the first context based on the first dialogue model; calculates a second score for the responses included in the response set for the first context based on the second dialogue model; and trains the second dialogue model based on the first and second scores.

[0012] Further details of the embodiments are included in the detailed description and drawings. [Effects of the Invention]

[0013] According to this disclosure, a search-based dialogue model can return answers that correspond to the answer quality of a generative dialogue model, which generates fluent answers based on extensive knowledge of a large-scale language model.

[0014] Furthermore, according to this disclosure, it is possible to solve the high latency problem of the generation-based dialogue model and improve the response quality of the search-based dialogue model.

[0015] The effects of the invention are not limited to those mentioned above, and any other effects not mentioned can be clearly understood by a person of the art ordinary from the description of the claims. [Brief explanation of the drawing]

[0016] [Figure 1] This is a flowchart illustrating a data-level dialogue model training method according to one embodiment. [Figure 2] This is a flowchart illustrating a model-level dialogue model training method according to one embodiment. [Figure 3] This graph shows latency-based interpersonal evaluation scores for open-domain dialogue models. [Figure 4] This is a diagram illustrating a data-level dialogue model training method according to one embodiment. [Figure 5] This is a diagram illustrating a model-level dialogue model training method according to one embodiment. [Figure 6] This is a block diagram showing an electronic device for training a model-level dialogue model according to one embodiment. [Modes for carrying out the invention]

[0017] The terminology used in the embodiments has been selected as widely used and general terms as possible, taking into account the function described herein, although this may change depending on the intent of the articulators, case law, the emergence of new technologies, etc. In certain cases, the applicant may have arbitrarily selected terms, in which case their meaning will be described in detail in the relevant sections of the description. Accordingly, the terminology used in this disclosure must be defined not merely as a name of a term, but based on the meaning of the term and the overall content of this disclosure.

[0018] Throughout the specification, when a part "includes" a component, this does not mean that other components are excluded, but rather that other components may be included, unless otherwise stated. Furthermore, terms such as "~part" and "~module" used in the specification refer to a unit that processes at least one function or operation, which may be embodied as hardware or software, or as a combination of hardware and software.

[0019] The expression “at least one of a, b, and c” as described throughout the specification may encompass “a alone,” “b alone,” “c alone,” “a and b,” “a and c,” “b and c,” or “all of a, b, and c.”

[0020] In the following, "terminal" can be embodied as a computer or mobile terminal capable of connecting to servers or other terminals via a network. Here, "computer" includes, for example, laptops, desktops, and laptops equipped with a web browser, and "mobile terminal" can include, for example, all types of handheld-based wireless communication devices such as IMT (International Mobile Telecommunication), CDMA (Code Division Multiple Access), W-CDMA (W-Code Division Multiple Access), LTE (Long Term Evolution) terminals, smartphones, and tablet PCs, as long as portability and mobility are guaranteed.

[0021] Work on open-domain dialogues, the technical field of this disclosure, has been studied using search models, generative models, or both. While a search model retrieves responses relevant to a given context from a predefined set of responses, a generative model generates responses based on the given context using automatic regression decoding. Search and generative models are known to have advantages in terms of inference efficiency and the quality of generated responses, respectively. To obtain both advantages, several example-based generative models have recently been proposed by combining search and generative models. The main difference between the training method proposed in this disclosure and the example-based generative model is that the example-based generative model provides the generative model with knowledge about the search model, while the proposed training method transmits knowledge about the generative model to the search model, focusing on the efficiency of the open-domain dialogue system.

[0022] More specifically, in open-domain dialogue, large-scale generative models, despite their remarkable performance, are known to be impractical for building real-time dialogue systems due to their high latency. On the other hand, retrieval models can return responses with much shorter latency, but their dialogue quality is limited by a predefined set of responses, resulting in inferior performance compared to large-scale generative models. To leverage both approaches, this disclosure proposes a novel learning method called G2R (Generative-to-Retrieval distillation) that injects knowledge about the generative model into the retrieval model to leverage the dialogue capabilities of the large-scale generative model while maintaining the efficiency of the retrieval model. G2R can be composed of two distinct distillation techniques. First, data-level G2R augments the dialogue dataset with additional responses generated by the large-scale generative model, while model-level G2R transfers the response quality scores evaluated by the generative model to the retrieval model's score through knowledge distillation loss. Through extensive experiments, including human evaluation, it can be confirmed that the search infrastructure dialogue system trained as G2R in this disclosure exhibits significantly improved performance compared to basic search models, while also showing far lower inference latency than large-scale generative models.

[0023] To this end, in this embodiment, new context-response pairs are generated by selecting at least some contexts from a dialogue dataset of context-response pairs for training the search model and generating responses using a generative model. The search model is then trained using the augmented dialogue dataset of the generated context-response pairs, thereby enabling the search model to generate a wider variety of answers.

[0024] Furthermore, in this embodiment, multiple response sets can be examined for a specific context, and the score of each response in each response set can be derived through different models. A single model can then be trained to reduce the score difference between the response sets of each model. More specifically, the performance of the search model can be further improved by using the generative model as the training model and the search model as the student model to derive the cross-entropy loss of the scores for the response sets, and then training the search model to reduce the difference between each score.

[0025] The embodiments of this disclosure will be described in detail below with reference to the attached drawings, so that they can be easily implemented by a person with ordinary skill in the art to which this disclosure belongs. However, this disclosure may be embodied in several different forms and is not limited to the embodiments described herein. Furthermore, expressions such as “First” and “Second” used in this disclosure are provided for distinction between terms and are not intended to limit the meaning.

[0026] Figure 1 is a flowchart illustrating a data-level dialogue model training method according to one embodiment.

[0027] In step S101, a first context can be selected from a first dialogue dataset containing one or more context and corresponding response pairs. According to one embodiment, the responses in the first dialogue dataset may be responses included in a predefined set of responses.

[0028] In step S102, a first response corresponding to a first context can be generated through the first dialogue model. According to one embodiment, the first dialogue model may be a generative-based dialogue model that generates a response for a given context.

[0029] In step S103, an augmented dialogue dataset can be generated, which includes a first context and a corresponding first response pair. According to one embodiment, the dialogue model training method of the present disclosure may generate an augmented response set, which includes the responses from the first dialogue dataset and the generated first responses, and utilize this as a fixed response set for subsequent input dialogue contexts.

[0030] In step S104, a second dialogue model can be trained based on the augmented dialogue dataset. According to one embodiment, the second dialogue model may be a search-based dialogue model that retrieves responses for a given context.

[0031] According to one embodiment, as part of training a second dialogue model, a second context included in an augmented dialogue dataset can be identified, and a response set can be obtained that includes a first response subset corresponding to the second context and an arbitrarily selected second response subset. Then, a first score can be calculated for the responses included in the response set for the second context based on the first dialogue model, and a second score can be calculated for the responses included in the response set for the second context based on the second dialogue model. The second dialogue model can then be trained based on the first and second scores. At this time, a loss can be calculated for the second score based on the first score, and the second dialogue model can be trained in a direction that minimizes the loss. Below, examples of scoring responses according to some embodiments of this disclosure will be described in more detail.

[0032] Figure 2 is a flowchart illustrating a model-level dialogue model training method according to one embodiment.

[0033] In step S201, a response set can be obtained from the first dialogue dataset, including a first response subset corresponding to the first context and an optionally selected second response subset. According to one embodiment, a second context can be selected from a second dialogue dataset containing one or more contexts and their corresponding response pairs, responses corresponding to the second context can be generated through the first dialogue model, and the first dialogue dataset can be generated by including the second context and its corresponding response pairs in the second dialogue dataset. As a result, the first dialogue dataset may be an augmented dialogue dataset.

[0034] In step S202, a first score can be calculated for the responses included in the response set for the first context based on the first dialogue model. According to one embodiment, the first score may be calculated using a normalized log-likelihood based on the length of each response included in the response set. Alternatively (or combined), the first score may be calculated based on the Mutual Information score for each response included in the response set regarding the first context.

[0035] In step S203, a second score can be calculated for the responses included in the response set for the first context based on the second dialogue model. According to one embodiment, the first context and the responses included in the response set are encoded as fixed-length embeddings, and a relevance score for each response in the response set with respect to the first context is calculated based on the embedding value corresponding to the first context and the embedding value corresponding to each response in the response set, and the second score can be calculated.

[0036] In step S204, the second dialogue model can be trained based on the first and second scores. According to one embodiment, a loss can be calculated based on the first and second scores, and the second dialogue model can be trained to minimize such a loss. This loss may include a cross-entropy loss for the scores corresponding to the first response subset and a knowledge distillation loss for the scores corresponding to the responses included in the response set. In some embodiments, the second dialogue model can be trained to maximize the scores corresponding to the first response subset to minimize the cross-entropy loss, and in some embodiments, it may be trained so that the first and second scores coincide to minimize the knowledge distillation loss.

[0037] In the following, the search-based dialogue model will be referred to as the "search model," and the generative-based dialogue model will be referred to as the "generative model."

[0038] Recently, generative models have achieved great success in open-domain dialogue, along with the development of large-scale language models, providing fluent and informative responses. However, generative models have latency and computational resource issues for building real-time dialogue systems, due to the automatic regression decoding required for response generation and the large GPU memory space.

[0039] On the other hand, search models such as bi-encoders and poly-encoders can build efficient open-domain dialogue systems by pre-defining response sets and searching for the most relevant responses to a given context from those response sets. Furthermore, bi-encoders can significantly reduce latency when employing efficient MIPS (Maximum Inner Product Search) libraries such as FAISS and ScaNN. Despite this superior efficiency, search models are shown to be somewhat less interactive than generative models. In particular, search models are known to return incorrect responses when the pre-defined response set does not contain the appropriate response for a given context, while generative models tend to handle such cases more flexibly.

[0040] To attempt to mitigate these problems, example-based generative models, which combine the advantages of two approaches, have been considered, but the inherent inefficiencies of generative models still remain. This is because example-based generative models use generative models for response generation. Therefore, this disclosure proposes a novel training method for retrieval models called G2R (Generative-to-Retrieval distillation) to create a fluent open-domain dialogue system while maintaining the efficiency preferred for practical applications.

[0041] According to one embodiment, G2R allows a search model to leverage knowledge about a large generative model at both the data and model levels. In some embodiments, data-level G2R can augment the original dialogue dataset with responses generated from the large generative model using the context of the original dialogue dataset, and the generated responses can be added to a predefined set of responses. The augmented dialogue dataset and response set are then used to train the search model in the training phase and to return responses in the inference phase, respectively. On the other hand, while data-level G2R allows the search model to leverage high-quality responses generated from the large generative model, it does not convey detailed knowledge of the generative model regarding the quality of individual responses. To address this, model-level G2R employs a knowledge distillation method that transmits response quality scores evaluated in the large teacher generative model as scores in the student search model. This method can guide the search model to select better responses in terms of response quality.

[0042] The transfer of knowledge from larger teacher networks to smaller student networks has been embodied in improving the performance of student models, including data augmentation and knowledge distillation. In terms of data augmentation, some studies utilize the output of pre-trained language models as labeled examples for text classification tasks. Other studies utilize the inference results of search and generation models as semi-negative datasets for training student search models. Knowledge distillation, on the other hand, transfers knowledge from teacher models to student models by matching student logits with softened teacher logits. Knowledge distillation exists specifically designed for particular tasks or model architectures, such as sequence generation tasks, search models, and transformer architectures.

[0043] The work most closely related to the dialogue model training method of this disclosure is dialogue distillation, which proposes data-level and model-level distillation of open-domain dialogue models. However, the dialogue model training method of this disclosure differs from dialogue distillation in three aspects. First, dialogue distillation requires additional unpaired text corpora, which can be difficult to obtain from certain situations. Instead, the dialogue distillation work focuses on leveraging knowledge of large-scale generative models to augment additional data. Also, dialogue distillation does not enrich predefined response sets. This is important for improving the performance of retrieval models, as evidenced by experimental results for the dialogue model training method of this disclosure. Finally, while dialogue distillation considers knowledge distillation only within homogeneous architectures such as generative-to-generative or retrieve-to-retval, the dialogue model training method of this disclosure focuses on model-level distillation between heterogeneous architectures, particularly generative-to-retval, in order to leverage the advantages of each architecture.

[0044] Figure 3 is a graph showing the latency-to-human evaluation scores for open-domain dialogue models. White circles represent generative-based dialogue models, black circles represent search-based dialogue models, and 310 stars represent dialogue models trained through the dialogue model training method of this disclosure according to some embodiments of this disclosure (e.g., G2R). Through Figure 3, it can be seen that the dialogue models of this disclosure show much better human evaluation scores than search-based dialogue models and much shorter latency than generative-based dialogue models, demonstrating that they achieve the optimal score, i.e., the "sweet spot," among a variety of models.

[0045] The search-based dialogue system, composed of a G2R-applied search model and MIPS library, empirically demonstrates considerable dialogue capability while exhibiting fast inference speed, as shown in Figure 3. For example, the search-based dialogue system disclosed herein shows approximately 20 times faster performance compared to the Blender model (90M parameters), while exhibiting similar human evaluation results for dialogue capability. Here, the Blender model is a state-of-the-art model for open-domain dialogue work, and various parameters exist, such as Blender 90M, Blender 2.7B, and Blender 9.4B. The Blender model uses decoding hyperparameters for response generation.

[0046] Figure 4 is a diagram illustrating a data-level dialogue model training method according to one embodiment.

[0047] To aid in understanding the method illustrated in Figure 4, we first consider a search model for open-domain dialogues. The following mathematical equation 1 represents a dialogue dataset containing n context-response pairs, where c i and r i These are gold responses, which are appropriate responses corresponding to the context and the i-th example, respectively. During the training phase, the search model compares the scores of negative responses with the given context c i Gold response to i It can be trained to maximize the score. Then, in the inference phase, the search model can return the response with the highest score for a given context c from a predefined set of responses R constructed in the dialogue dataset D. Mathematical formula 2 shows a predefined set of responses R containing n responses.

[0048]

number

[0049]

number

[0050] Next, knowledge distillation is performed on student model z s Logit and teacher model z t This method transfers knowledge from the teacher model to the student model by adding a loss that matches the logit. For a classification task with one class, the knowledge distillation loss can be defined as the cross-entropy between the softened output probabilities of the student model and the teacher model. Mathematical equation 3 is given by the knowledge distillation loss L KD This indicates.

[0051]

number

[0052] Here, p(y|x) and z(x,y) are the softening probability and logit value of the model for the input x and class y, respectively, and T is the temperature parameter for smoothing the logit value.

[0053] The goal of this disclosure is to create an efficient open-domain dialogue system based on a search model. However, simply utilizing a search model can be inefficient when the size of the response set R is large, because the search model must calculate scores for all response candidates. To address this, a process according to a particular embodiment of this disclosure employs a biencoder model with an efficient MIPS library (e.g., FAISS) to efficiently select an appropriate response without calculating scores for all response candidates. Specifically, the biencoder uses a Transformer architecture to encode the context c and response r as fixed-length embeddings, and the correlation score between c and r can be defined as the inner product of the two embeddings. Through this, the speed of the search process can be increased.

[0054] On the other hand, utilizing an additional high-quality dialogue dataset helps improve the performance of the retrieval model. In addition, enhancing the predefined response set R to include more diverse responses expands opportunities for selecting appropriate responses, which can help the retrieval model respond appropriately to diverse input contexts. However, obtaining such high-quality dialogue datasets or responses through human-in-the-loop annotation is labor-intensive and costly.

[0055] A properly tuned large generative model can achieve human-like conversational capabilities, so the dialogue model training method of the present disclosure can utilize the generation results of the large generative model to expand not only the dialogue dataset for training the retrieval model, but also the response set.

[0056] First, for each context c in dialogue dataset D i , the large generative model G generates m responses r G i,j . Mathematical formula 4 represents the response r G i,j as shown below.

[0057]

Mathematical Expression

[0058] Then, the generated response can be regarded as the gold response for the given context c i and added to the dialogue dataset D and the predefined response set R as shown in Mathematical formula 5. Here, D G and R G represent the augmented dialogue dataset and the augmented response set, respectively.

[0059]

Mathematical Expression

[0060] After the dialogue dataset and response set are augmented, the search model R uses randomly sampled speech responses R. - The cross-entropy loss L maximizes the probability of selecting the correct response r from the set of values. CE It can be learned to minimize the cross-entropy loss L. Mathematical formula 6 is given by L CE This indicates.

[0061]

number

[0062] Here, R(c, r) is the score calculated by the search model R for a given context c and response r. In mathematical formula 6, R - R G Responses are randomly sampled from the source to be generated differently for each iteration. As an example, the Blender 9.4B model, the largest open-domain dialogue model available as a large-scale generative model G, can be used. Since beam search algorithms tend to generate similar responses from the same context, top-k sampling can be applied for response diversity. Additionally, responses can be sampled multiple times with different minimum length constraints to diversify the specificity and length of the generated responses.

[0063] Referring to the data-level dialogue model training model 400 in Figure 4, first, the generative model G430 can be trained using a dialogue dataset D410 that contains one or more pairs of given contexts and corresponding responses. For this purpose, an arbitrary context c420 can be extracted from the dialogue dataset D410 and used as input to the generative model G430. According to one embodiment, the contexts and corresponding responses in the dialogue dataset D410 may be responses returned by an existing search model R470 through a search for a given context. The generative model G430 may be a generative model based on a large-scale language model.

[0064] According to one embodiment, the generation model G430 takes the context c420 as input and generates a new response r G 440 can be generated. This allows the generation model G430 to generate a new response r corresponding to the context c420. G A new dialogue dataset 450 containing one or more pairs of 440 can be generated. Next, the dialogue dataset D410 and the new dialogue dataset 450 are combined to create an augmented dialogue dataset D G It can generate 460. Augmented dialogue dataset D G 460 can be used as input to a model-level dialogue model training model, as shown in Figure 4 below. The search model R470 is used with the augmented dialogue dataset D G It can be learned using 460. On the other hand, a new response r G 440 is added to the existing response set, creating an augmented pre-defined response set. G Composed of 480, the search model R470 will subsequently provide an enhanced and predefined set of responses R, regardless of the context input to the dialogue model. G It can search for 480 and return the appropriate response.

[0065] Figure 5 is a diagram illustrating a model-level dialogue model training method according to one embodiment.

[0066] While data-level dialogue model training methods provide additional high-quality dialogue data and diverse responses, they do not transmit detailed knowledge about the quality of individual responses to the large-scale generative model G. The model-level dialogue model training method of this disclosure is designed to solve this problem. According to one embodiment, the data-level dialogue model training method can solve this problem by transmitting quality scores of individual response levels, evaluated by the large-scale supervising generative model G, to the student search model R.

[0067] Specifically, first, from the perspective of the teacher-generating model G, the response quality score is defined as follows: G(c,r). Next, the student retrieval model can be trained, similar to existing knowledge distillation techniques, so that the student retrieval model score R(c,r) and the teacher-generating model score G(c,r) match.

[0068] According to one embodiment, the score G(c,r) of the teacher-generated model can be defined as the log-likelihood normalized as the length of the response, as shown in mathematical formula 7.

[0069]

number

[0070] Here, P G (r|c) is the probability of the response r for a given context c in the generative model G, and |r| is the number of tokens in the response r. The log-likelihood can be normalized as the length of the response to mitigate the problem of preferring shorter responses. The distillation loss L is calculated by considering the scores G(c,r) of the teacher generative model and R(c,r) of the student search model as logits for the teacher and student models, respectively. KD This allows us to derive the following. As a result, mathematical formula 6 is changed to the following mathematical formula 8.

[0071]

number

[0072] Here, R i D G Context c i This is a set of affirmative responses corresponding to [the given phrase].

[0073] On the other hand, the score G(c) of the supervising generation model for speech responses i ,r - ) requires many additional calculations, therefore, randomly sampled speech responses r - ∈R - The calculation can be simplified by approximating it as shown in mathematical formula 9.

[0074]

number

[0075] Finally, the final loss L for the model-level dialogue model training method is, referring to mathematical equation 10, the original cross-entropy loss L in mathematical equation 6. CE The hyperparameter α controls the weighting of each term in the knowledge distillation loss L. KD It can be shown as a sum of .

[0076]

number

[0077] Referring to the model-level dialogue model training model 500 in Figure 5, first, an arbitrary context c i Response set R corresponding to 510 i 520 can be configured. According to one embodiment, context c i 510 is the enhanced dialogue dataset D in Figure 4. G This may be a context included in 460. Also, response set R i 520 is a new response r G In 440, context c iAppropriate affirmative response to 510 and new response r G In response 440, excluding affirmative responses, one or more arbitrarily selected optional responses (or voice responses) may be included. Response set R i In question 520, only one affirmative response is required.

[0078] Next, response set R i Based on 520, the teacher model, generative model G530, and the student model, search model R550, are in context c i A response can be returned to 510. The generative model G530 is in context c i The score G(c) for the response returned to 510 i ,r)540 response set R i Each response included in 520 can be examined. Furthermore, the search model R550 is in context c i R(c) is the score for the response returned to 510. i ,r)560 response set R i You can examine each response included in 520.

[0079] According to one embodiment, the generative model G530 calculates the log-likelihood normalized as the length of the response r and then G(c i ,r)540 can be calculated. Also, the search model R550 is context c i Encode 510 and response r as fixed-length embeddings, and take the dot product of the two embeddings as context c. i Define the correlation score between 510 and response r and R(c i, It can be defined as r)560.

[0080] According to one embodiment, G(c i ,r)540 and R(c i ,r)560 Based on each, response set R i Cross-entropy loss L represents the probability of selecting a positive response from 520. CE570 can be calculated. In particular, the generative model G530, being based on a large-scale language model, has a higher probability of selecting affirmative responses than speech responses. As a result, the search model R550 maximizes the probability of selecting the correct response for randomly sampled speech responses by using the cross-entropy loss L CE It will learn to minimize 570. Next, G(c i ,r)540 and R(c i Considering r)560 as the logits of the teacher and student models respectively, the distillation loss L KD 580 can be derived. The search model R550 is distillation loss L KD To minimize 580, G(c i It is possible to train the models so that R(ci,r)540 and R(ci,r)560 match. On the other hand, in the embodiment, explaining that the models are trained to match scores is to explain the direction of training, and the scores of the two models do not necessarily have to match as a result of the training.

[0081] The following describes an evaluation and results regarding the use of the dialogue model training method described herein to conduct open-domain dialogues.

[0082] First, the dataset used will be an open-domain dialogue dataset consisting of Blended Skill Talk, ConvAI2, Empathetic Dialogues, and Wizard of Wikipedia. In the experiment, all four datasets will be used together, and the merged dataset can be referred to as BST+.

[0083] Human evaluation was performed on 200 example questions randomly sampled from the BST+ test dataset. Human reviewers assessed the quality of the generated responses using two criteria on a 0-2 scale. First, they assessed Appropriacy (Appr.) to determine whether the generated responses were fluent, logical, and appropriate to the given context, and second, they assessed Informationality (Info.) to determine whether the generated responses contained meaningful information relevant to the given context. Each example was evaluated by a minimum of three unique human reviewers, and all human evaluations were performed via Amazon Mechanical Turk.

[0084] Furthermore, the experiment can report a variety of automated metrics. MaUdE is an unreferenced dialogue response evaluation metric calculated by a model trained to score syntactically and semantically negative responses as 0 and positive responses as 1 using the ConvAI2 dataset. Because MaUdE shows a high correlation with human judgments of response fluency and interest, it is used as a proxy metric to evaluate the overall quality of responses generated by each model. In the experiment, the Dist-2 and Dist-3 models can also be used to measure the lexical diversity of the generated responses, where Dist-n represents the ratio of unique n-grams to the total number of n-grams of all responses generated by each model. The length, which is the average number of tokens in the generated responses, is reported for reference. Finally, in the experiment, the latency for generating responses to a single input context is measured and reported to validate the efficiency of the models in this disclosure. Generally, latency measured in a GPU-assisted environment is reported, but latency measured using only the CPU may also be reported.

[0085] This disclosure compares the results of a dialogue model training method at the model level with those of a generative model using knowledge distillation techniques, using a smaller Blender model extracted from a larger generative model. Here, a 400M parametric Blender model distilled in Blender 2.7B is used, along with TinyBERT-style distillation represented as Distilled Blender. Meanwhile, a 256M parametric biencoder and polyencoder, pre-trained on the Pushshift Reddit annotation dataset and fine-tuned on the BST+ dataset, can serve as a baseline for the search model. As previously mentioned, the biencoder model integrated with the MIPS library is represented as Bi-encoder(w / FAISS). RetNRef is an example-based generative model that integrates the responses of the search model into the input of the generative model. Unlike G2R, one of the dialogue model training models in this disclosure, RetNRef leverages the search model to improve the generative model, while G2R leverages knowledge about the generative model to improve the search model. In particular, G2R uses a dialogue search model trained as an alpha-blending technique. In one embodiment, human responses to the dialogue model indicate empirically noted labels annotated on the BST+ dataset.

[0086] In the dialogue model training method of this disclosure, the biencoder R is trained as G2R using Blender 9.4B as the supervising model G. G2R-DM represents a model trained as both data-level G2R and model-level G2R. In this disclosure, two types of variations are considered for comprehensive analysis. For example, G2R-D is trained only as data-level G2R, and G2R-D(excluding FAISS) is G2R-D with the additional removal of the use of FAISS, which is a MIPS library.

[0087] Table 1 shows the results of human evaluation and automated metrics for several dialogue models of open-domain dialogue. Here, the Latency (Speedup) column shows the relative speed improvement of each model compared to the latency of Blender 90M. In Table 1, it can be seen that the system trained as the dialogue model training method (G2R) of this disclosure achieved the "sweet spot" between dialogue ability and efficiency. The system of this disclosure significantly improves human evaluation results while maintaining the low latency of a bi-encoder (w / FAISS), achieving human evaluation scores that are similar to or better than those of Blender 90M and human responses, respectively.

[0088] [Table 1]

[0089] Further examination reveals that, as seen from the Dist-2 and Dist-3 scores, the blender generation model and the distilled blender model exhibit high human evaluation scores but relatively long latency along with a lack of diversity. The Retrieval baselines (Bi-encoder and Poly-encoder) show the opposite trend, exhibiting much lower latency and relatively high response diversity, but relatively lower conversational ability in terms of human evaluation scores. Contrary to the human evaluation results, the MaUdE scores of the Bi-encoder and Poly-encoder are unexpectedly high. However, this result is because the MaUdE items were trained on the ConvAI2 dataset, a subset of the BST+ dataset, and these retrieval models have similar training goals. The G2R-based model in this disclosure achieves far better human evaluation results than the original model, the Bi-encoder (w / FAISS). Applying the G2R-specific data level (G2R-D) significantly improves model performance, allowing it to be compared to the Gold Human Response in terms of human evaluation. Using the G2R data level, a predefined set of responses R GBecause the number of responses increases by more than 10 times, using a bi-encoder without FAISS (G2R-D (w / o FAISS)) will increase latency. When the size of the response set is small (in the case of a bi-encoder (w / FAISS)), using FAISS will introduce latency overhead, but when using FAISS with a larger response set, such as G2R-D, low latency can be maintained. However, the response quality may be slightly lower compared to the version without FAISS.

[0090] The additional application of model-level G2R can further improve the performance of search models. G2R-DM, trained as both data-level G2R and model-level G2R, exhibits even higher human evaluation scores and MaUdE scores than G2R-D, trained only as data-level G2R, and is significantly faster to run while achieving human evaluation scores comparable to the Blender 90M model. G2R-DM shows slightly lower human evaluation scores compared to larger Blender generative models, but exhibits considerably lower latency (23.0 times faster than distilled Blender models, and 44.7 times faster than Blender 2.7B). Furthermore, G2R-DM shows significantly higher response diversity compared to Blender generative models. On the other hand, the RetNRef model performs worse than the G2R-DM model and offers significantly higher latency.

[0091] Table 2 shows the original response set R and the new response set R generated by the data level G2R. G The basic statistics are shown. After applying the G2R data level, R G This will have approximately 11 times more candidates compared to the original response set R. G To determine if the responses exhibit greater diversity compared to the original response set R, we can calculate the number of unique tokens and bigrams / trigrams shown in each response set. See Table 2 for the augmented response set R. GBecause it has far more unique tokens and bigrams / trigrams than the original response set, it can handle a wider variety of subjects and entities and exhibit greater diversity in terms of syntax and expression.

[0092] [Table 2]

[0093] In the following, we will conduct an ablation study to analyze in detail how the model's performance changes depending on how the responses generated from the G2R method at the data level are used. The responses generated from the G2R at the data level are used for training the search model R on the dialogue dataset D. G Increases the enhanced response set R G This is used to construct [the model]. Through ablation research, these two types of ablation methods can be separated, and the model can be evaluated when using only each method.

[0094] Table 3 shows the evaluation results of such ablation models. To evaluate the performance of the search model along with human evaluation metrics and automated metrics, we utilize the Hits@1 / K and Hits@5 / K biencoder models trained on the widely adopted BST+ test set. In Table 3, from the top row to the second to last row, the results are as follows: training the search model using the existing dialogue dataset D to construct the original response set R (i.e., the existing dialogue model); training the search model using the existing dialogue dataset D to augment the response set R by model level G2R. G When constructing (i.e., an ablation model), the search model is augmented by the G2R data level in the dialogue dataset D G When training using and constructing the original response set R (i.e., the ablation model), and when training the search model on the dialogue dataset D augmented by the data level G2R, G The response set R was trained using the model level G2R and augmented by the model level G2R. GThis is the result of human evaluation and automated metrics related to constructing (i.e., the G2R-DM of this disclosure).

[0095] [Table 3]

[0096] As can be seen in Table 3, models using only one method do not show better performance than models using both methods. G Leveraging the responses generated for construction improves the model's fit score, supporting the dialogue model training method of this disclosure that using a diverse set of responses helps the model respond better. The augmented dialogue dataset D for building R. G Using it helps improve human evaluation scores for all appropriateness and informationality metrics. Also, the augmented dialogue dataset D G Training using this method significantly improves the Hits metric of the search model. Nevertheless, since using both methods together results in the best human evaluation performance among all ablation models, it is clear that using new examples to train the search model and build the response set is crucial for achieving good performance.

[0097] In Table 3, the augmented dialogue dataset generated by training a biencoder model using the top m responses of an already trained search model is D. R This can be done. Augmented Dialogue Dataset D R Dialogue model using and dialogue dataset D augmented by data level G2R G Comparing dialogue models using these methods, we can confirm that using large-scale generative models generates better quality educational datasets than simply using search models. As seen in Table 3, D improves all human evaluation scores and query metrics. G Unlike when using D RUsing this as a training dataset does not result in corresponding performance improvements across all metrics. These results strongly suggest that using large-scale generative models for dialogue augmentation, as in data-level G2R, is a far more effective augmentation strategy than using search models.

[0098] In one embodiment, the log-likelihood score (LL score) is used to define the score G(c,r) of the supervised generation model in the G2R model level, but other methods can also be used. One example is the use of the mutual information score (MI score). The MI score is the point-wise mutual information between a given context c and response r, and is known to assign lower values ​​to general responses while increasing the score of responses that are more specific to the given context. Using the MI score generates more specific and diverse responses compared to the LL score, but at the same time, there is a slightly higher risk of returning responses that contain inappropriate details regarding the input context. Therefore, below we compare the performance of a G2R model level that uses the MI score as G(c,r) and a model that uses the LL score.

[0099] Table 4 shows the results of human evaluation and automated metrics for the G2R model at model levels, which uses MI scores to define the score G(c,r) of the teacher-generated model.

[0100] [Table 4]

[0101] Using MI scores for model-level G2R results in slightly lower human assessment scores than using LL scores, particularly for suitability scores, meaning that MI scores may not generate appropriate and accurate responses. However, in terms of automated metrics, MI scores show higher MaUdE scores. That is, using MI scores shows greater response diversity compared to LL scores, indicating that MI scores can be used for a wider range of responses in dialogue systems.

[0102] Table 5 provides examples of responses returned by the baseline and G2R models. In the experiments related to Table 5, it was assumed that context B ("Oh yeah that is a fact") was input to the syntax of chatbot A ("Amazon is a great place to order books") as the input context.

[0103] [Table 5]

[0104] Referring to Table 5, in this example, the Bi-encoder (w / FAISS) returns a response irrelevant to the given context ("Comics is at the top of the list in interest"). The Blender model's response is logically correct, but it either simply changes the topic (Blender 90M, distilled Blender: "Do you have any other hobbies that you like to do on the weekends? I like to read", "What else do you like to do in your spare time? I like to go to the beach") or provides a result that is relatively lacking in certain details (Blender 2.7B, Blender 9.4B: "What kind of books do you buy on Amazon? I like to read science fiction and fantasy", "What kind of books do you like to read? I like romance novels and mystery novels"). The data level G2R (G2R-D) attempts to respond in detail, but the response may contain somewhat irrelevant phrases related to the subject (e.g., "Do you ever order groceries? I love Amazon's selection of books and videos"). In contrast, G2R-DM can be seen as providing an appropriate response with specific details (e.g., "What is your favorite book? Mine is 'the cat in the hat' by Dr. Suess").

[0105] On the other hand, when generating a dialogue model through a search model, the methods described in Figures 4 and 5 can be used in parallel.

[0106] Figure 6 is a block diagram showing an electronic device 10 for training a model-level dialogue model according to one embodiment.

[0107] The electronic device 10 may, in one embodiment, include a storage device 12 and a processor 14, but is not limited thereto.

[0108] The storage device 12 can store the context and responses necessary for open-domain interaction. The storage device 12 is hardware that stores various data processed within the electronic device 10 and can store programs for processing and control of the processor 14. The storage device 12 may include RAM (random access memory) such as DRAM (dynamic random access memory) and SRAM (static random access memory), ROM (read-only memory), EEPROM (electrically erasable programmable read-only memory), CD-ROM, Blu-ray or other optical disc storage, HDD (hard disk drive), SSD (solid state drive), or flash memory.

[0109] The processor 14 can control the overall operation of the electronic device 10 and process data and signals. In one embodiment, the processor 14 may include at least one other processor. In one embodiment, the processor 14 can acquire a response set from a first dialogue dataset via a storage device 12, which includes a first response subset corresponding to a first context and an optionally selected second response subset. The processor 14 can also calculate a first score for the responses included in the response set for the first context based on a first dialogue model, and a second score for the responses included in the response set for the first context based on a second dialogue model. The processor 14 can then train the second dialogue model based on the first and second scores.

[0110] The electronic device 10 of this disclosure may further include a communication device (not shown). The communication device can communicate with an external electronic device using wired or wireless communication technology and may include a transceiver. The external electronic device may be a terminal or a server. The communication technology used by the communication device may include, but is not limited to, GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), and others.

[0111] The electronic device according to the embodiment described above may include a memory for storing and executing program data, a permanent storage unit such as a disk drive, a communication port for communicating with external devices, and user interface devices such as a touch panel, keys, and buttons. The method, embodied as a software module or algorithm, may be stored on a computer-readable recording medium as computer-readable code or program instructions executable on the processor. Here, computer-readable recording media include magnetic storage media (e.g., ROM (read-only memory), RAM (random-access memory), floppy disks, hard disks, etc.) and optical reading media (e.g., CD-ROM, DVD (Digital Versatile Disc)). Computer-readable recording media may be distributed across a network of connected computer systems, and computer-readable code may be stored and executed in a distributed manner. The medium is computer-readable, can be stored in memory, and can be executed on the processor.

[0112] This embodiment may be presented as a functional block configuration and various processing stages. Such a functional block may be embodied as a variety of hardware and / or software configurations that perform a particular function. For example, the embodiment may employ an integrated circuit configuration such as memory, processing, logic, look-up table, etc., which can perform a variety of functions under the control of one or more microprocessors or other control devices. Just as the components may be executed as software programming or software elements, this embodiment includes a variety of algorithms that may be embodied as data structures, processes, routines, or combinations of other programming configurations, and may be embodied as programming or scripting languages ​​such as C, C++, Java, assembler, etc. The functional aspects may be embodied as algorithms executed on one or more processors. Furthermore, this embodiment may employ prior art for electronic environment configuration, signal processing, and / or data processing, etc. Terms such as “mechanism,” “element,” “means,” and “configuration” may be used broadly and are not limited to mechanical and physical configurations. The terms may also include the meaning of a series of software processes (routines) in conjunction with a processor, etc.

[0113] The embodiments described above are merely illustrative examples, and other embodiments may be embodied within the scope of the claims described later.

Claims

1. A method for training a dialogue model in an electronic device, A step of selecting a first context from a first dialogue data set containing one or more pairs of contexts and corresponding responses, A step of generating a first response corresponding to the first context through a first dialogue model, A step of generating an augmented dialogue dataset by including the first context and the corresponding first response pair in the first dialogue dataset, A method for training a dialogue model, comprising the step of training a second dialogue model based on the augmented dialogue dataset.

2. The dialogue model training method according to claim 1, wherein the first dialogue model is a generative dialogue model that generates a response to a given context, and the second dialogue model is a retrieval dialogue model that retrieves a response to the given context.

3. The dialogue model training method according to claim 1, further comprising the step of generating an augmented response set comprising the responses of the first dialogue dataset and the first responses.

4. The aforementioned learning stage is, The steps include obtaining a response set that includes a first response subset corresponding to a second context included in the augmented dialogue dataset, and an arbitrarily selected second response subset, A step of calculating a first score for the responses included in the response set for the second context based on the first dialogue model, A step of calculating a second score for the responses included in the response set for the second context based on the second dialogue model, A dialogue model training method according to claim 1, comprising the step of training the second dialogue model based on the first score and the second score.

5. The step of training the second dialogue model based on the first and second scores is: A step of calculating loss based on the first score and the second score, A dialogue model training method according to claim 4, comprising the step of training the second dialogue model so as to minimize the loss.

6. A method for training a dialogue model in an electronic device, The process involves obtaining a response set from a first dialogue dataset that includes a first response subset corresponding to a first context, and an arbitrarily selected second response subset. A step of calculating a first score for the responses included in the response set for the first context based on a first dialogue model, A step of calculating a second score for the responses included in the response set for the first context based on a second dialogue model, A dialogue model training method comprising the step of training the second dialogue model based on the first score and the second score.

7. The dialogue model training method according to claim 6, wherein the first dialogue model is a generative dialogue model that generates a response to a given context, and the second dialogue model is a retrieval dialogue model that retrieves a response to the given context.

8. The steps include selecting a second context from a second dialogue dataset containing one or more contexts and corresponding response pairs, A step of generating a response corresponding to the second context through the first dialogue model, A method for training a dialogue model according to claim 6, comprising the step of generating a first dialogue dataset by including the second context and corresponding response pairs in the second dialogue dataset.

9. The step of calculating the second score is: The steps include encoding the responses included in the first context and the response set as fixed-length embedding, A dialogue model training method according to claim 6, comprising the step of calculating a relevance score for each response in the response set to the first context based on an embedding value corresponding to the first context and an embedding value corresponding to each response in the response set.

10. The dialogue model training method according to claim 6, wherein the first score is calculated using a log-likelihood normalized based on the length of each response included in the response set.

11. The dialogue model training method according to claim 6, wherein the first score is calculated based on the Mutual Information score relating to the first context of each response included in the response set.

12. The stage of training the second dialogue model described above is: A step of calculating loss based on the first score and the second score, A dialogue model training method according to claim 6, comprising the step of training the second dialogue model so as to minimize the loss.

13. The dialogue model training method according to claim 12, wherein the loss includes a cross-entropy loss for scores corresponding to the first response subset and a knowledge distillation loss for scores corresponding to responses included in the response set.

14. The dialogue model training method according to claim 13, wherein the step of training the second dialogue model to minimize the loss is to train it to minimize the cross-entropy loss by maximizing the score corresponding to the first response subset.

15. The dialogue model training method according to claim 13, wherein the step of training the second dialogue model to minimize the loss is to train it so that the first score and the second score match and the knowledge distillation loss is minimized.

16. A computer-readable, non-temporary recording medium that stores a program for causing a computer to execute the dialogue model training method of claim 6.

17. An electronic device for training dialogue models, Storage device and It includes a control unit, The control unit, Through the storage device, a response set is obtained from the first dialogue dataset, including a first response subset corresponding to the first context and an arbitrarily selected second response subset. Based on the first dialogue model, calculate a first score for the responses included in the response set for the first context. Based on the second dialogue model, a second score is calculated for the responses included in the response set for the first context. An electronic device for training the second dialogue model based on the first score and the second score.

18. The control unit, Select a second context from a second dialogue dataset that contains one or more contexts and corresponding response pairs. Through the first dialogue model, a response corresponding to the second context is generated, The electronic device according to claim 17, wherein the first dialogue dataset is generated by including the second context and the corresponding response pair in the second dialogue dataset.

19. The control unit, An enhanced response set is generated, which includes the responses of the second dialogue dataset and the responses corresponding to the second context. The electronic device according to claim 18, which stores an enhanced set of responses through the storage device.

20. The control unit, in order to train the second dialogue model, Based on the first score and the second score, calculate the loss. The second dialogue model is trained to minimize the aforementioned loss. The electronic apparatus according to claim 17, wherein the loss includes a cross-entropy loss for the scores corresponding to the first response subset and a knowledge distillation loss for the scores corresponding to the responses included in the response set.

Citation Information

Patent Citations

  • Method for processing information, information processor, and program

    JP2019082987A