Dialogue delivery method and dialogue processing system
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NAVER CORP
- Filing Date
- 2023-06-21
- Publication Date
- 2026-08-07
AI Technical Summary
【0013】 上述のように、本発明に係る対話提供方法及び対話処理システムは、メモリに格納されたユーザー履歴を用いてユーザーと対話を行うことで、ユーザーにカスタマイズされた対話を提供することができる。
Smart Images

Figure 0007902295000010 
Figure 0007902295000011 
Figure 0007902295000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method for providing dialogue, a method for processing dialogue, and a system for conducting dialogue based on user information. [Background technology]
[0002] The dictionary definition of artificial intelligence is a technology that uses computer programs to realize human learning abilities, reasoning abilities, perceptual abilities, and natural language comprehension abilities. Such artificial intelligence has made remarkable progress through deep learning.
[0003] In particular, with the development of artificial intelligence, various language models have been developed. These language models not only recognize text and understand its meaning, but have also reached a level where they can extract and classify information from large amounts of text-based data, such as documents, and even directly generate text.
[0004] Such language models are actively used in a variety of fields, including search services, document creation (e.g., resume writing, report writing, and post creation), free-flowing dialogue across various categories, data analysis on given text (e.g., data summarization and classification), provision of expertise, programming, and converting given texts into appropriately styled texts.
[0005] Recently, agents that provide conversational capabilities are being used to proactively offer services to users in various fields such as shopping, search, healthcare, and counseling. However, such agents only consider the content of the conversation with the user in the current conversation session, and there are limitations to considering past conversations between the agent and the user. As a result, users have to take proactive actions such as repeatedly providing their information in each different conversation session with the agent or correcting what the agent says without considering the user's situation, which causes inconvenience to the user. [Overview of the project] [Problems that the invention aims to solve]
[0006] The present invention provides a dialogue provision method and a dialogue processing system that can perform appropriate dialogue between a user and an agent, reflecting the user's information.
[0007] Specifically, the present invention provides a dialogue provision method and a dialogue processing system that enable an agent to guide an appropriate dialogue according to the user's state or situation using the user's history information.
[0008] Furthermore, the present invention provides a dialogue provision method and a dialogue processing system that can store important information about a user and use that stored important information to interact with the user.
[0009] Furthermore, the present invention provides a dialogue analysis method and system that can systematically manage a user's state using the content of past conversations between the user and an agent, as well as a user monitoring method and system using the same. [Means for solving the problem]
[0010] To solve the above problems, the dialogue provision method according to the present invention may include the steps of: forming a dialogue session between an agent and a user; generating an utterance for the agent using user history relating to a previous dialogue session formed before the dialogue session; and engaging in dialogue with the user by providing the utterance for the agent to the user.
[0011] Furthermore, the dialogue processing system according to the present invention may include: a memory for storing user history related to past dialogue sessions; a summarizer that receives user utterances in the current dialogue session formed between an agent and a user and summarizes at least a portion of the user utterances in sentence form; and a memory operator that uses the summary information summarized by the summarizer and the user history to specify an action to be taken on the memory.
[0012] Furthermore, a program executed by one or more processes on an electronic device and stored on a computer-readable recording medium may include commands for performing the steps of: forming an interactive session between an agent and a user; generating an utterance for the agent using a user history relating to a previous interactive session formed before the current one and stored in association with the user's account; and interacting with the user by providing the user with the agent's utterance. [Effects of the Invention]
[0013] As described above, the dialogue provision method and dialogue processing system according to the present invention can provide a customized dialogue to the user by engaging in dialogue with the user using the user history stored in memory.
[0014] More specifically, the dialogue provision method and dialogue processing system according to the present invention store the user's utterances from previous dialogue sessions as user history and use this to engage in dialogue with the user, thereby enabling natural dialogue with the user based on the latest information corresponding to the user history.
[0015] Furthermore, in this invention, by interacting with the user based on their user history, it is possible to monitor or check the user's situation or status according to their user history.
[0016] On the other hand, the dialogue provision method and dialogue processing system according to the present invention can summarize user utterances using a summarization unit that has been trained to summarize only important user utterances from among the user's utterances in a dialogue session between the user and the agent. This prevents the indiscriminate consumption of memory resources and enables the provision of a new dialogue session with the user based on important information about the user. [Brief explanation of the drawing]
[0017] [Figure 1] This is a conceptual diagram illustrating the dialogue processing method and dialogue processing system according to the present invention. [Figure 2a] This is a conceptual diagram illustrating the dialogue processing method and dialogue processing system according to the present invention. [Figure 2b] This is a conceptual diagram illustrating the dialogue processing method and dialogue processing system according to the present invention. [Figure 2c] This is a conceptual diagram illustrating the dialogue processing method and dialogue processing system according to the present invention. [Figure 3] This is a conceptual diagram illustrating the dialogue processing method and dialogue processing system according to the present invention. [Figure 4] This is a conceptual diagram illustrating a method for processing dialogue in the summarization unit of the dialogue processing system according to the present invention. [Figure 5] This is a conceptual diagram illustrating a method for processing dialogue using a memory operator in the dialogue processing system according to the present invention. [Figure 6] It is a conceptual diagram for explaining a method of processing an interaction by a memory operator of an interaction processing system according to the present invention. [Figure 7] It is a conceptual diagram for explaining a method of processing an interaction by a memory operator of an interaction processing system according to the present invention. [Figure 8] It is a conceptual diagram for explaining a method of processing an interaction by a memory operator of an interaction processing system according to the present invention. [Figure 9] It is a conceptual diagram for explaining a method of processing an interaction by a memory operator of an interaction processing system according to the present invention. [Figure 10] It is a conceptual diagram for explaining a method of generating an interaction by a generation unit of an interaction processing system according to the present invention. [Figure 11] It is a conceptual diagram for explaining a method of generating an interaction by a generation unit of an interaction processing system according to the present invention. [Figure 12] It is a conceptual diagram for explaining a user monitoring system 1200 according to the present invention.
Embodiments for Carrying Out the Invention
[0018] Hereinafter, the embodiments disclosed in this specification will be described in detail with reference to the accompanying drawings. However, the same or similar components are given the same reference numerals regardless of the reference signs, and redundant descriptions thereof are omitted. The suffixes "module" and "unit" for the components used in the following description are merely given or mixed for the purpose of easily creating the specification, and do not have meanings or roles that are distinguished from each other by themselves. Further, when it is determined that a specific description of related known technologies makes the gist of the embodiments disclosed in this specification unclear when explaining the embodiments disclosed in this specification, the detailed description thereof is omitted. Further, the accompanying drawings are merely for easily understanding the embodiments disclosed in this specification, and the technical idea disclosed in this specification is not limited by the accompanying drawings, and should be understood to include all modifications, equivalents, and alternatives included in the idea and technical scope of the present invention.
[0019] Terms including ordinal numbers such as "1st," "2nd," etc., can be used to describe various components, but the components are not limited by these terms. These terms are used solely for the purpose of distinguishing one component from another.
[0020] When it is mentioned that one component is “connected” or “linked” to another component, it should be understood that it may be directly connected or linked to the other component, but there may also be other components between them. On the other hand, when it is mentioned that one component is “directly connected” or “directly linked” to another component, it should be understood that there are no other components between them.
[0021] A singular expression includes plural forms unless the context clearly indicates otherwise.
[0022] In this application, terms such as “includes” or “having” should be understood to indicate the presence of features, figures, steps, actions, components, parts, or combinations thereof as described in the specification, and not to preemptively exclude the possibility of the presence or addition of one or more other features, figures, steps, actions, components, parts, or combinations thereof.
[0023] The present invention provides a dialogue provision method and a dialogue processing system that can conduct appropriate dialogue between a user and an agent, reflecting the user's information. Specifically, the present invention provides a dialogue provision method and a dialogue processing system that can use the user's history information to enable an agent to guide an appropriate dialogue according to the user's state or situation.
[0024] As shown in Figure 1, the agent may be included as a function of various types of electronic devices 20, or as a function of a website, application, or software that provides various services in which interaction between the user and the agent may take place, such as conversational services, care services, or counseling services.
[0025] The format of the dialogue 30 between the user and the agent can vary; for example, the dialogue may take the form of voice or chat. For the sake of clarity, we will not distinguish whether the dialogue is voice or text (e.g., chat). Furthermore, regardless of the format of the dialogue, dialogue generated by the user will be referred to as user utterances or user speech, and dialogue generated by the agent will be referred to as agent utterances or agent speech. On the other hand, agents that interact with users are sometimes called "bots" or "chatbots."
[0026] The dialogue processing system 100 according to the present invention is a memory-management-based dialogue system that can be used for long-term dialogues when multiple dialogues are conducted between a user and an agent with time differences. As shown in Figures 2a, 2b, and 2c, when multiple dialogue sessions (e.g., Session 1, Session 2, Session 3) are formed between a user and an agent, the system provides a method for constructing the dialogue of the current dialogue session using the dialogue content of previously formed dialogue sessions.
[0027] The dialogue processing system 100 according to the present invention can receive a dialogue between a user and an agent and perform a series of processes to store information regarding the content of the dialogue in a memory 130. In the present invention, the information regarding existing dialogue content stored in the memory 130 can be referred to as "user history".
[0028] For example, let's assume that the first dialogue session (e.g., Session 1, Figure 2a), the second dialogue session (e.g., Session 2, Figure 2b), and the third dialogue session (e.g., Session 3, Figure 2c) are dialogue sessions that occurred in order from the first dialogue session.
[0029] If a second dialogue session is in progress, user history 221, 222 based on user-agent dialogues 201, 202 in a previous first dialogue session (e.g., Session 1, Figure 2a) may be used in the second dialogue session agent's utterance 204. Furthermore, if a dialogue session corresponding to a third dialogue session is in progress, user history 221, 222, 223, 224 based on user-agent dialogues 201, 202, 203, 204 corresponding to at least one of the first and second dialogue sessions may be used in the third dialogue session agent's utterance 206.
[0030] As shown in the figure, at least a portion of the user's utterances 201 in the first dialogue session may be stored in memory 130 as user history 221, 222. Furthermore, the dialogue processing system 100 may use the user history stored in memory 130 to generate the agent's utterances 204 in the second dialogue session formed between the user and the agent after the first dialogue session.
[0031] When the first dialogue session ends, the dialogue processing system 100 may store the contents of at least a portion of the dialogue from the first dialogue session in text format in the memory 130. Furthermore, if a second dialogue session takes place between the user and the agent after the first dialogue session, the dialogue processing system 100 may use one of the texts corresponding to the user history to generate an utterance from the agent regarding this.
[0032] For example, the agent may check the user's status or condition in relation to the user history "having a sore throat due to a cold" 221 stored in a previous conversation session (or past conversation session), and generate the utterance "How is your sore throat?" 204a.
[0033] As another example, in response to user history 222 which corresponds to "planning to go to the hospital," an agent utterance 204b could be generated asking "What did the hospital tell you?" to check whether the user has gone to the hospital.
[0034] Similarly, when the second dialogue session ends, the dialogue processing system 100 may store information regarding at least a portion of the content of the dialogue that took place in the second dialogue session as user history 223, 224 in the memory 130. The user history stored in the memory 130 may also be used in a third dialogue session that takes place after the second dialogue session.
[0035] Thus, the dialogue processing system 100 according to the present invention manages and utilizes the content of dialogues that take place in multiple dialogue sessions between the user and the agent in memory, thereby enabling continuous monitoring and management of various states (e.g., health, sleep, etc.) or situations (e.g., living situation, employment situation, etc.) of the user, and enabling more natural and appropriate dialogue with the user.
[0036] On the other hand, in the dialogue processing system 100 according to the present invention, the memory 130 may be updated so that the user's latest information on the same topic or category is maintained. That is, the user history stored based on past dialogue sessions may be updated based on the dialogue of the current dialogue session.
[0037] For example, in the first dialogue session, the user history contains the content (or text; hereafter, for the sake of explanation, the term "text" will be used, but it does not necessarily have to be in the form of text) "has a sore throat due to a cold" 221. If, in the second dialogue session, it is analyzed that the user's throat condition has improved, the user no longer has a sore throat, so the text "has a sore throat due to a cold" 221 may be deleted, and memory 130 may be updated.
[0038] As a similar example, the user history contains the entry "I have a hospital appointment" 222. In this case, if the dialogue 203b from the second dialogue session analyzes that the user has already gone to the hospital, the entry "I have a hospital appointment" 222 no longer needs to be stored in the user history, and this sentence can be deleted.
[0039] As described above, the dialogue processing system according to the present invention stores information that should be remembered in relation to the user from the dialogue session between the user and the agent as user history in memory, and may delete unnecessary information. Furthermore, in the next dialogue session, the agent's utterances can be generated using the user history, enabling natural dialogue based on the user's latest situation or state.
[0040] For this purpose, the dialogue processing system 100 according to the present invention may include a Summarizer 110, a Memory Operator 120, a Memory 130, and a Generator 140. Furthermore, the dialogue processing system 100 may be configured to further include a retriever 150.
[0041] As shown in Figures 3 and 4, the summarization unit 110 can receive the dialogue content D of a dialogue session between an agent and a user and generate a summary 115. The dialogue of the Nth dialogue session may be transmitted to the summarization unit 110 for processing after the Nth dialogue session has ended. The entity transmitting the dialogue to the summarization unit 110 may be a service server that provides dialogue services, but the present invention is not particularly limited to this.
[0042] As shown in the figure, the summarization unit 110 receives a dialogue D that includes the agent's utterance and the user's utterance, and the summarization unit 110 can generate a summary 115 based on the dialogue D.
[0043] More specifically, the summarization unit 110 can summarize the information from the dialogue that should be remembered in relation to the user in natural language sentences.
[0044] The summarization unit 110 may consist of a language model trained to summarize information from dialogue D that should be remembered in relation to the user in natural language sentence form. For example, a pre-trained language model that has already been trained on various types of information may be used to generate a summary (for example, a summary sentence (hereinafter, for the sake of explanation, the term "summary sentence" will be used, but it does not necessarily have to be in the form of a summary sentence)) when dialogue is input. Preferably, the language model may be trained to generate the summary sentence using newline as a delimiter.
[0045] Specifically, the summarization model that summarizes user information to be stored in the dialogue record D using various natural language sentences S={S1, S2, ..., Sk} is the correct summary sentence (gold summary sentence:
number
number
[0046] The summarization unit 110 may be trained to generate summary texts only for pre-set categories (category or topic). For example, the pre-set categories may be categories relating to various states or situations of the user. As an example, the pre-set categories may be related to health, sleep, exercise, diet, work, etc.
[0047] In this case, based on the dialogue content in dialogue D related to the health category, such as "My throat is fine now, but I have a slight headache" 301, the summarization unit 110 can generate summary information such as "My throat is fine, but I have a headache" 311 or "My throat was sore, but it's better now, and I still have a headache."
[0048] Furthermore, based on dialogue content from dialogue D, such as "I said let's wait and see a little longer. I have another hospital appointment next week" 302, the summarization unit 110 can generate summary information such as "has a hospital appointment" 312.
[0049] Furthermore, based on dialogue content in dialogue D, such as "I haven't been able to sleep at all lately" 303, the summarization unit 110 can generate summary information such as "a state of difficulty sleeping" 313.
[0050] On the other hand, the summarization unit 110 has been trained to generate summary texts only for user utterances in a dialogue session that correspond to pre-set categories. As a result, it does not need to generate summary texts for user utterances that correspond to categories other than the pre-set categories. For example, as shown in Figure 4, the summarization unit 110 can generate summary texts 421, 422, 431, 432, and 433 for user utterances 401, 402, 411, 412, and 413 in a dialogue session (such as Session 1 or Session 2) that correspond to pre-set categories related to health, sleep, exercise, diet, or work. Furthermore, it does not need to generate summary texts for utterances in categories other than the pre-set categories, such as the "weather" category (e.g., utterance 403: "It's been really hot lately, I really hate the heat").
[0051] On the other hand, when a dialogue is received, the summarization unit 110 can use a summarization model to generate summary texts of the user utterances and agent utterances that constitute the dialogue for pre-set categories. Specifically, the summarization model may be a language model that has been trained to receive dialogue and category information (e.g., "health," "sleep") as input and generate summary information about the corresponding category from the dialogue content. Therefore, the summary texts summarized by the summarization unit 110 can exist with information matching which category the content of the summary text corresponds to. The memory where the summary texts are stored may contain summary texts for each pre-set category. Before summarizing the texts, the summarization unit 110 may first classify the categories of the texts, and as a result, it can generate summary texts only for texts classified into pre-set categories. Therefore, data resources can be saved by not generating summary texts for texts that do not require summarization.
[0052] Next, the Memory Operator 120 can control the operation of the memory 130 so that the user history (or user information) stored in the memory 130 maintains the most up-to-date information about the user.
[0053] The memory 130 may be located inside or outside the dialogue processing system 100 (for example, an external server, cloud server, or cloud storage). As shown in Figure 3, the memory operator 120 can determine an action to be taken on the memory 130 using the summary text (or summary information 311, 312, 313) summarized by the summarization unit 110 and the user history (specifically, the texts 321, 322, 323 that constitute the user history) that are pre-stored in the memory 130.
[0054] As shown in Figure 3, the user history stored in memory 130 may consist of the dialogue content of previous dialogue sessions formed between the user and the agent before the Nth dialogue session was formed. The user history stored in memory 130 may consist of summary sentences 321, 322, and 323, which are summaries of at least a portion of the dialogue from previous dialogue sessions, as summarized by the summarization unit 110. The user history may also include content related to the user's state or circumstances.
[0055] Memory 130 can be updated by an operation specified by the memory operator 120. By the specified operation, memory 130 can either i) store at least a portion of the summary information in memory 130, or ii) delete at least a portion of the stored user history.
[0056] The memory operator 120 can be controlled to perform one of two different operations on memory with respect to pairs of summary sentences summarized from the dialogue of an interaction session and summary sentences included in the user history stored in memory. When an operation is performed on the summary sentence for the dialogue of the Nth interaction session, the user history stored in memory 130 may be updated to reflect the content of the dialogue of the Nth interaction session.
[0057] We will examine the different operations (operations) performed on the memory 130 as defined in this invention for the summary text m included in the user history stored in memory, and the summary text s for a new dialogue session.
[0058] The first action may mean maintaining the storage of m in memory 130, but not storing s in memory 130 (PASS). The first action may occur even if the contents of the two sentences are identical or similar, or if the contents of s are included in the contents of m. In this way, the first action can be performed when there is no need to update memory.
[0059] For example, as shown in Figure 3, the memory operator 120 can maintain the user history stored in memory 130 as is, for example, the summary sentence "Planning to go to the hospital" 322, which corresponds to the user history stored in memory 130, and the summary sentence of the current dialogue session, "Hospital appointment made" 312.
[0060] The second action could mean maintaining the storage of m in memory 130 and also storing s in memory 130 (APPEND). The second action can handle cases where the contents of m and s are unrelated or where s contains additional information.
[0061] For example, as shown in Figure 3, there is no relationship between the summary sentence "Went jogging" 323, which corresponds to the user history stored in memory 130, and the summary sentence "Has trouble sleeping" 313, which is the summary sentence of the current dialogue session. Therefore, the memory operator 120 can control the operation of memory 130 so that the summary sentence "Has trouble sleeping" 313 is newly added to memory 130.
[0062] The third operation could mean deleting m from memory 130 and storing s in memory 130 (REPLACE). In other words, m in memory 130 can be replaced with s. The third operation occurs when the contents of two sentences do not match or contradict each other, in which case the memory operator 120 deletes the information previously stored in memory 130 in order to maintain the user history as the most up-to-date information for the user. For example, as shown in Figure 3, given the summary sentence "I have a sore throat from a cold" 321 and the summary sentence of the dialogue session "My throat is fine, but I have a headache" 311, which are stored in memory 130, in the content of the dialogue in the Nth dialogue session, the user no longer has a sore throat and has switched to a headache state, so the memory operator 120 can control the operation of memory 130 so that "My throat is fine, but I have a headache" 311 is stored in memory 130 instead of "I have a sore throat from a cold" 321.
[0063] The fourth action may mean deleting m from memory 130 and not storing s in memory 130 either (DELETE). In the case of the fourth action, the content of the text may no longer reflect the user's state or situation. For example, if the user history contains the summary sentence "I took cold medicine" and the summary sentence "My cold is gone" from the Nth dialogue session, the user has recovered from their cold and no longer needs cold medicine. In this case, memory 130 no longer needs to store information about the user related to the cold.
[0064] As shown in Figure 5, the memory operator 120 can identify a memory operation for the summary text summarized in the dialogue session and the user history stored in memory, based on any of the first to fourth operations.
[0065] If the first dialogue session (Session 1) is the first (initial or first) dialogue session associated with the user, then no user history may exist in Memory (Memory 1). In this case, the result of the operation of the memory operator 120 on both the first dialogue session (Session 1) and the user history may both be the second operation, "APPEND". Therefore, the summary text (Summary 1) summarized in the first dialogue session (Session 1) may be stored in Memory (Memory 2) as is.
[0066] On the other hand, if a second dialogue session (Session 2) takes place after a first dialogue session (Session 1), the summarization unit 110 can receive the dialogue of the second dialogue session and generate a summary document (Summary 2) concerning the dialogue of the second dialogue session. Furthermore, the memory operator 120 can update memory 130 using the user history (Memory 2) stored in memory 130 and the summary document (Summary 2) for the second dialogue session. By performing specific operations on memory 130 in relation to the user history (Memory 2) and the summary document (Summary 2) concerning the dialogue of the second dialogue session, a user history (Memory 3) reflecting the second dialogue session (Session 2) can be constructed.
[0067] Figure 6 shows a simplified memory update algorithm in the memory operator 120. The memory update process of the present invention can maintain the user's latest information by combining existing information and new information using the above-described operation (operator).
[0068] n memory documents stored in memory M:
number
number
number
[0069] To find M', m i ∈M, s j ∈S sentence pairs (m i ,s j A method can be used to classify the relationship between (m). The memory operator 120 is (m i ,s j For this, one of the first to fourth actions described above, {"PASS", "REPLACE", "APPEND", "DELETE"}, is determined.
[0070] The memory update unit can update the memory to M'. On the other hand, according to one embodiment of the present invention, instead of comparing all pairs between user history and summary texts, the operation of the memory can be specified only for texts that correspond to the same category.
[0071] The summary texts may be stored in memory 130, categorized according to their respective categories. Therefore, the memory operator 120 can specify the operation of memory 130 only for summary texts that belong to the same category.
[0072] As shown in Figure 7, if there are pre-configured categories 1 through 4, the memory operator 120 can compare the sentences corresponding to each category as a pair. For each category, the memory operator 120 can identify one of the first through 4 actions (PASS, APPEND, REPLACE, DELETE) described above for the input sentences as a pair. As a result, the user history stored in memory 130 can be updated for each category.
[0073] On the other hand, the memory 130 may contain summary texts as user history for each pre-configured category. This is to maintain only the most recent information regarding the user's status or condition for each category.
[0074] As described above, the memory operator 120 may be configured using a classification model trained to predict or identify the operation of memory 130 corresponding to one of the first to fourth operations for a pair of sentences. The dataset for training the model may consist of a pair of sentences corresponding to m (or premise sentence) and s (or hypothesis sentence), and labels indicating which of the first to fourth operations (PASS, APPEND, REPLACE, DELETE) the pair of sentences corresponds to, as shown in Figures 8 and 9.
[0075] The memory operator 120 may be learned based on a pair of sentences and one of the first to fourth actions corresponding to the pair of sentences (for example, each of which is mapped to a single token corresponding to the digits 0 to 3).
[0076] As a result of learning by the method described above, the memory operator 120 becomes able to predict or identify the operation of memory 130 that corresponds to one of the first to fourth operations for a pair of sentences.
[0077] The generating unit (Generator) 140 is configured to generate the agent's utterance using the user history stored in the memory 130.
[0078] More specifically, as shown in FIG. 10, in the present invention, the process (S1010) of forming a dialogue session between the agent and the user can be performed. Further, the process of generating the agent's utterance using the user history can be performed (S1020). As discussed above, the user history may be configured based on information extracted from previous dialogue sessions formed between the agent and the user.
[0079] On the other hand, in the memory 130, there may be corresponding user histories for each user account. The generating unit 140 can generate the agent's utterance by referring to the user history stored in association with the user account of the currently conversing user.
[0080] The generating unit 140 can generate the agent's utterance using at least a part of the user history stored in the memory 130 and the dialogue history in the current session. The dialogue history D at time step t t can be expressed as follows (c is the agent's utterance and u is the user's utterance).
Number
Number
Number
number
[0081] If the user history contains multiple summary sentences corresponding to multiple different categories related to the user's state or situation (see reference numerals 1111, 1112, 1113, and 1114 in Figure 11), the generation unit 140 can generate the agent's utterance using one of the multiple summary sentences based on the context of the dialogue in the currently ongoing dialogue session. As shown in Figure 11, when the dialogue session D2 is started between the user and the agent, all or part of the multiple summary sentences 1111, 1112, 1113, and 1114 corresponding to the user history of the user stored in the memory 130 are transmitted to the generation unit 140 and may be used to generate the agent's utterance in the currently ongoing dialogue session.
[0082] The search unit 150 can select a portion of the multiple summary sentences stored in the memory 130 and transmit them to the generation unit 140. In some cases, the configuration of the search unit 150 may be omitted, and all of the multiple summary sentences stored in memory may be transmitted to the generation unit 140.
[0083] The generation unit 140 does not have to use the multiple summary sentences that make up the user history to generate the agent's utterances if there is no summary sentence among them that corresponds to the context of the dialogue in the currently ongoing dialogue session. In other words, even if a user history exists, the generation unit 140 does not have to use content unrelated to the context of the dialogue in the agent's utterances. Thus, in the present invention, the process of engaging in dialogue with the user can be carried out by providing the user with agent utterances generated based on the user history (S1030).
[0084] Thus, according to the present invention, user utterances from previous dialogue sessions are stored as user history, and by using this to engage in dialogue with the user, it is possible to have a natural dialogue with the user based on the latest information corresponding to the user history. Furthermore, in the present invention, by engaging in dialogue with the user based on the user history, it is possible to monitor or check the user's situation or state according to the user history.
[0085] According to one embodiment of the present invention, a user monitoring method and system can be provided that manages the status of a user by periodically interacting with a specific user through a call connection or the like. Referring to Figure 12, the user monitoring system 1200 according to the present invention may be configured to include at least one of a call processing system 1210, a management system 1220, a dialogue analysis system 1230, and a storage unit 1240. Each of these components may be operated independently, and conceptually, the functions performed by a combination of these can be described as being performed by a user monitoring method or user monitoring system.
[0086] The dialogue analysis system 1230 can analyze the user's state or situation using the acquired dialogue. The analysis results can be provided to the administrator by the management system 1220, enabling user monitoring.
[0087] The call processing system 1210 is responsible for initiating calls to users and interacting with users through the connected calls. It can initiate calls to users and obtain conversations according to policies set in the management system 1220 (for example, call management policies or call initiation policies).
[0088] The call processing system 1210 may include a dialogue processing unit 1211, a call connection unit 1212, a speech synthesis unit 1213, and a speech recognition unit 1214.
[0089] The dialogue processing unit 1211 provides a dialogue function with the user connected to the call. Based on a language model that has been trained on various information, the dialogue processing unit 1211 can generate appropriate responses to the user's utterances and engage in dialogue with the user. In this invention, various user utterances can be collected through unstandardized, open dialogues utilizing the language generation model, and these can be analyzed to confirm the user's state.
[0090] The specific method for generating a dialogue in the dialogue processing unit 1211 according to the present invention is as described in the generation unit 140 of the dialogue processing system 100, and in this case, other components of the dialogue processing system 100 can correspond to other components of the user monitoring system 1200 (for example, the storage unit 1240, the memory model 1233, etc.).
[0091] Furthermore, the dialogue processing unit 1211 can set up an agent persona so that the user feels they are talking to a real person who empathizes with and cares about their story. It can also train a language model to speak according to scenarios designed to correspond to the set persona. Additionally, the dialogue processing unit 1211 may be designed to use listening techniques employed in the dialogue and to ask follow-up questions at an appropriate level to the user's responses in order to express that the agent is listening to the user.
[0092] The call connection unit 1212 may be configured to initiate calls to users. The call connection unit 1212 can initiate calls to users based on policies related to call initiation. Call initiation policies can be configured through the management system 1220.
[0093] The speech synthesis unit 1213 can convert text into speech so that the agent's utterances generated by the dialogue processing unit are output as speech. The speech synthesis unit 1213 can express natural speech using speech processing technology (for example, utilizing NES (Natural End-to-end Speech Synthesis) and HDTs (High-quality DNN Text-to-Speech) technologies in a hybrid manner). The speech synthesis unit 1213 can learn counselor voices that respond to various call situations. For example, a bright and cheerful voice may be learned as the basic voice, or it may be learned to speak in a voice that shows empathy and concern for the user's situation depending on the situation.
[0094] The speech recognition unit 1214 may be configured to recognize the user's speech utterances and convert them into text. For example, the speech recognition unit 1214 can utilize speech recognition technology that leverages an advanced big language model trained on a large amount of diverse and extensive data. Furthermore, it can be trained to have performance tailored to the user's age characteristics and regional characteristics, taking into account the user's characteristics.
[0095] The speech recognition unit 1214 can recognize the user's voice using one of several speech recognition models specialized for different characteristics (for example, characteristics defined by criteria such as region and age group). For example, it can recognize the user's voice better by using one of several speech recognition models specialized for specific regional dialects or the inaccurate pronunciation of elderly people. On the other hand, which of the multiple models to use can be determined based on the administrator's selection in the management system 1220 or the target user of the call.
[0096] The conversation acquired by the call processing system 1210 can be transmitted to at least one of the management system 1220, the conversation analysis system 1230, and the storage unit 1240.
[0097] The management system 1220 can set policies for calls made to users and provide information about the current status of calls made to users or the user's status. The management system 1220 can acquire information about matters that should be checked to understand the user's status (e.g., health, sleep, diet, exercise, going out, etc.) and provide the acquired information to the administrator. The management system 1220 can receive information analyzed from the dialogue analysis system 1230 and provide it to the administrator, and if special matters that require checking or monitoring (e.g., health abnormality signals) are detected, it can notify the administrator or guardian. For example, based on the dialogue between the user and the agent, the management system 1220 can monitor users who require management, understand abnormal situations, emergencies, etc., and provide a function to respond quickly (for example, confirming information that an elderly person is waiting for 119 and contacting them separately, or confirming information that a meal has not been delivered and taking appropriate action).
[0098] The management system 1220 may include a management policy setting unit 1221, an analysis model setting unit 1222, and a screen processing unit 1223.
[0099] The management policy setting unit 1221 can manage policies (or "call outgoing policies") for calls made to users. The management policy setting unit 1221 can set policies for at least one of the following: the user to whom a call is made, the call outgoing time, and the call outgoing cycle. In this invention, a policy is a unit of execution for making a call, and one policy may include one or more users (recipients), outgoing settings (e.g., outgoing time, outgoing frequency, outgoing cycle, etc.), and reporting targets. Furthermore, one or more groups may be added to one policy (e.g., outgoing groups by day of the week) to manage users separately.
[0100] The management policy setting unit 1221 can configure multiple policies, and each policy may be specified to apply to at least one user. Policies may be configured based on various criteria (e.g., the scope of a specific area) at the administrator's discretion. For example, a policy may be configured based on "Magok-dong, Gangseo-gu, Seoul" and applied to users living there. Furthermore, a particular policy may have multiple subgroups based on it (for example, a policy based on "Gangseo-gu" may have subgroups such as the "Magok-dong" group and the "Balsan-dong" group, which are subgroups based on multiple areas included in Gangseo-gu).
[0101] The analysis model setting unit 1222 can set an analysis model for analyzing the dialogue acquired from the call processing system 1210. Once the analysis model is set by the analysis model setting unit 1222, the information of the set analysis model can be transmitted to the dialogue analysis system 1230. The dialogue analysis system 1230 can analyze the dialogue using the analysis model based on the received information and transmit the analysis results to the management system 1220.
[0102] The screen processing unit 1223 can provide various information related to calls and users based on information received from the call processing system 1210 and the dialogue analysis system 1230. For example, it can visually provide current status information or statistical information of outgoing calls (e.g., total number of outgoing calls, number of completed calls, number of answered calls, number of unanswered calls, etc.) and status information for the user. In addition, user history generated by the memory model 1233 may be provided.
[0103] Furthermore, the administration page allows administrators to configure settings based on their choices, including the policy settings, group settings, and analysis model settings mentioned above.
[0104] The dialogue analysis system 1230 may include various functions for analyzing dialogue, such as a user state model 1231, an emergency notification model 1232, and a memory model 1233. The dialogue analysis system 1230 can obtain dialogue analysis results by inputting the dialogue into each model. When the dialogue session between the user and the agent ends, the call processing system 1210 transmits the dialogue obtained from the ended dialogue session to the dialogue analysis system 1230, which can then analyze it.
[0105] The user state model 1231 can analyze (determine or perceive) the user's state from the content of the dialogue. The user state model 1231 may consist of a classification model that has been trained to determine the user's state for a specific category. For example, the user state model 1231 may be trained to classify the user's state as positive, negative, or unknown (or irrelevant) for each category such as health, diet, sleep, exercise, and going out. The categories subject to state determination (or classification) can be set based on the administrator's selection in the management system 1220.
[0106] The emergency notification model 1232 is configured to identify the user's emergency (or abnormal) situation from the interaction. The emergency notification model 1232 may also be configured to extract key abnormal signals, such as emergency situations that require monitoring by an administrator. For example, the emergency notification model 1232 can be implemented by using a deep learning model trained to classify predefined emergency situations (e.g., health-related danger utterances), or by extracting summary information (or summary sentences) from the user's utterances through slot processing. Information regarding the emergency situation determined by the emergency notification model 1232 can be transmitted to the management system 1220 and provided to the administrator.
[0107] The memory model 1233 ensures that information to be remembered about the user is stored as user history during interactions between the user and the agent. This user history may also be used by the call processing system 1210 when generating agent utterances. This allows the agent to reduce the repetition of the same questions in each interaction session and to increase familiarity with the user by conducting interactions based on the user's information. The memory model 1233 can update the user history based on the most recent interaction session between the user and the agent to maintain the user's most up-to-date information on the same topic or category.
[0108] As a specific embodiment of the memory model 1233, the method used in the above-described dialogue processing system 100 to generate a dialogue using the user history can be used.
[0109] On the other hand, the present invention described above can be implemented as a program that is executed by one or more processes on a computer and can be stored on a medium (or recording medium) that is readable by such a computer.
[0110] Furthermore, the present invention described above can be implemented as computer-readable code or instruction words on a medium on which a program is recorded. In other words, the present invention can be provided in the form of a program.
[0111] On the other hand, computer-readable media include all types of recording devices that store data readable by a computer system. Examples of computer-readable media include HDDs (Hard Disk Drives), SSDs (Solid State Disks), SSDs (Silicon Disk Drives), ROMs, RAMs, CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices.
[0112] Furthermore, the computer-readable medium may include storage and may be a server or cloud storage accessible by electronic devices via communication. In this case, the computer can download the program according to the present invention from the server or cloud storage via wired or wireless communication.
[0113] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, i.e., a CPU (Central Processing Unit), and its type is not particularly limited.
[0114] On the other hand, the above detailed description should not be interpreted as restrictive in all respects, but rather should be considered illustrative. The scope of the invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the scope of the equivalents of the invention are included within the scope of the invention.
Claims
1. A method of providing dialogue, The steps include: forming a dialogue session between the agent and the user, The steps include generating the agent's utterance using the user history of a previous dialogue session formed before the aforementioned dialogue session, The step includes providing the user with the agent's utterances, thereby engaging in dialogue with the user. The user history consists of a summary of the user's utterances during a previous dialogue session, which was formed between the agent and the user prior to the dialogue session. The summary content includes information relating to the user's state or circumstances, The method of providing the dialogue is When the aforementioned dialogue session ends, the dialogue from the dialogue session is transmitted to a summarizer. The summarization unit further includes the step of summarizing in sentence form specific utterances from the user's utterances included in the dialogue of the dialogue session that correspond to a pre-set category, The summarization unit is trained to generate summary content only for the user utterances included in the dialogue of the dialogue session that correspond to the pre-set categories. The method of providing the dialogue is The steps include obtaining an output value corresponding to one of the different operations related to updating the user history stored in memory, based on the relationship between the summaries of the summary content of the dialogue session and the summary content of the user history stored in memory, The step further includes updating the user history stored in the memory with the output value, Methods for providing dialogue.
2. In the aforementioned user history, If there are multiple summary contents corresponding to multiple different categories related to the user's state or situation, In the step of generating the agent's utterance, The dialogue provision method according to claim 1, characterized in that, based on the context of the dialogue in the dialogue session, the agent's utterance is generated using any of the multiple summary contents that correspond to the context of the dialogue in the dialogue session.
3. The aforementioned user history is The method for providing dialogue according to claim 1, characterized in that summary content exists for each of several different categories.
4. Of the aforementioned distinct operations, the first operation is: The operation maintains that the summary content corresponding to the user history among the pairs of summary content mentioned above is stored in the memory, but the summary content summarized from the dialogue of the dialogue session is not stored in the memory. Of the aforementioned mutually distinct operations, the second operation is: This operation maintains that the summary content corresponding to the user history among the pairs of summary content mentioned above is stored in the memory, and also stores the summary content summarized from the dialogue of the dialogue session in the memory. Of the aforementioned distinct operations, the third operation is: This operation involves deleting the summary content corresponding to the user history from the memory among the pairs of summary content, and storing the summary content summarized from the dialogue of the dialogue session in the memory. Of the aforementioned distinct operations, the fourth operation is: The dialogue provision method according to claim 1, characterized in that, among the pairs of summary contents, the summary content corresponding to the user history is deleted from the memory, and the summary content summarized from the dialogue of the dialogue session is also not stored in the memory.
5. A memory that stores user history related to past dialogue sessions formed between an agent and a user, wherein the user history consists of a summary of the user's utterances during the past dialogue session, and the summary includes content related to the user's state or situation, and A summarizer receives the dialogue of the current conversation session formed between the agent and the user, and summarizes at least a portion of the user utterances included in the dialogue in text form. A memory operator that uses the summary information summarized by the summarization unit and the user history to identify one of several different operations relating to updating the user history on the memory, The summary information and the user history include the content summarized by the summarization unit, The memory operator uses pairs of specific content included in the summary information and specific content included in the user history to determine the operation on the memory based on the relationship between each of the specific contents. Dialogue processing system.
6. The dialogue processing system according to claim 5, characterized in that the pair of contents described above corresponds to the same category.
7. The dialogue processing system according to claim 6, characterized in that, based on the operation on the memory, the memory is updated so that the user history reflects the dialogue of the current dialogue session.
8. The system further includes a Generator that generates the utterances of the agent, The generating unit is The dialogue processing system according to claim 7, characterized in that, after the current dialogue session has ended, the agent generates an utterance using the updated user history in a newly formed dialogue session between the agent and the user.
9. A program executed by one or more processes on an electronic device and stored on a computer-readable recording medium, This program is The steps include: forming a dialogue session between the agent and the user, The steps include generating the agent's utterance using the user history of a previous dialogue session formed before the aforementioned dialogue session, The steps include: a step of interacting with the user by providing the user with the agent's utterances; and a command word for performing this step. The user history consists of a summary of the user's utterances during a previous dialogue session, which was formed between the agent and the user prior to the dialogue session. The summary content includes information relating to the user's state or circumstances, This program is When the aforementioned dialogue session ends, the dialogue from the dialogue session is transmitted to a summarizer. The summarization unit further includes a command for performing the step of summarizing in sentence form a specific utterance from the user's utterances included in the dialogue of the dialogue session that corresponds to a pre-set category, The summarization unit is trained to generate summary content only for the user utterances included in the dialogue of the dialogue session that correspond to the pre-set categories. This program is The steps include obtaining an output value corresponding to one of the different operations related to updating the user history stored in memory, based on the relationship between the summaries of the summary content of the dialogue session and the summary content of the user history stored in memory, The instruction further includes the step of updating the user history stored in the memory based on the output value, program.
Citation Information
Patent Citations
Dialogue management device, method, and program, and consciousness extraction system
JP2009193532A
Interaction device and interaction method
JP2020118842A
JPP6882975B
Method and apparatus for providing counseling information
KR1020200072315A