Dialog providing method and dialog processing system

The dialogue system addresses the limitation of existing systems by using user history to generate appropriate dialogues, ensuring natural and efficient interactions by managing and updating user information, thus improving user experience.

JP2025522519AActive Publication Date: 2025-07-15NAVER CORP

Patent Information

Application Number
JP2024575055
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-20
Filing Date
2023-06-21
Publication Date
2025-07-15
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

Existing dialogue systems fail to consider past dialogue content between users and agents, requiring users to repeatedly provide information and correct agent content without considering their situation, causing inconvenience.

Method used

A dialogue providing method and system that utilizes user history from previous sessions to generate appropriate dialogues, including a memory to store user history, a summarizer to summarize important user utterances, and a memory operator to manage and update this information for natural and efficient dialogues.

Benefits of technology

Enables customized and natural dialogues by reflecting user history, reducing resource consumption and allowing continuous monitoring of user states and situations, thereby enhancing user interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025522519000001_ABST
    Figure 2025522519000001_ABST
Patent Text Reader

Abstract

The present invention relates to a dialogue providing method, a dialogue processing method, and a system for conducting a dialogue based on user information. The dialogue providing method according to the present invention may include steps of forming a dialogue session between an agent and a user, generating an utterance of the agent using user history regarding a previous dialogue session formed before the dialogue session, and conducting a dialogue with the user by providing the utterance of the agent to the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dialogue providing method, a dialogue processing method, and a system for conducting a dialogue based on user information.

Background Art

[0002] The dictionary definition of artificial intelligence is a technology that realizes human learning ability, reasoning ability, perception ability, natural language understanding ability, etc. by computer programs. Such artificial intelligence has made a leap in development through deep learning.

[0003] In particular, with the development of artificial intelligence, various language models have been developed. Such language models not only recognize text and understand its meaning, but also extract and classify information from data containing a large amount of text such as documents, and even reach the level of directly generating text.

[0004] Such language models are actively utilized in various fields. For example, search services, document creation (such as resume creation, report creation, post creation, etc.), free dialogue for various categories, data analysis with given text (such as data summarization, classification, etc.), provision of specialized knowledge, programming, conversion of given articles into articles with appropriate styles, etc. There are various fields that can be carried out based on text.

[0005] Recently, agents that provide a dialogue function have been actively providing services to users in various fields such as shopping, search, healthcare, and counseling services. However, in the case of such agents, only the dialogue content with the user conducted in the current dialogue session is considered, and there are limitations in considering the past dialogue content between the agent and the user. Therefore, the user has to take active actions such as repeatedly providing their information for each different dialogue session with the agent, or correcting the content spoken by the agent without considering the user's situation, which causes inconvenience to the user.

Summary of the Invention

Problems to be Solved by the Invention

[0006] The present invention aims to provide a dialogue providing method and a dialogue processing system that can conduct appropriate dialogue between a user and an agent by reflecting the user's information.

[0007] Specifically, the present invention aims to provide a dialogue providing method and a dialogue processing system that can lead an agent to conduct appropriate dialogue according to the state or situation of the user by using the user's history information.

[0008] Furthermore, the present invention aims to provide a dialogue providing method and a dialogue processing system that can remember important information about the user and conduct dialogue with the user by using the remembered important information.

[0009] Furthermore, the present invention aims to provide a dialogue analysis method and system that can systematically manage the state of the user by using the past dialogue content conducted between the user and the agent, and a user monitoring method and system using the same.

Means for Solving the Problems

[0010] In order to solve the above problems, the dialogue providing method according to the present invention may include: a step of forming a dialogue session between an agent and a user; a step of generating an utterance of the agent by using a user history regarding a previous dialogue session formed before the dialogue session; and a step of having a dialogue with the user by providing the utterance of the agent to the user.

[0011] Furthermore, the dialogue processing system according to the present invention may include: a memory that stores a user history regarding past dialogue sessions; a summarizer that receives a user utterance in a current dialogue session formed between an agent and a user and summarizes at least a part of the user utterance in a text form; and a memory operator that specifies an operation on the memory by using the summary information summarized by the summarizer and the user history.

[0012] Furthermore, a program executed by one or more processes in an electronic device and stored in a computer-readable recording medium may include instruction words for: a step of forming a dialogue session between an agent and a user; a step of generating an utterance of the agent by using a user history regarding a previous dialogue session formed before the dialogue session and associated with the user's account; and a step of having a dialogue with the user by providing the utterance of the agent to the user.

Advantages of the Invention

[0013] As described above, the dialogue providing method and the dialogue processing system according to the present invention can provide a customized dialogue to the user by having a dialogue with the user by using the user history stored in the memory.

[0014] More specifically, the dialogue providing method and the dialogue processing system according to the present invention store the user's utterances in previous dialogue sessions as user history, and by using this to conduct a dialogue with the user, a natural dialogue with the user can be carried out based on the latest information according to the user history.

[0015] Furthermore, in the present invention, by conducting a dialogue with the user based on the user history, the situation or state of the user according to the user history can be monitored or checked.

[0016] On the other hand, the dialogue providing method and the dialogue processing system according to the present invention can summarize the user's utterances by using a summarization unit learned to summarize only important user utterances among the user's utterances in the dialogue session between the user and the agent. Thereby, it is possible to prevent the consumption of memory resources indiscriminately, and based on the important information about the user, a new dialogue session with the user can be provided.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2a

Figure 2b

Figure 2c

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0018] Hereinafter, the embodiments disclosed in this specification will be described in detail with reference to the accompanying drawings. However, the same or similar components are given the same reference numerals regardless of the drawing reference numerals, and duplicate descriptions thereof are omitted. The suffixes "module" and "unit" for the components used in the following description are merely given or mixed for the purpose of easily creating the specification, and do not have meanings or roles that are distinguishable from each other. Also, when explaining the embodiments disclosed in this specification, if it is determined that the specific description of related known technologies obscures the gist of the embodiments disclosed in this specification, the detailed description thereof is omitted. Also, the accompanying drawings are merely for facilitating the understanding of the embodiments disclosed in this specification, and the technical idea disclosed in this specification is not limited by the accompanying drawings, and should be understood to include all modifications, equivalents, and alternatives included in the idea and technical scope of the present invention.

[0019] Terms including ordinal numbers such as the first, the second, etc. can be used to describe various components, but the above components are not limited by the above terms. The above terms are only used for the purpose of distinguishing one component from another component.

[0020] When it is mentioned that a certain component is "connected" or "coupled" to another component, it should be understood that it may be directly connected or coupled to the other component, but there may also be other components between them. On the other hand, when it is mentioned that a certain component is "directly connected" or "directly coupled" to another component, it should be understood that there are no other components between them.

[0021] Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0022] In this application, terms such as "comprising" or "having" are used to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0023] The present invention is for providing a dialogue providing method and a dialogue processing system capable of performing an appropriate dialogue between a user and an agent by reflecting user information. Specifically, the present invention is for providing a dialogue providing method and a dialogue processing system capable of guiding an appropriate dialogue according to the state or situation of a user by using the history information of the user.

[0024] As shown in FIG. 1, the agent may be included as a function of various types of electronic devices 20, or may be included as a function of a website, an application, or software that provides various services such as a dialogue service, a care service, or a counseling service, where dialogue between the user and the agent can be conducted.

[0025] The forms of the dialogue 30 conducted between the user and the agent are various. For example, the dialogue may be conducted in the form of voice or chat. For the sake of convenience of explanation, the form of the dialogue is not distinguished as to whether it is voice or text (e.g., chat). Furthermore, regardless of the form of the dialogue, the dialogue generated by the user is expressed as user utterance or the user's utterance, etc., and the dialogue generated by the agent side is expressed as agent utterance or the agent's utterance, etc. On the other hand, the agent that conducts dialogue with the user may also be called a "bot" or a "chatbot".

[0026] The dialogue processing system 100 according to the present invention is a dialogue system based on memory management that can be used for long-term dialogue, in which dialogue is conducted multiple times with a time difference between the user and the agent. As shown in FIGS. 2a, 2b, and 2c, when a plurality of dialogue sessions (e.g., Session 1, Session 2, Session 3) are formed between the user and the agent, a method is provided for constructing the dialogue of the current dialogue session using the dialogue content of the previously formed dialogue session.

[0027] The dialogue processing system 100 according to the present invention can perform a series of processes of receiving the dialogue conducted between the user and the agent and storing information regarding the content of the dialogue in the memory 130. In the present invention, the information regarding the existing dialogue content stored in the memory 130 can be expressed as a "user history".

[0028] For example, assume that a first dialogue session (e.g., Session 1, Figure 2a), a second dialogue session (e.g., Session 2, Figure 2b), and a third dialogue session (e.g., Session 3, Figure 2c) are dialogue sessions that occur in order from the first dialogue session.

[0029] When the second dialogue session is in progress, user histories 221 and 222 based on the dialogues 201 and 202 between the user and the agent in the previously conducted first dialogue session (e.g., Session 1, Figure 2a) may be used for the utterance 204 of the second dialogue session agent. Further, when the dialogue session corresponding to the third dialogue session is in progress, user histories 221, 222, 223, and 224 based on the dialogues 201, 202, 203, and 204 between the user and the agent corresponding to at least one of the first and second dialogue sessions may be used for the utterance 206 of the third dialogue session agent.

[0030] As shown in the figure, the content related to at least a part of the user's utterance 201 in the first dialogue session may be stored in the memory 130 as user histories 221 and 222. Further, the dialogue processing system 100 may generate the utterance 204 of the agent in the second dialogue session formed between the user and the agent after the first dialogue session by using the user histories stored in the memory 130.

[0031] When the first dialogue session ends, the dialogue processing system 100 may store the content related to at least a part of the dialogue of the first dialogue session in the memory 130 in text form. Also, when a second dialogue session is conducted between the user and the agent after the first dialogue session, the dialogue processing system 100 may generate the utterance of the agent related thereto by using any one of the texts corresponding to the user histories.

[0032] For example, the agent may generate an utterance "How is your sore throat cold?" 204a to check the user's state or situation with respect to the user history "having a sore throat due to a cold" 221 memorized in a previous dialogue session (or past dialogue sessions).

[0033] As another example, corresponding to the user history 222 corresponding to "planning to go to the hospital", the agent's utterance 204b "What did they tell you at the hospital?" may be generated to check whether the user has been to the hospital.

[0034] Similarly, when the second dialogue session ends, the dialogue processing system 100 may store information regarding at least a part of the content of the dialogue conducted in the second dialogue session in the memory 130 as user histories 223, 224. Also, the user histories stored in the memory 130 may be utilized in a third dialogue session conducted after the second dialogue session.

[0035] In this way, the dialogue processing system 100 according to the present invention manages and utilizes the content of the dialogues conducted in a plurality of dialogue sessions between the user and the agent in the memory, thereby enabling continuous monitoring and management of various states (e.g., health, sleep, etc.) or situations (e.g., living situation, employment situation, etc.) of the user, and enabling a more natural and appropriate dialogue with the user.

[0036] On the other hand, in the dialogue processing system 100 according to the present invention, the memory 130 may be updated so that the latest information of the user is maintained for the same topic or category. That is, the user history stored based on past dialogue sessions may be updated based on the dialogue of the current dialogue session.

[0037] For example, in the first dialogue session, as user history, content (or text (hereinafter, for convenience of explanation, the term "text" is used, but it does not necessarily have to be in the form of text)) such as "having a sore throat due to a cold" 221 is stored. At this time, if it is analyzed from the dialogue conducted in the second dialogue session that the user's throat condition has improved, since the user is no longer in a state of having a sore throat, the text "having a sore throat due to a cold" 221 may be deleted and the memory 130 may be updated.

[0038] As a similar example, content such as "planning to go to the hospital" 222 is stored as user history. At this time, if it is analyzed from the dialogue 203b conducted in the second dialogue session that the user has gone to the hospital, since there is no need to store the content "planning to go to the hospital" 222 in the user history anymore, this text may be deleted.

[0039] As described above, in the dialogue processing system according to the present invention, among the dialogues of the dialogue session between the user and the agent, information to be memorized in relation to the user is stored in the memory as user history, and unnecessary information may be deleted. Also, in the next dialogue session, by generating the agent's utterance using the user history, a natural dialogue based on the user's latest situation or state can be conducted.

[0040] For this purpose, the dialogue processing system 100 according to the present invention may include a summarizer 110, a memory operator 120, a memory 130, and a generator 140. Further, the dialogue processing system 100 may be configured to further include a retriever 150.

[0041] As shown in FIGS. 3 and 4, the summary unit 110 can receive the dialogue content D of the dialogue session conducted between the agent and the user and generate a summary 115. The dialogue of the Nth dialogue session may be transmitted to the summary unit 110 for processing after the Nth dialogue session ends. The entity that transmits the dialogue to the summary unit 110 may be a service server that provides a dialogue service, but the present invention is not particularly limited thereto.

[0042] As shown in the figure, a dialogue D including the utterances of the agent and the user is input to the summary unit 110, and the summary unit 110 can generate a summary 115 based on the dialogue D.

[0043] More specifically, the summary unit 110 can summarize, in the form of a natural language sentence, the information related to the user that should be memorized among the dialogue content.

[0044] The summary unit 110 may be composed of a language model trained to summarize, in the form of a natural language sentence, the information related to the user that should be memorized among the dialogue D. For example, using a pre-trained language model (Pre-trained Language Model) that has already been trained on various information and a language model tuned with a training dataset composed of the dialogue session and the main information that should be memorized in the dialogue session, when a dialogue is input, a summary content (for example, a summary sentence (hereinafter, for convenience of explanation, the term "summary sentence" is used, but it does not necessarily have to be in the form of a summary sentence)) may be generated. Preferably, the language model may be trained to generate a summary sentence using newline as a delimiter.

[0045] Specifically, the summary model that summarizes the user information to be memorized in the dialogue record D with various natural language sentences S = {S1, S2,..., Sk} is a gold summary sentence:

Number

Number

[0046] The summary part 110 may be learned to generate a summary text only for a preset category (or topic). For example, the preset category may be a category for various states or situations of the user. As an example, the preset category may be related to health, sleep, exercise, diet, work, etc.

[0047] In this case, based on the conversation content such as "My throat doesn't hurt now, but my head aches a little" 301 related to the health category in the conversation D, the summary part 110 can generate summary information such as "The throat is okay but the head aches" 311 or "The throat was painful but has improved and the head aches".

[0048] Furthermore, based on the conversation content such as "I said I would wait and see a little longer. I have a hospital reservation next week" 302 in the conversation D, the summary part 110 can generate summary information such as "In the state of having reserved a hospital" 312.

[0049] Also, based on the conversation content such as "I haven't been able to sleep at all recently" 303 related to the sleep category in the conversation D, the summary part 110 can generate summary information such as "In the state of not being able to sleep well" 313.

[0050] On the one hand, since the summarization unit 110 is trained to generate summary texts only for the content corresponding to a preset category among the user's utterances included in the conversation of the dialogue session, it may not be necessary to generate a summary text for the content corresponding to a category different from the preset category among the user's utterances. For example, as shown in FIG. 4, among the conversations of the dialogue session (such as Session 1 or Session 2), for the user utterances 401, 402, 411, 412, 413 regarding health, sleep, exercise, diet, or work corresponding to the preset category, the summarization unit 110 can generate summary texts 421, 422, 431, 432, 433. Also, for a conversation 403 regarding a category other than the preset category, for example, the category of "weather" ("It has been really hot and tough recently. I really hate the heat"), it may not be necessary to generate a summary text.

[0051] On the other hand, when a dialogue is received, the summarization unit 110 can use a summarization model to generate a summary text of the user utterances and agent utterances constituting the dialogue for a preset category. Specifically, the summarization model may be a language model that receives a dialogue and category information (such as "health", "sleep", etc.) as input and is trained to generate summary information regarding the corresponding category in the dialogue content. Therefore, information regarding which category the summary text summarized by the summarization unit 110 corresponds to can be matched and present. In the memory where the summary texts are stored, summary texts may be stored for each preset category. The summarization unit 110 may first classify the category of the text before summarizing the text, and as a result, generate a summary text only for the text classified into the preset category. Therefore, by not generating a summary text for texts for which summarization is unnecessary, data resources can be saved.

[0052] Next, the Memory Operator 120 can control the operation of the Memory 130 so that the user history (or user information) stored in the Memory 130 maintains the latest information about the user.

[0053] The Memory 130 may be located in at least one of the inside and outside of the dialogue processing system 100 (for example, an external server, a cloud server, or a cloud storage). As shown in FIG. 3, the Memory Operator 120 can specify the operation on the Memory 130 by using the summary text (or summary information 311, 312, 313) summarized by the Summarization Unit 110 and the user history (specifically, the texts 321, 322, 323 constituting the user history) pre-stored in the Memory 130.

[0054] As shown in FIG. 3, the user history stored in the Memory 130 may be constituted by the dialogue content of the previous dialogue session formed between the user and the agent before the Nth dialogue session is formed. The user history stored in the Memory 130 may be constituted by the summary texts 321, 322, 323 in which the Summarization Unit 110 summarizes at least a part of the dialogue of the previous dialogue session. The user history may include content related to the state or situation of the user.

[0055] The Memory 130 can be updated by the operation specified by the Memory Operator 120. By the specified operation, the Memory 130 can i) store at least a part of the summary information in the Memory 130, or ii) delete at least a part of the stored user history.

[0056] The memory operator 120 can be controlled to perform any one of different operations on the memory with respect to a pair of a summary text summarized from the conversation of an interactive session and a summary text included in the user history stored in the memory. When an operation on the summary text regarding the conversation of the N-th interactive session is performed, the user history stored in the memory 130 can be updated to reflect the content of the conversation of the N-th interactive session.

[0057] Consider different operations on the memory 130 defined in the present invention for the summary text m included in the user history stored in the memory and the summary text s for a new interactive session.

[0058] The first operation may mean an operation (PASS) of maintaining storing m in the memory 130 but not storing s in the memory 130. The first operation may be the case where the contents of the two texts are the same or similar, or the case where the content of s is included in the content of m. Thus, the first operation can be performed when there is no need to update the memory.

[0059] For example, as shown in FIG. 3, for the summary text "Planning to go to the hospital" 322 corresponding to the user history stored in the memory 130 and the summary text "Currently in the state of reserving a hospital" 312 of the current interactive session, the memory operator 120 can maintain the user history stored in the memory 130 as it is.

[0060] The second operation may mean an operation (APPEND) of maintaining storing m in the memory 130 and also storing s in the memory 130. The second operation can correspond to the case where the content of m and the content of s are not related to each other or the case where it is additional information.

[0061] For example, as shown in FIG. 3, since there is no relevance between the summary sentence "Jogging was done" 323 corresponding to the user history stored in the memory 130 and the summary sentence "Insomnia state" 313 of the current dialogue session, the memory operator 120 can control the operation of the memory 130 so that the summary sentence "Insomnia state" 313 of the dialogue session is newly added to the memory 130.

[0062] The third operation may mean an operation (REPLACE) of deleting m from the memory 130 and storing s in the memory 130. That is, m in the memory 130 can be replaced with s. The third operation is a case where the contents of the two sentences do not match or contradict each other, and the memory operator 120 deletes the information previously stored in the memory 130 in order to maintain the user history as the latest information of the user. For example, as shown in FIG. 3, for the summary sentence "Having a sore throat due to a cold" 321 corresponding to the user history stored in the memory 130 and the summary sentence "My throat is okay but my head hurts" 311 which is the summary sentence of the dialogue session, in terms of the content of the dialogue in the Nth dialogue session, the user's throat is no longer sore and has switched to a state where the head hurts, so the memory operator 120 can control the operation of the memory 130 so that "My throat is okay but my head hurts" 311 is stored in the memory 130 instead of "Having a sore throat due to a cold" 321.

[0063] The fourth operation may mean an operation (DELETE) of deleting m from the memory 130 and not storing s in the memory 130 either. The case corresponding to the fourth operation may be a case where the content of the sentence no longer reflects the user's state or situation. For example, as user history, there are a summary sentence "Took cold medicine" and a summary sentence "The cold has healed" in the Nth dialogue session. Since the user's cold has completely healed, cold medicine is no longer needed. In this case, there is no need to store information about the user related to the cold in the memory 130 anymore.

[0064] As shown in FIG. 5, the memory operator 120 can identify a memory operation by any of the first to fourth operations for the summary text summarized in the dialogue session and the user history stored in the memory.

[0065] If the first dialogue session (Session 1) is the first (initial or first) dialogue session related to the user, it may be assumed that there is no user history in the memory (Memory 1). In this case, the operation results of the memory operator 120 for the first dialogue session (Session 1) and the user history may both be "APPEND" of the second operation. Therefore, the summary text (Summary 1) summarized in the first dialogue session (Session 1) may be stored as it is in the memory (Memory 2).

[0066] On the other hand, when the second dialogue session (Session 2) is held after the first dialogue session (Session 1), the summary unit 110 can receive the dialogue of the second dialogue session and generate a summary text (Summary 2) regarding the dialogue of the second dialogue session. Also, the memory operator 120 can update the memory 130 using the user history (Memory 2) stored in the memory 130 and the summary text (Summary 2) for the second dialogue session. By the operation of the memory 130 identified for the user history (Memory 2) and the summary text (Summary 2) regarding the dialogue by the second dialogue session, a user history (Memory 3) reflecting the second dialogue session (Session 2) can be constructed.

[0067] FIG. 6 briefly shows a memory update algorithm in the memory operator 120. The memory update process of the present invention can maintain the latest information of the user by combining the existing information and the new information using the above-described operation (operator).

[0068] n memorized texts stored in the memory M: [Number] and k new summary sentences: [Number] when given these, memory operator 120 uses them to create a set of sentences that has no information loss, is consistent, and has no duplicates: [Number] and can find.

[0069] To find M', for m i ∈ M, s j ∈ S, a method for classifying the relationship of the sentence pair (m i , s j ) can be used. Memory operator 120 determines, for (m i , s j ), one of the values {"PASS", "REPLACE", "APPEND", "DELETE"} which are the above-described first through fourth operations.

[0070] The memory update unit can update the memory to M'. On the other hand, according to one embodiment of the present invention, instead of comparing all pairs between the user history and the summary sentences, it is possible to perform the operation of specifying the memory only for sentences corresponding to the same category as each other.

[0071] The summary sentences may be classified and stored in the memory 130 for each corresponding category. Therefore, the memory operator 120 can specify the operation of the memory 130 only for summary sentences between categories that are the same as each other.

[0072] As shown in FIG. 7, when there are preset first to fourth categories, the memory operator 120 can compare the sentences corresponding to each category in pairs. The memory operator 120 can identify any one of the above-described first to fourth operations (PASS, APPEND, REPLACE, DELETE) for the sentences input in pairs for each category. As a result, the user history stored in the memory 130 can be updated for each category.

[0073] On the other hand, a summary sentence may exist as a user history for each preset category in the memory 130. This is to maintain only the information regarding the latest situation or state of the user for each category.

[0074] As described above, the memory operator 120 may be configured using a classification model learned to predict or identify the operation of the memory 130 corresponding to any one of the first to fourth operations for a pair of sentences. As shown in FIGS. 8 and 9, the dataset for model learning may be composed of a pair of sentences corresponding to m (or premise sentence) and s (or hypothesis sentence), and a label indicating which operation among the first to fourth operations (PASS, APPEND, REPLACE, DELETE) the pair of sentences corresponds to.

[0075] The memory operator 120 may be learned based on a pair of sentences and a label corresponding to any one of the first to fourth operations for the pair of sentences (for example, mapped to a single token corresponding to each of the numbers 0 to 3).

[0076] As a result of being learned by the above method, the memory operator 120 can predict or identify the operation of the memory 130 corresponding to any one of the first to fourth operations for a pair of sentences.

[0077] The generator 140 is configured to generate the agent's utterance using the user history stored in the memory 130.

[0078] More specifically, as shown in FIG. 10, in the present invention, the process (S1010) of forming a dialogue session between the agent and the user can be performed. Further, the process of generating the agent's utterance using the user history can be performed (S1020). As discussed above, the user history may be configured based on information extracted from previous dialogue sessions formed between the agent and the user.

[0079] On the other hand, in the memory 130, there may be corresponding user histories for each user account. The generator 140 can generate the agent's utterance by referring to the user history stored in association with the user account of the currently conversing user.

[0080] The generator 140 can generate the agent's utterance using at least a part of the user history stored in the memory 130 and the dialogue history in the current session. The dialogue history D at time step t t can be represented as follows (where c is the agent's utterance and u is the user's utterance).

Number

Number

Number

Equation

[0081] When there are a plurality of summary sentences corresponding to different categories related to the user's state or situation as the user history (refer to reference numerals 1111, 1112, 1113, 1114 in FIG. 11), the generation unit 140 can generate the agent's utterance using any one of the plurality of summary sentences based on the context of the conversation in the currently ongoing dialogue session. As shown in FIG. 11, when a dialogue session D2 is started between the user and the agent, all or part of the plurality of summary sentences 1111, 1112, 1113, 1114 corresponding to the user history of the user stored in the memory 130 are transmitted to the generation unit 140 and may be used to generate the agent's utterance in the currently ongoing dialogue session.

[0082] The search unit 150 can select a part of the plurality of summary sentences stored in the memory 130 and transmit it to the generation unit 140. In some cases, the configuration of the search unit 150 may be omitted, and all of the plurality of summary sentences stored in the memory may be transmitted to the generation unit 140.

[0083] If there is no summary sentence corresponding to the context of the current dialogue session among the multiple summary sentences that make up the user history, the generation unit 140 may not use the multiple summary sentences to generate the agent's utterance. That is, even if there is a user history, the generation unit 140 may not use content that has no relation to the context of the dialogue for the agent's utterance. Thus, in the present invention, by providing the user with the agent's utterance generated based on the user history, the process of conducting a dialogue with the user can be carried out (S1030).

[0084] Thus, according to the present invention, by storing the user's utterance in the previous dialogue session as the user history and using this to conduct a dialogue with the user, a natural dialogue with the user can be carried out based on the latest information according to the user history. Furthermore, in the present invention, by conducting a dialogue with the user based on the user history, the situation or state of the user according to the user history can be monitored or checked.

[0085] According to an embodiment of the present invention, it is possible to provide a user monitoring method and system for managing the state of a user by conducting a dialogue periodically through a call connection or the like with a specific user. Referring to FIG. 12, the user monitoring system 1200 according to the present invention may be configured to include at least one of a call processing system 1210, a management system 1220, a dialogue analysis system 1230, and a storage unit 1240. Each configuration may be operated independently, and conceptually, the functions exhibited by the combination of these can be expressed as being executed by the user monitoring method or the user monitoring system.

[0086] The dialogue analysis system 1230 can analyze the state or situation of the user using the obtained dialogue. The analyzed result can be provided to the administrator by the management system 1220 to conduct user monitoring.

[0087] The call processing system 1210 plays a role of making a call to a user and interacting with the user through the call connected to the user, and can make a call to the user and obtain an interaction according to a policy (for example, a call management policy or a call origination policy) set in the management system 1220.

[0088] The call processing system 1210 may include an interaction processing unit 1211, a call connection unit 1212, a voice synthesis unit 1213, and a voice recognition unit 1214.

[0089] The interaction processing unit 1211 provides an interaction function with the user connected to the call. Based on a language model (Language Model) that has been learned for various information, the interaction processing unit 1211 can generate an appropriate response to the user's utterance and interact with the user. In the present invention, various user utterances can be collected through an open and non-standardized interaction using a language generation model, and the user's state can be confirmed by analyzing this.

[0090] The specific method of generating an interaction by the interaction processing unit 1211 according to the present invention is as described in the generation unit 140 of the interaction processing system 100. At this time, the other configurations of the interaction processing system 100 can correspond to the other configurations of the user monitoring system 1200 (for example, the storage unit 1240, the memory model 1233, etc.).

[0091] Furthermore, the interaction processing unit 1211 can set the persona of the agent so that the user feels as if they are talking to an actual person who empathizes with and cares about their own words. Also, the language model can be learned to speak according to a scenario designed to correspond to the set persona. Also, the interaction processing unit 1211 may be designed to use listening techniques utilized in the interaction or ask an appropriate level of follow-up questions to the user's answer in order to express that the agent is listening to the user's words.

[0092] The call connection unit 1212 may be configured to initiate a call to the user. The call connection unit 1212 can initiate a call to the user based on a policy related to call initiation. The call initiation policy can be set through the management system 1220.

[0093] The voice synthesis unit 1213 can play a role in converting text into voice so that the speech of the agent generated by the dialogue processing unit is output as voice. The voice synthesis unit 1213 can use voice processing technologies (for example, a hybrid of NES (Natural End-to-end Speech Synthesis) and HDTs (High-quality DNN Text-to-Speech) technologies) to express natural voices. The voice synthesis unit 1213 can learn the voices of counselors according to various call situations. For example, a bright and energetic voice may be learned as the basic voice, or it may be learned to speak with a voice that empathizes with and worries about the user's situation according to the situation.

[0094] The voice recognition unit 1214 may be configured to recognize the user's voice speech and convert it into text. For example, as the voice recognition unit 1214, voice recognition technology using an advanced large language model trained with diverse and large amounts of data can be used. Furthermore, considering the characteristics of the user, it can be trained to have performance suitable for the user's age characteristics and regional characteristics.

[0095] The voice recognition unit 1214 can recognize the user's voice using any one of the voice recognition models specialized for different characteristics (for example, characteristics defined by criteria such as region, age group, etc.). For example, the user's voice can be recognized better using any one of various voice recognition models specialized for dialects in a specific region, inaccurate pronunciations of the elderly, etc. On the other hand, which model among the multiple models to use can be specified based on the selection of the administrator in the management system 1220 or the user to whom the call is to be made.

[0096] The conversation obtained by the call processing system 1210 can be transmitted to at least one of the management system 1220, the conversation analysis system 1230, and the storage unit 1240.

[0097] The management system 1220 can set a policy for calls made to a user and provide information on the current situation regarding calls made to the user or information regarding the user's state. The management system 1220 can obtain information regarding matters to be checked (for example, health, sleep, diet, exercise, going out, etc.) in order to grasp the user's state, and provide the obtained information to the administrator. When the management system 1220 is transmitted the information analyzed by the conversation analysis system 1230 and provided to the administrator, and a special matter (for example, a health abnormality signal) that is determined to require checking or monitoring is detected, it can notify the administrator, the caregiver, etc. For example, based on the conversation conducted between the user and the agent, the management system 1220 can monitor users who require management to grasp abnormal situations, emergency situations, etc., and quickly respond (for example, confirm information that an elderly person is waiting for 119 and make a separate call, or confirm information that a bento has not been delivered and take relevant actions, etc.).

[0098] The management system 1220 may include a management policy setting unit 1221, an analysis model setting unit 1222, and a screen processing unit 1223.

[0099] The call policy setting unit 1221 can manage the policy (or "call transmission policy") for calls sent to users. The call policy setting unit 1221 can set policies for at least one of the users to whom calls are to be transmitted, the call transmission time, the transmission cycle, etc. The policy in the present invention is, as an execution unit of call transmission, one policy may include one or more users (recipients), transmission settings (e.g., transmission time, transmission frequency, transmission cycle, etc.), the person to whom reports are to be made, etc. Further, one or more groups may be added to one policy (e.g., transmission groups by day of the week), and the users may be set to be managed separately.

[0100] The call policy setting unit 1221 can set a plurality of policies, and each policy may be specified to be applied to at least one user. The policy may be set according to various criteria (e.g., the scope of a specific region) based on the selection of the administrator. For example, a policy may be set based on "Magok-dong, Gangseo-gu, Seoul", and may be set to apply the policy to the users living there. Also, a specific policy may have a plurality of groups further classified based on this (e.g., for the policy based on "Gangseo-gu", there may be a "Magok-dong" group, a "Palbong-dong" group, etc. classified based on a plurality of regions included in Gangseo-gu).

[0101] The analysis model setting unit 1222 can play a role in setting an analysis model for analyzing the conversations acquired from the call processing system 1210. When the analysis model is set by the analysis model setting unit 1222, the information of the set analysis model can be transmitted to the conversation analysis system 1230. In the conversation analysis system 1230, the conversations can be analyzed using the analysis model based on the received information, and the analysis results can be transmitted to the management system 1220.

[0102] The screen processing unit 1223 can provide various information related to calls and users based on the information received from the call processing system 1210 and the dialogue analysis system 1230. For example, it can visually provide the status information or statistical information of the outgoing calls (such as the total number of outgoing calls, the number of completed calls, the number of answered calls, the number of unanswered calls, etc.), and the status information of the users. Also, the user history generated by the memory model 1233 may be provided.

[0103] Furthermore, on the management page, settings based on the administrator's selection, such as the above-mentioned policy settings, group settings, and analysis model settings, are possible.

[0104] The dialogue analysis system 1230 may include various functions for analyzing dialogues, such as a user state (USER STATE) model 1231, an emergency notification model 1232, and a memory model 1233. The dialogue analysis system 1230 can obtain the dialogue analysis results by inputting the dialogue into each model. When the dialogue session between the user and the agent ends, the call processing system 1210 transmits the dialogue obtained from the ended dialogue session to the dialogue analysis system 1230, and the dialogue analysis system 1230 can analyze this.

[0105] The user state model 1231 can analyze (judge or sense) the user's state from the content of the dialogue. The user state model 1231 may be composed of a classification model learned to judge the user's state for a specific category. For example, the user state model 1231 is learned to classify the user's state as positive, negative, or unknown (or irrelevant) for each category such as health, diet, sleep, exercise, going out, etc. The category to be the object of state judgment (or classification) can be set based on the administrator's selection in the management system 1220.

[0106] The emergency notification model 1232 is configured to grasp the user's emergency situation (or abnormal situation) from the conversation. The emergency notification model 1232 may be configured to extract major abnormal signals such as emergency situations that require administrator monitoring. For example, the emergency notification model 1232 can use a deep learning model trained to classify defined emergency situations (e.g., health-related dangerous utterances), or can be realized by a method such as slot-processing summary information (or summary text) for the user's utterance to extract it. Information on the emergency situation determined by the emergency notification model 1232 can be transmitted to the management system 1220 and provided to the administrator.

[0107] The memory model 1233 is configured to store information to be remembered for the user as the user history in the conversation conducted between the user and the agent. The user history may be used when generating the agent's utterance in the call processing system 1210. Thereby, the agent can reduce the repetition of the same question for each user dialogue session, and by conducting the dialogue based on the user's information, the intimacy with the user can be further enhanced. The memory model 1233 can update the user history based on the latest dialogue session between the user and the agent so that the latest information of the user is maintained for the same topic or category.

[0108] As a specific embodiment of the memory model 1233, the method used to generate a dialogue using the user history in the above-described dialogue processing system 100 can be used.

[0109] On the other hand, the present invention described above can be executed by one or more processes in a computer and can be realized as a program storable in a computer-readable medium (or recording medium) such as this.

[0110] Furthermore, the present invention described above can be realized as computer-readable code or instruction words on a medium on which a program is recorded. That is, the present invention can be provided in the form of a program.

[0111] On the other hand, a computer-readable medium includes all kinds of recording devices in which data readable by a computer system is stored. Examples of computer-readable media include HDD (Hard Disk Drive), SSD (Solid State Disk), SDD (Silicon Disk Drive), ROM, RAM, CD-ROM, magnetic tape, floppy (registered trademark) disk, optical data storage device, and the like.

[0112] Furthermore, the computer-readable medium may include storage and may be a server or cloud storage accessible by an electronic device via communication. In this case, the computer can download the program according to the present invention from the server or cloud storage via wired or wireless communication.

[0113] Furthermore, in the present invention, the computer described above is an electronic device equipped with a processor, that is, a CPU (Central Processing Unit, central processing unit), and its type is not particularly limited.

[0114] On the other hand, the above detailed description should not be construed as limiting in all respects, but should be considered as exemplary. The scope of the present invention should be determined by a reasonable interpretation of the appended claims, and all modifications within the equivalent scope of the present invention are included in the scope of the present invention.

Claims

1. A step of forming an interactive session between an agent and a user; A step of generating an utterance of the agent by using a user history regarding a previous interactive session formed before the interactive session; A step of interacting with the user by providing the utterance of the agent to the user, the method for providing interaction.

2. The user history is Composed of a summary content obtained by summarizing the utterances of the user during the previous interactive session formed between the agent and the user before the interactive session, The summary content includes content regarding the state or situation of the user, the method for providing interaction according to Claim 1.

3. In the user history, When there are a plurality of summary contents respectively corresponding to a plurality of different categories related to the state or situation of the user, In the step of generating an utterance of the agent, Based on the context of the interaction of the interactive session, among the plurality of summary contents, an utterance of the agent is generated by using any one of the summary contents corresponding to the context of the interaction of the interactive session, the method for providing interaction according to Claim 2.

4. The user history is Characterized in that there are summary contents for each of a plurality of different categories, the method for providing interaction according to Claim 1.

5. When the interactive session ends, a step of transmitting the interaction of the interactive session to a summarizer; In the summarizer, a step of summarizing, in a text form, specific utterances corresponding to a preset category among the utterances of the user included in the interaction of the interactive session, the method for providing interaction according to Claim 1.

6. The summarizer is Characterized in that it is learned to generate summary contents only for the content corresponding to the preset category among the utterances of the user included in the interaction of the interactive session, the method for providing interaction according to Claim 5.

7. A step of obtaining an output value corresponding to a specific operation corresponding to any one of different operations on the memory for a pair of the summary content summarized from the interaction of the interactive session and the summary content included in the user history stored in the memory; The method for providing a dialogue according to claim 6, further comprising the step of updating the user history stored in the memory according to the output value.

8. Among the different operations, the first operation is Among the pairs of summary contents, it is an operation of maintaining storing the summary content corresponding to the user history in the memory, but not storing the summary content summarized from the dialogue of the dialogue session in the memory. Among the different operations, the second operation is Among the pairs of summary contents, it is an operation of maintaining storing the summary content corresponding to the user history in the memory and also storing the summary content summarized from the dialogue of the dialogue session in the memory. Among the different operations, the third operation is Among the pairs of summary contents, it is an operation of deleting the summary content corresponding to the user history from the memory and storing the summary content summarized from the dialogue of the dialogue session in the memory. Among the different operations, the fourth operation is The method for providing a dialogue according to claim 7, characterized in that, among the pairs of summary contents, it is an operation of deleting the summary content corresponding to the user history from the memory and not storing the summary content summarized from the dialogue of the dialogue session in the memory either.

9. A memory (Memory) for storing user history regarding past dialogue sessions, A summarizer that receives the dialogue of the current dialogue session formed between the agent and the user and summarizes at least a part of the user utterances included in the dialogue in text form, A dialogue processing system including a memory operator that specifies an operation on the memory using the summary information summarized by the summarizer and the user history.

10. The summary information and the user history include the content summarized by the summarizer, In the memory operator, The dialogue processing system according to claim 9, characterized in that an operation on the memory is specified using a pair of specific content included in the summary information and specific content included in the user history.

11. The dialogue processing system according to claim 10, characterized in that the pair of the content corresponds to the same category.

12. The dialogue processing system according to claim 11, wherein based on operations on the memory, the memory is updated such that the user history reflects the dialogue of the current dialogue session.

13. further comprising a generator (Generator) for generating utterances of the agent, wherein the generator, after the current dialogue session ends, generates an utterance of the agent using the updated user history in a newly formed dialogue session between the agent and the user, the dialogue processing system according to claim 12.

14. A program executed by one or more processes in an electronic device and stored in a computer-readable recording medium, wherein the program, includes steps of forming a dialogue session between an agent and a user, generating an utterance of the agent using a user history regarding a previous dialogue session formed before the dialogue session, and steps of interacting with the user by providing the utterance of the agent to the user, a program stored in a computer-readable recording medium, characterized by including instruction words for performing the steps.

Citation Information

Patent Citations

  • Dialogue management device, method, and program, and consciousness extraction system

    JP2009193532A

  • Interaction device and interaction method

    JP2020118842A

  • Dialogue scenario generation device, program, and method capable of determining context from dialogue logs

    JP6882975B2

  • Golf putter that irradiate laser to guide swing angle

    KR102367966B1

Cited By

  • Program, information processing device, method, and system

    JP7828517B1

  • Method for providing service by artificial intelligence agent with multi-turn and artificial intelligence agent using the same

    US12536164B1