Electronic device and method for providing chatbot service by using multi-session conversation model based on chronological dynamics

The multi-session conversation model addresses the limitations of existing chatbots by generating coherent and dynamic conversations through a directed event graph and summary module, ensuring consistency and diversity across multiple sessions.

WO2025150613A1PCT designated stage expired Publication Date: 2025-07-17UNIST (ULSAN NAT INST OF SCI & TECH)
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/004149
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-10
Filing Date
2024-04-01
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing chatbot technologies struggle to provide rich and realistic conversations across multiple sessions due to limitations in handling long-term interactions, temporal inconsistencies, and dynamic speaker relationships, especially in open-domain scenarios.

Method used

A multi-session conversation model that generates a conversation chronicle with consistent but diverse conversations by constructing a directed event graph, applying various time intervals, and considering speaker relationships, using a summary and response generation module to ensure coherent and dynamic interactions.

Benefits of technology

The model efficiently stores and summarizes previous sessions, minimizing information loss while generating responses that are consistent with previous sessions and reflect dynamic interactions, thus providing realistic and diverse conversations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024004149_17072025_PF_FP_ABST
    Figure KR2024004149_17072025_PF_FP_ABST
Patent Text Reader

Abstract

In an electronic device and a method for providing a chatbot service by using a multi-session conversation model according to an embodiment, the electronic device may comprise: a memory that stores the multi-session conversation model which is trained with conversation chronicles in order to respond to an input query and comprises a summarization module and a response generation module; and a processor connected to the memory to generate conversation summaries of respective sessions in chronological order by using the summarization module on the basis of one or more previous sessions and the current session, and generate a response in the current session by using the response generation module on the basis of the generated conversation summaries of the one or more previous sessions, relationships between conversation speakers, time intervals between sessions, and the last query of the current session.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method for providing chatbot service using a multi-session conversation model based on chronological dynamics

[0001] Below, a technology for a multi-session conversation environment implementing chronological dynamics is provided.

[0002] A chatbot is a chat robot that understands human language, such as voice and text, and can carry on a conversation. When a user asks a chatbot a question, it understands the intent of the question and provides an appropriate response. If the chatbot lacks sufficient information to provide a response, it proactively repeats the question-and-answer process with the user to supplement the missing information. This allows users to quickly and easily access the information they need, even without prior knowledge of the subject matter, by following the chatbot's guidance. Due to these characteristics, chatbots are being utilized in diverse fields such as technical support, marketing, online shopping, and law.

[0003] The background technology described above is technology that the inventor possessed or acquired in the process of deriving the disclosure of the present application, and cannot necessarily be said to be publicly known technology disclosed to the general public prior to the present application.

[0004] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment can provide a chatbot service based on multiple sessions using a multi-session conversation model.

[0005] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment can generate a conversation chronicle including consistent but diverse conversations according to various time intervals and speaker relationships, and train a multi-session conversation model using the conversation chronicle.

[0006] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment collects events by constructing a directional event graph to create a conversation chronicle, applies various time intervals from hours to years to the collected events, and can consider responses that may vary over time and relationships between speakers.

[0007] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment may include a summary module.

[0008] However, technical challenges are not limited to the technical challenges described above, and other technical challenges may exist.

[0009] An electronic device providing a chatbot service using a multi-session conversation model according to one embodiment may include a memory in which the multi-session conversation model is stored, which is trained with conversation chronicles to respond to an input query and includes a summarization module and a response generation module, and a processor connected to the memory to generate a conversation summary for each session in chronological order using the summarization module based on one or more previous sessions and a current session, and to generate a response in the current session using the response generation module based on the conversation summary of one or more previous sessions generated, a relationship between conversation speakers, a time interval between sessions, and the last query of the current session.

[0010] The above processor can generate the conversation chronicle composed of multiple episodes consisting of multiple sessions by assigning relationships between speakers within each episode and time intervals between each event to multiple events per episode.

[0011] The above processor can form an event list including a plurality of events, select event pairs based on their interrelationships from the event list, and construct a directed event graph in which the order of premises and hypotheses of the selected event pairs is specified, thereby filtering the final events used in the conversation chronicle.

[0012] The above processor can calculate relationships between events in the event list, and select an event pair that is in an entailment relationship among the entailment, neutral, and contradiction relationships between the events.

[0013] The processor can connect each event of the selected event pair as a node based on the order of premises and hypotheses, extract a plurality of event sequences of a predetermined length, and select one event sequence from among the plurality of extracted event sequences.

[0014] The above processor can assign a time interval of any one of hour, day, week, month, and year to each event of the selected event sequence.

[0015] The processor may select one of Classmates, Neighbors, Co-workers, Mentee and Mentor, Husband and Wife, Patient and Doctor, Parent and Child, Student and Teacher, Employee and Boss, and Athlete and Coach as the relationship between the speakers based on the events of the selected one event sequence.

[0016] The above processor can retrieve social common sense from common sense data including any two elements and relationships between the elements, and form an event in the form of a sentence for one session based on the retrieved social common sense and include it in an event list.

[0017] The processor may delete an episode that includes at least one of a session in which two or more speakers participated, a session in which there is inconsistency or ambiguity between utterances and speakers, a session in which speakers other than those in a predetermined relationship appear, and a session with unnecessary information among the generated conversation chronicles.

[0018] The processor trains the response generation module by randomly noise-ing the text of the conversation chronicle to reconstruct the original text, and the trained response generation module can calculate a conditional probability to generate a response in the current session.

[0019] A method for providing a chatbot service using a multi-session conversation model according to one embodiment may include a step of generating a conversation summary for each session in chronological order based on one or more previous sessions and a current session by a summary module among the multi-session conversation models, and a step of generating a response in a current session based on a conversation summary of one or more previous sessions generated by a response generation module among the multi-session conversation models, a relationship between conversation speakers, a time interval between sessions, and a last query of the current session.

[0020] The method may further include, before generating the conversation summary, a step of generating a conversation chronicle comprising a plurality of episodes composed of a plurality of sessions by assigning relationships between speakers in the episode and time intervals between each event to a plurality of events per episode, and a step of deleting an episode including at least one of a session in which two or more speakers participated, a session in which there is inconsistency or ambiguity between utterances and speakers, a session in which a speaker other than a predetermined relationship between speakers appears, and a session with unnecessary information, among the generated conversation chronicles.

[0021] The step of generating the above conversation chronicle may include the steps of forming an event list including a plurality of events, selecting event pairs based on their interrelationships from the event list, and filtering the final events used in the conversation chronicle by constructing a directed event graph in which the order of premises and hypotheses of the selected event pairs is specified.

[0022] The step of selecting the above event pair may include a step of calculating relationships between events in the event list and a step of selecting an event pair that is in a relationship among the relationships between the events, which are entailment, neutrality, and contradiction relationships.

[0023] The step of filtering the above final event may include the step of connecting each event of the selected event pair as a node based on the order of the premise and hypothesis, the step of extracting a plurality of event sequences of a predetermined length, and the step of selecting one event sequence from the plurality of extracted event sequences.

[0024] The step of generating the above conversation chronicle may further include the step of assigning a time interval of any one of hours, days, weeks, months, and years to each event of the selected event sequence, and the step of selecting one of the following as the relationship between the speakers: classmate, neighbor, colleague, mentee and mentor, husband and wife, patient and doctor, parent and child, student and teacher, subordinate and superior, and athlete and coach, based on the events of the selected event sequence.

[0025] The step of forming the above event list may include a step of searching for social common sense from common sense data including any two elements and relationships between the elements, and a step of forming an event in the form of a sentence for one session based on the searched social common sense and including the event in the event list.

[0026] The method may further include a step of training the response generation module by randomly noise-ing the text of the conversation chronicle to reconstruct the original text before generating the conversation summary.

[0027] The step of generating the above response may include a step of the trained response generating module calculating a conditional probability to generate a response in the current session.

[0028] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment can provide a rich and realistic conversation by providing a chatbot service based on multiple sessions using a multi-session conversation model.

[0029] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment generates a conversation chronicle including consistent but diverse conversations according to various time intervals and speaker relationships, and trains the multi-session conversation model using the conversation chronicle, thereby generating responses that are consistent with previous sessions.

[0030] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment construct a directional event graph to collect events to create a conversation chronicle, apply various time intervals from hours to years to the collected events, and consider responses that may vary over time and relationships between speakers, thereby preventing temporal contradictions, allowing diversity, and reflecting dynamic interactions, thereby generating a conversation similar to an actual scenario.

[0031] An electronic device and method for providing a chatbot service using a multi-session conversation model according to one embodiment can efficiently store the entire conversation of a previous session while minimizing information loss by including a summary module in the multi-session conversation model.

[0032] FIG. 1 illustrates a block diagram of an electronic device that provides a chatbot service using a multi-session conversation model according to one embodiment.

[0033] FIG. 2 illustrates an example of generating a response from an electronic device providing a chatbot service using a multi-session conversation model according to one embodiment.

[0034] FIG. 3 illustrates an example of generating a conversation chronicle of an electronic device providing a chatbot service using a multi-session conversation model according to one embodiment.

[0035] FIG. 4 illustrates a flowchart of a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0036] FIG. 5 illustrates a graph of performance evaluation results between a method of providing a chatbot service using a multi-session conversation model according to one embodiment and another method.

[0037] FIG. 6 illustrates an example of any one of the conversation chronicles generated from a method of providing a chatbot service using a multi-session conversation model according to one embodiment.

[0038] FIG. 7 illustrates an example of a method for providing a chatbot service using a multi-session conversation model according to one embodiment, which illustrates the relationship between speakers in a corresponding scenario generated using a summary module, the time interval between the previous session and the current session, and the summary of each session.

[0039] FIG. 8 illustrates an example of a conversation having various flows depending on the relationship between speakers generated using a response generation module in a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0040] FIG. 9 illustrates an example of a conversation reflecting a time interval generated using a response generation module in a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0041] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Therefore, the actual implementation is not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or alternatives within the technical concepts described in the embodiments.

[0042] Although terms such as "first" or "second" may be used to describe various components, these terms should be interpreted solely to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.

[0043] When it is said that a component is "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.

[0044] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprises" or "has" should be understood to indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but not to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.

[0045] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.

[0046] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0047] Natural Language Processing (NLP) can refer to the analysis, understanding, or generation of natural language, the language humans use to communicate. Furthermore, NLP is a major development area in artificial intelligence and can be performed using technologies such as machine learning (ML) and / or deep learning. NLP is the foundation technology for chatbots, a type of conversational AI. Chatbots respond to user queries on behalf of users, and are therefore being applied in diverse fields such as finance, healthcare, and education, leading to rapid technological advancements.

[0048] Chatbots can be categorized into rule-based chatbots, machine learning chatbots, hybrid chatbots, custom chatbots, virtual assistants, domain-specific chatbots, and social media chatbots, depending on their intended use and / or purpose. Rule-based chatbots operate based on predefined rules and patterns, making them suitable for simple tasks but less suitable for more complex conversations. Domain-specific chatbots, also known as closed-domain chatbots, are a type of chatbot that improves on these shortcomings. Their models are trained with conversational data based on specialized knowledge in a specific industry or domain. This allows them to identify patterns in user queries and generate responses based on these patterns.

[0049] Open-domain chatbots, the opposite concept, are an improved form of closed-domain chatbots. They are not limited to a specific domain and can communicate with users as if they were having a human conversation. Open-domain chatbots offer the advantage of not only providing information through natural conversations but also offering a wide range of other services, and related technologies are actively being researched. For example, open-domain chatbots can be applied to various emotional services, such as counseling or language tutoring, to promote emotional support and mental health. However, while open-domain chatbots can generate fluent, human-like responses, they generate responses based on short-term conversation sessions (e.g., single sessions) that only include the current conversation. This may limit their practical application to emotional services that require responses based on longer-term conversations.

[0050] Accordingly, an electronic device (100) that provides a chatbot service using a multi-session conversation model according to one embodiment can provide a rich and realistic conversation by providing a chatbot service based on multiple sessions using the multi-session conversation model.

[0051] FIG. 1 illustrates a block diagram of an electronic device that provides a chatbot service using a multi-session conversation model according to one embodiment.

[0052] The electronic device (100) may include a memory (110) and a processor (120). For example, the electronic device (100) may be, but is not limited to, a server, a smartphone, an AI speaker, etc.

[0053] The memory (110) may be trained with conversation chronicles to respond to input queries, and may store a multi-session conversation model including a summarization module and a response generation module. A chatbot to which the multi-session conversation model is applied is a chatting robot, and may include an agent for providing at least one service through an electronic device (100) based on communication with a user, or the entire data constituting the agent.

[0054] The memory (110) can store instructions (e.g., programs) executable by the processor (120). For example, the instructions may include instructions for executing operations of the processor (120) and / or operations of each component of the processor (120).

[0055] According to various embodiments, the memory (110) may be implemented as a volatile memory device or a nonvolatile memory device. The volatile memory device may be implemented as a dynamic random access memory (DRAM), a static random access memory (SRAM), a thyristor RAM (T-RAM), a zero capacitor RAM (Z-RAM), or a twin transistor RAM (TTRAM). The nonvolatile memory device may be implemented as an Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory, Magnetic RAM (MRAM), Spin-Transfer Torque (STT)-MRAM, Conductive Bridging RAM (CBRAM), Ferroelectric RAM (FeRAM), Phase change RAM (PRAM), Resistive RAM (RRAM), Nanotube RRAM, Polymer RAM (PoRAM), Nano Floating Gate Memory (NFGM), holographic memory, Molecular Electronic Memory Device, and / or Insulator Resistance Change Memory.

[0056] FIG. 2 illustrates an example of generating a response from an electronic device providing a chatbot service using a multi-session conversation model according to one embodiment.

[0057] The processor (120) is connected to the memory (110) and generates a conversation summary for each session in chronological order based on one or more previous sessions and the current session by the summary module (210), and generates a response (221) in the current session (e.g., session N) by the response generation module (220) based on the conversation summary (213) of one or more previous sessions (e.g., session 1 to session N-1), the relationship (211) between conversation speakers, the time interval (212) between sessions, and the last query (e.g., current conversation context) of the current session (e.g., session N). When generating the response, the processor (120) must store and utilize the entire conversation history of one or more previous sessions as a context in the memory (110) because it is necessary to consider the chronological connection between one or more previous sessions and the current session. However, simply storing and utilizing the entire conversation history of the previous session has problems in securing storage capacity and inefficient data processing. The aforementioned problem can be solved by utilizing a summary module (210) that briefly summarizes the entire conversation history of one or more previous sessions while minimizing information loss. Accordingly, the processor (120) can utilize the trained summary module (210) to summarize the entire conversation of the current session, including the previous session, into chronological events (213).

[0058] Before using the summary module (210) and the response generation module (220), the processor (120) can train the summary module (210) and the response generation module (220). The processor (120) can train the summary module (210) to specify conditions and output new text summarizing the contents of the original text based on the conditions. The processor (120) can train the response generation module (220) by randomly noise-inducing the text of the conversation chronicle to reconstruct the original text. Thereafter, the trained response generation module (220) can calculate a conditional probability to generate a response in the current session. The conditional probability can be represented as P(c|r,t,s,h), where c can represent the next utterance (e.g., response) of the chatbot. r can represent the relationship between speakers (e.g., student and teacher, patient and doctor, etc.). t represents the time interval between each event. In situations where more than one session has been conducted, the chatbot must consider the time interval between the previous session and the ongoing session, so t can be considered as an input for generating a response. s represents a summary of the conversation from the previous session generated by the summary module, and h can represent the current conversation context (e.g., all conversational utterances in progress in the current session).

[0059] First, the processor (120) can generate a conversation chronicle to train a multi-session conversation model. The conversation chronicle can contain consistent yet diverse conversations across various time intervals (e.g., multiple sessions) and speaker relationships. Therefore, the processor (120) can train the multi-session conversation model using the conversation chronicle, thereby generating responses consistent with previous sessions. If the processor (120) trains the multi-session conversation model using general chatbot data (e.g., data with a single session), the single-session data can cause the multi-session conversation model to engage in conversations centered on specific events, ignoring the past context between speakers. This may mean that the multi-session conversation model cannot appropriately generate responses in the current session following the previous session when multiple sessions are chronologically connected in a single episode related to a single topic. In addition, even if a multi-session conversation model is trained using data with only multi-sessions, it may only use data with multi-sessions with short time intervals, so responses that may vary depending on the time elapsed from the previous session are not taken into account, which may hinder the dynamism of conversational interactions and may have the disadvantage of not being similar to real-world scenarios.

[0060] FIG. 3 illustrates an example of generating a conversation chronicle of an electronic device providing a chatbot service using a multi-session conversation model according to one embodiment.

[0061] The processor (120) can generate a conversation chronicle (340) that comprises a plurality of episodes composed of a plurality of sessions by assigning a relationship (331) between speakers within the episode and a time interval (332) between each event to a plurality of events (333) per episode. The processor (120) can generate the conversation chronicle (340) using a conversation chronicle generation model (e.g., ChatGPT).

[0062] Prior to generating a conversation chronicle (340), the processor (120) may first form an event list (310) including multiple events. The processor (120) may retrieve social common sense from common sense data including any two elements and relationships between the elements, and may form a sentence-type event for a single session based on the retrieved social common sense and include the event list.

[0063] Common sense data can be, for example, a Commonsense Knowledge Graph, and the Commonsense Knowledge Graph can be expressed as any two elements and the relationship between the elements. For example, the relationship can be divided into causal xAttr, xReact, xEffect, xIntent, xWant, and xNeed, etc., and the first element X can mean perception, reaction, effect, intention, want, and need, respectively. The Commonsense Knowledge Graph includes social (e.g., intention, desire, reaction) commonsense and event-based (e.g., event sequence) commonsense, but since the processor (120) needs to filter social interactions, it can retrieve the social commonsense. The retrieved social commonsense can be a Commonsense Knowledge Graph based on social commonsense. Accordingly, the processor (120) can form a sentence based on the two elements and the relationship between the elements. Then, the processor (120) can complete an event by adding a short example to the formed sentence.

[0064] The processor (120) can select event pairs based on their interrelationships from the event list (310) and construct a directed event graph (320) in which the order of the premises and hypotheses of the selected event pairs is specified, thereby filtering the final events used in the conversation chronicle. The processor (120) can utilize a Natural Language Inference (NLI) method to connect two related events. Accordingly, the processor (120) can calculate the relationships between events in the event list and select an event pair that is in an entailment relationship (e.g., true) among the entailment, neutral, and contradiction relationships between events. The processor (120) can connect each event of the selected event pair as a node based on the order of the premises and hypotheses. Through this, a directed event graph (320) can be constructed, and temporal contradictions can be prevented due to the directionality. The processor (120) can extract a plurality of event sequences of a predetermined length (e.g., five events connected) and select one event sequence from among the extracted plurality of event sequences. In the selection method, if there is an overlapping path of a predetermined length (e.g., three events connected) among the plurality of event sequences, the processor (120) can randomly select only one event sequence with overlapping paths. For example, if there are a plurality of event sequences having event pairs of the 5-6-7-9-10 path, the 4-8-5-9-10 path, the 1-2-3-4-5 path, and the 1-2-3-8-5 path, the path 1-2-3 can overlap in the 1-2-3-4-5 path and the 1-2-3-8-5 path. The processor (120) can randomly select the 1-2-3-4-5 path from among the overlapping paths.

[0065] The processor (120) can divide the event graph form of the selected event sequence into events (333) according to the flow of time and classify them into one episode (330). One episode (330) can be assigned a relationship (331) between speakers and a time interval (332) between each event (333). The processor (120) can roughly randomly assign a time interval between the last conversation of the previous session of one session and the first conversation of one session. For example, the processor (120) can assign a time interval (332) of any one of hour, day, week, month, and year to each of the events (333) of the selected event sequence. This can mean that the processor (120) randomly selects and assigns one of a total of five time intervals. The time interval (332) may mean an approximate unit of time (e.g., 3 to 6 days) rather than a numerical time (e.g., 3 days).

[0066] The processor (120) may predefine a plurality of relationship lists so that they can be assigned. The processor (120) may generate a relationship (331) between speakers using a relationship determination model (e.g., ChatGPT). Based on events (333) of a selected event sequence, the processor (120) may select one of Classmates, Neighbors, Co-workers, Mentee and Mentor, Husband and Wife, Patient and Doctor, Parent and Child, Student and Teacher, Employee and Boss, and Athlete and Coach as the relationship (331) between speakers. Accordingly, the processor (120) applies various time intervals (332) from hours to years and considers the relationships between responses and speakers that may change over time, thereby preventing temporal contradictions, allowing diversity, and reflecting dynamic interactions, thereby generating a conversation similar to an actual scenario.

[0067] Additionally, the processor (120) may delete episodes containing at least one of the following: sessions involving two or more speakers, sessions with inconsistencies or ambiguities between utterances and speakers, sessions featuring speakers other than those with predetermined speaker relationships, and sessions containing unnecessary information (e.g., information about actions or situations). This ensures consistent quality of the conversation chronicles.

[0068] According to various embodiments, the processor (120) may execute computer-readable code (e.g., software) stored in the memory (110) and instructions generated by the processor (120). The processor (120) may be a hardware-implemented data processing device having a circuit having a physical structure for executing desired operations. The desired operations may include, for example, code or instructions included in a program. The hardware-implemented data processing device may include, for example, a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an application-specific integrated circuit (ASIC), and a field programmable gate array (FPGA).

[0069] FIG. 4 illustrates a flowchart of a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0070] In method (410), the processor may generate a conversation summary for each session in chronological order based on one or more previous sessions and the current session using a summary module among the multi-session conversation models. Before generating the conversation summary, the processor may train a response generation module by randomly denoising the text of the conversation chronicle to reconstruct the original text.

[0071] In method (420), the processor may generate a response in the current session based on a conversation summary of one or more previous sessions generated by a response generation module among multi-session conversation models, relationships between conversation speakers, time intervals between sessions, and the last query of the current session. In generating the response, the processor may generate the response in the current session by having the trained response generation module calculate a conditional probability.

[0072] Before generating a conversation summary in method (410), the processor may generate a conversation chronicle comprising multiple episodes, each consisting of multiple sessions, by assigning relationships between speakers within each episode and time intervals between each event to multiple events within each episode. Furthermore, the processor may delete, from the generated conversation chronicle, episodes that include at least one of the following: sessions involving two or more speakers, sessions with inconsistencies or ambiguities between utterances and speakers, sessions involving speakers other than those in a predetermined relationship between speakers, and sessions containing unnecessary information.

[0073] In a specific method of generating a conversation chronicle, the processor may first form an event list including multiple events. When forming the event list, the processor may first retrieve social common sense from common sense data including any two elements and their relationships, and, based on the retrieved social common sense, form a sentence-type event for a session and include it in the event list. Then, the processor may select event pairs from the event list based on their interrelationships and construct a directed event graph specifying the order of the premises and hypotheses of the selected event pairs to filter the final events used in the conversation chronicle. When selecting event pairs, the processor may calculate relationships between events in the event list and select event pairs that are entailed among the calculated relationships between events, such as entailment, neutrality, and contradiction. When filtering the final events, the processor may connect each event in the selected event pair as a node based on the order of the premises and hypotheses, extract multiple event sequences of a predetermined length, and select one event sequence from among the extracted multiple event sequences.

[0074] Thereafter, the processor assigns a time interval of one of hours, days, weeks, months, and years to each event of the selected event sequence, and selects one of classmates, neighbors, colleagues, mentees and mentors, husbands and wives, patients and doctors, parents and children, students and teachers, subordinates and superiors, and athletes and coaches as the relationship between the speakers based on the events of the selected event sequence, thereby generating a conversation chronicle.

[0075] FIG. 5 illustrates a graph of performance evaluation results between a method of providing a chatbot service using a multi-session conversation model according to one embodiment and another method.

[0076] Performance evaluations between the multi-session dialogue model (520) and other models (510) are performed through human evaluations and can be based on consistency, coherence, time intervals, and relationships between speakers. Consistency refers to the fact that two speakers do not make contradictory statements from previous sessions, and coherence refers to the fact that the conversation between two speakers has a natural flow in terms of event transitions. The graph shows the performance evaluation results based on consistency, coherence, time intervals, and the overall average value. In Figure 5, the performance of the multi-session dialogue model (520) can be seen to be superior to that of the other model (510).

[0077] FIG. 6 illustrates an example of any one of the conversation chronicles generated from a method of providing a chatbot service using a multi-session conversation model according to one embodiment.

[0078] An episode in a conversation chronicle is a conversation where the speakers are colleagues and the time interval is measured in years. In session N-1, the conversation might include the phrases "After working all day in the heat" and "Having a beer." Years later, the conversation in session N might consider the conversation in session N-1 and include the phrase "Remember when you had a relaxing moment with a few beers after working all day in the sun?" A multi-session conversation model trained using these conversation chronicles can also generate responses reflecting previous sessions.

[0079] FIG. 7 illustrates an example of a method for providing a chatbot service using a multi-session conversation model according to one embodiment, which illustrates the relationship between speakers in a corresponding scenario generated using a summary module, the time interval between the previous session and the current session, and the summary of each session.

[0080] The processor can use the summary module to generate relationships between speakers in a given scenario, the time interval between the previous and current sessions, and summaries (e.g., events) of each session. The processor can use the summary module to generate relationships between the speakers, who are athletes and coaches, and the time interval in hours for the episode. The processor can use the summary module to generate a summary of a previous session, in which the coach and athlete are scheduled to begin training in the morning, and a summary of a current session, in which the coach and athlete decide to change their training schedule due to their class schedule. The summary module can effectively capture these state changes between the previous and current sessions, allowing it to detail the changes in the training schedule from morning to afternoon.

[0081] FIG. 8 illustrates an example of a conversation having various flows depending on the relationship between speakers generated using a response generation module in a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0082] A processor (e.g., a chatbot) can have different conversational flows based on the relationships between speakers on the same conversation topic, and can use a response generation module to generate responses based on the relationships between speakers. For example, if a user asks, "I feel like my salary is too low compared to my workload," the response may vary depending on the relationship. If the chatbot and the user are husband and wife, the chatbot may provide a response that suggests a solution, such as talking to their boss. Alternatively, if the chatbot and the user are subordinate and boss, the chatbot may provide a response that emotionally accepts the user's opinion about the company.

[0083] FIG. 9 illustrates an example of a conversation reflecting a time interval generated using a response generation module in a method for providing a chatbot service using a multi-session conversation model according to one embodiment.

[0084] A processor (e.g., a chatbot) can use a response generation module to generate responses based on relationships and time intervals. In this example, the processor can generate responses based on past sessions (the first and second conversations), even if they are not the immediate session. Furthermore, the chatbot can accurately reflect the cumulative effect of past time in its responses. For example, the chatbot can include an event related to going to the beach together several years ago in the response for the current session. Furthermore, the chatbot can provide a vacation-related event from a conversation a few weeks after going to the beach as a response related to the vacation from several years ago in the current session.

[0085] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, those skilled in the art will appreciate that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, the processing device may include multiple processors, or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0086] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.

[0087] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0088] The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.

[0089] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0090] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In an electronic device that provides a chatbot service using a multi-session conversation model, A memory storing the multi-session conversation model, which is trained with CONVERSATION CHRONICLES to respond to input queries and includes a summarization module and a response generation module; and A processor connected to the memory, generating a conversation summary for each session in chronological order by the summary module based on one or more previous sessions and the current session, and generating a response in the current session by the response generation module based on the conversation summary of one or more previous sessions generated, the relationship between conversation speakers, the time interval between sessions, and the last query of the current session. An electronic device comprising:

2. In paragraph 1, The above processor, Generating the conversation chronicle composed of multiple episodes consisting of multiple sessions by assigning the relationships between speakers in each episode and the time intervals between each event to multiple events in each episode. Electronic devices.

3. In paragraph 2, The above processor, Forming an event list including multiple events, selecting event pairs based on their mutual relationships from the event list, and constructing a directed event graph in which the order of premises and hypotheses of the selected event pairs is specified, thereby filtering the final events used in the conversation chronicle. Electronic devices.

4. In paragraph 3, The above processor, Compute the relationships between events in the above event list, and select a pair of events that are in an entailment relationship among the relationships between the computed events, which are entailment, neutral, and contradiction relationships. Electronic devices.

5. In paragraph 3, The above processor, Connecting each event of the above-mentioned selected event pair as a node based on the order of the premise and hypothesis, extracting multiple event sequences of a predetermined length, and selecting one event sequence from among the multiple extracted event sequences. Electronic devices.

6. In paragraph 5, The above processor, Assigning a time interval of one of hour, day, week, month, and year to each event of the above-mentioned selected event sequence. Electronic devices.

7. In paragraph 5, The above processor, Based on the events of the above-mentioned selected event sequence, one of the following is selected as the relationship between the speakers: Classmates, Neighbors, Co-workers, Mentee and Mentor, Husband and Wife, Patient and Doctor, Parent and Child, Student and Teacher, Employee and Boss, and Athlete and Coach. Electronic devices.

8. In paragraph 2, The above processor, Retrieving social common sense from common sense data including any two elements and the relationship between the elements, and forming a sentence-type event for one session based on the retrieved social common sense and including it in the event list. Electronic devices.

9. In paragraph 2, The above processor, Among the above generated conversation chronicles, an episode that includes at least one of the following: a session in which two or more speakers participated, a session in which there is inconsistency or ambiguity between utterances and speakers, a session in which speakers other than those in a predetermined relationship appear, and a session with unnecessary information is deleted. Electronic devices.

10. In paragraph 1, The above processor, The response generation module is trained by randomly noise-ing the text of the above conversation chronicle to reconstruct the original text, The trained response generation module calculates the conditional probability to generate a response in the current session. Electronic devices.

11. A method for providing a chatbot service using a multi-session conversation model, A step of generating a conversation summary for each session in chronological order based on one or more previous sessions and the current session as a summary module in a multi-session conversation model; and A step for generating a response in the current session based on the conversation summary of one or more previous sessions generated by the response generation module in the multi-session conversation model, the relationship between conversation speakers, the time interval between sessions, and the last query of the current session. A method including:

12. In paragraph 11, Before generating the above conversation summary, a step of generating a conversation chronicle consisting of multiple episodes composed of multiple sessions by assigning relationships between speakers in the episode and time intervals between each event to multiple events per episode; and Step of deleting an episode that includes at least one of the following: a session in which two or more speakers participated, a session in which there is inconsistency or ambiguity between utterances and speakers, a session in which speakers other than those in a predetermined relationship appear, and a session with unnecessary information among the generated conversation chronicles. How to include more.

13. In paragraph 12, The steps to create the above conversation chronicle are: A step of forming an event list including multiple events; A step of selecting a pair of events based on a mutual relationship from the above event list; and A step of filtering the final events used in the conversation chronicle by constructing a directed event graph in which the order of the premises and hypotheses of the above-mentioned selected event pairs is specified. A method including:

14. In paragraph 13, The step of selecting the above event pair is: A step of calculating relationships between events in the above event list; and A step for selecting a pair of events that are in an accompanying relationship among the relationships among the calculated events, which are entailment, neutrality, and contradiction relationships. A method including:

15. In paragraph 13, The step of filtering the above final event is: A step of connecting each event of the above-mentioned selected event pair as a node based on the order of the premise and hypothesis; a step of extracting a plurality of event sequences of a predetermined length; and A step of selecting one event sequence from among the above extracted multiple event sequences. A method including:

16. In paragraph 13, The steps to create the above conversation chronicle are: A step of assigning a time interval of any one of hours, days, weeks, months and years to each event of a selected event sequence; and A step of selecting one of the following as the relationship between the speakers based on the events of the above-mentioned selected event sequence: classmate, neighbor, colleague, mentee and mentor, husband and wife, patient and doctor, parent and child, student and teacher, subordinate and boss, and athlete and coach. How to include more.

17. In paragraph 13, The steps for forming the above event list are: A step of retrieving social common sense from common sense data including any two elements and the relationship between the elements; and A step of forming a sentence-type event for one session based on the searched social common sense and including it in the event list. A method including:

18. In paragraph 11, Before generating the above conversation summary, a step of training the response generation module by randomly noise-ing the text of the conversation chronicle to reconstruct the original text. How to include more.

19. In paragraph 11, The steps for generating the above response are: A step in which the trained response generation module calculates conditional probability to generate a response in the current session. A method including:

Citation Information

Patent Citations

  • Discriminating ambiguous expressions to enhance user experience

    KR1020170099917A

  • Backside power rail to deep vias

    KR1020230034902A

  • Apparatus for supplying droplet

    KR102277982B1

  • Database systems and methods for conversation-driven dynamic updates

    US20210247957A1