Electronic device and method for providing dialogue service, and computer-readable recording medium

By automatically recognizing users' facial IDs and storing conversation history in electronic devices, the problem of not being able to provide personalized answers in existing technologies is solved, thus achieving continuity and accuracy in conversation services.

CN120910221APending Publication Date: 2025-11-07SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511361895.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-08-26
Filing Date
2020-03-02
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies in devices that provide dialogue services to multiple users cannot provide accurate personalized responses when users utter words related to their past conversation history.

Method used

By automatically recognizing a user's facial ID in electronic devices, storing and managing their conversation history, and using this history to generate response messages, the continuity and personalization of conversations can be achieved.

Benefits of technology

It improves the accuracy and continuity of conversational services without user identification, reduces the inconvenience of user account registration, and provides personalized responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910221A_ABST
    Figure CN120910221A_ABST
Patent Text Reader

Abstract

The application provides an electronic device, a method of providing a dialogue service performed by the electronic device, and a computer-readable recording medium. The method comprises the following steps: receiving utterance input; identifying a time expression representing a time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from a plurality of databases storing dialogue history information of the user using the dialogue service; interpreting the text based on dialogue history information of the user acquired from the selected database; generating a response message to the utterance input based on a result of the interpretation; and outputting the generated response message.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The disclosure relates to an electronic device and a method for providing a conversation service, and for example, to a method and device for interpreting a user utterance input and outputting a response message based on a conversation history of a user. BACKGROUND

[0002] With recent development of electronic devices, such as smartphones, for performing various functions in a complex manner, electronic devices equipped with a voice recognition function have been launched to improve operability. The voice recognition technology can be applied to a conversational user interface for outputting a response message to a question input by a user's voice in a daily natural language to provide a user-friendly conversation service.

[0003] A conversational user interface refers to a smart user interface that operates while talking in a user's language. The conversational user interface can be used in a question answering (QA) system for outputting an answer to a user's question. In contrast to an information retrieval technology for simply retrieving and presenting list information related to a user's question, the difference of the QA system is that the QA system searches for and provides an answer to a user's question.

[0004] For example, personal electronic devices such as smartphones, computers, personal digital assistants (PDAs), portable multimedia players (PMPs), smart home appliances, navigation devices, wearable devices, etc. can provide a conversation service by connecting to a server or executing an application.

[0005] As another example, public electronic devices such as unmanned guidance information terminals, unmanned kiosks, unmanned checkout counters, etc. installed in a store or a public institution can also provide a conversation service. Public electronic devices installed in public places need to store and use a conversation history for each user in order to accurately analyze a user utterance input and provide a personalized answer thereto. SUMMARY

[0006] TECHNICAL PROBLEM

[0007] When using a device for providing a conversation service to a plurality of users, there is a need for a method capable of receiving an accurate personalized answer from the device even when a user utters a speech related to a past conversation history.

[0008] TECHNICAL SOLUTION

[0009] Embodiments of the disclosure provide a method and device for providing a conversation service by performing a process of retrieving a stored conversation history associated with a user account in a more user-friendly manner.

[0010] Additional aspects will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art.

[0011] A method of providing a conversation service, performed by an electronic device according to an example embodiment of the disclosure, includes receiving an utterance input, identifying a time expression representing a time in text obtained from the utterance input, determining a time point related to the utterance input based on the time expression, selecting a database corresponding to the determined time point from among a plurality of databases, wherein the plurality of databases store conversation history information of a user using the conversation service, interpreting the text based on the conversation history information of the user obtained from the selected database, generating a response message to the utterance input based on a result of the interpreting, and outputting the generated response message. BRIEF DESCRIPTION OF DRAWINGS

[0012] The above and other aspects, features, and advantages of certain embodiments of the disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0013] Figure 1 FIG. 1 is a diagram illustrating an example in which an electronic device provides a conversation service based on a conversation history according to an embodiment of the disclosure;

[0014] Figure 2a FIG. 2 is a diagram illustrating an example of an exemplary system for providing a conversation service according to an embodiment of the disclosure;

[0015] Figure 2b FIG. 3 is a diagram illustrating an example of an exemplary system for providing a conversation service according to an embodiment of the disclosure;

[0016] Figure 3 FIG. 4 is a flowchart illustrating an exemplary method of providing a conversation service, performed by an electronic device according to an embodiment of the disclosure;

[0017] Figure 4a FIG. 5 is a diagram illustrating an example in which an electronic device provides a conversation service based on a conversation history according to an embodiment of the disclosure;

[0018] Figure 4b FIG. 6 is a diagram illustrating an example in which an electronic device provides a conversation service based on a conversation history according to an embodiment of the disclosure;

[0019] Figure 5 FIG. 7 is a diagram illustrating an exemplary process of providing a conversation service, performed by an electronic device according to an embodiment of the disclosure;

[0020] Figure 6 FIG. 8 is a diagram illustrating an example of stored conversation history information according to an embodiment of the disclosure;

[0021] Figure 7a FIG. 9 is a flowchart illustrating an exemplary method of providing a conversation service, performed by an electronic device according to an embodiment of the disclosure;

[0022] Figure 7bis a flowchart illustrating an exemplary method of providing a conversation service, performed by an electronic device, according to an embodiment of the disclosure;

[0023] Figure 8 is a flowchart illustrating an exemplary method of determining whether an electronic device is to generate a response using conversation history information, according to an embodiment of the disclosure;

[0024] Figure 9 is a flowchart illustrating an exemplary method of selecting a database based on a user utterance input, performed by an electronic device, according to an embodiment of the disclosure;

[0025] Figure 10 is an exemplary probability graph of determining a time point related to a user utterance input, by an electronic device, according to an embodiment of the disclosure;

[0026] Figure 11 is a diagram illustrating an exemplary method of switching a database in which a user's conversation history is stored, performed by an electronic device, according to an embodiment of the disclosure;

[0027] Figure 12a is a diagram illustrating an exemplary process of a plurality of electronic devices sharing a user's conversation history with each other, according to an embodiment of the disclosure;

[0028] Figure 12b is a diagram illustrating an exemplary process of a plurality of electronic devices sharing a user's conversation history with each other, according to an embodiment of the disclosure;

[0029] Figure 12c is a diagram illustrating an exemplary process of a plurality of electronic devices sharing a user's conversation history with each other, according to an embodiment of the disclosure;

[0030] Figure 12d is a diagram illustrating an exemplary process of a plurality of electronic devices sharing a user's conversation history with each other, according to an embodiment of the disclosure;

[0031] Figure 13a is a block diagram illustrating an exemplary configuration of an exemplary electronic device, according to an embodiment of the disclosure;

[0032] Figure 13b is a block diagram illustrating an exemplary configuration of an exemplary electronic device, according to another embodiment of the disclosure;

[0033] Figure 14 is a block diagram of an exemplary electronic device, according to an embodiment of the disclosure;

[0034] Figure 15a is a block diagram of an exemplary processor included in an exemplary electronic device, according to an embodiment of the disclosure;

[0035] Figure 15b is a block diagram illustrating an exemplary processor included in an exemplary electronic device according to an embodiment of the disclosure;

[0036] Figure 16 is a block diagram illustrating an exemplary speech recognition module according to an embodiment of the disclosure;

[0037] Figure 17 is a diagram illustrating an exemplary temporal expression extraction model according to an embodiment of the disclosure; and

[0038] Figure 18 is a diagram illustrating an exemplary time point prediction model according to an embodiment of the disclosure. DETAILED DESCRIPTION

[0039] According to one exemplary embodiment of the disclosure, a method of providing a conversation service performed by an electronic device includes receiving an utterance input, identifying a temporal expression representing a time in text obtained from the utterance input, determining a time point related to the utterance input based on the temporal expression, selecting a database corresponding to the determined time point from among a plurality of databases storing conversation history information of a user using the conversation service, interpreting the text based on the conversation history information of the user obtained from the selected database, generating a response message to the utterance input based on a result of the interpretation, and outputting the generated response message.

[0040] According to another exemplary embodiment of the disclosure, an electronic device configured to provide a conversation service includes a memory storing one or more instructions and at least one processor configured to execute the one or more instructions to provide the conversation service to a user, wherein the at least one processor is further configured to execute the one or more instructions to control the electronic device to receive an utterance input, identify a temporal expression representing a time in text obtained from the utterance input, determine a time point related to the utterance input based on the temporal expression, select a database corresponding to the determined time point from among a plurality of databases storing information about a conversation history of a user using the conversation service, interpret the text based on conversation history information of the user obtained from the selected database, generate a response message to the utterance input based on a result of the interpretation, and output the generated response message.

[0041] According to another exemplary embodiment of the disclosure, a program for executing a method of providing a conversation service is stored in one or more non-transitory computer-readable recording media, and the method includes receiving an utterance input, identifying a time expression representing a time in text obtained from the utterance input, determining a time point related to the utterance input based on the time expression, selecting a database corresponding to the determined time point from among a plurality of databases storing conversation history information of a user using the conversation service, interpreting the text based on the conversation history information of the user obtained from the selected database, generating a response message to the utterance input based on a result of the interpretation, and outputting the generated response message.

[0042] Various exemplary embodiments of the disclosure will be described below in greater detail with reference to the accompanying drawings. However, the embodiments of the disclosure can have various forms, and should not be construed as being limited to the various exemplary embodiments set forth herein. In addition, portions irrelevant to the disclosure can be omitted in order to make the description of the disclosure clear, and the same reference numerals denote the same elements throughout the drawings.

[0043] Throughout the disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or a variation thereof.

[0044] Some embodiments of the disclosure can be described in terms of functional block components and various processing steps. All or some of the functional blocks can be implemented using any number of hardware and / or software components configured to perform the specified functions. For example, functional blocks of the disclosure can be implemented using one or more microprocessors or circuit components for performing the predetermined functions. Also, for example, functional blocks of the disclosure can be implemented in various programming or scripting languages. The functional blocks can be implemented in algorithms running on one or more processors. Also, the disclosure can employ techniques related to electronic configuration, signal processing, and / or data processing in the relevant art.

[0045] In addition, the connection lines or connectors shown in the drawings between elements are intended to represent only an exemplary functional relationship and / or physical or logical coupling between the elements. It should be noted that there can be many alternative or additional functional relationships, physical connections or logical connections in actual devices.

[0046] Various exemplary embodiments of the disclosure will be described below in greater detail with reference to the accompanying drawings.

[0047] Generally, a device for providing a conversation service to a plurality of users can manage a conversation on a per-user account basis or on a per-session basis without recognizing the users. A session can refer to, for example, a time period from the start to the end of a conversation service that performs speech recognition for a user's query and outputs a response thereto.

[0048] Generally, in the case of a device that manages a conversation on a per-user account basis, a user needs to register his or her account by inputting user information such as a user ID or a name. In addition, each time the user uses the device, the user can suffer from the inconvenience of having to input the user information and retrieve his or her account.

[0049] To reduce the inconvenience of the user, a device that manages a conversation on a per-session basis without recognizing the user can be used. In the case of a device that manages a conversation on a per-session basis, when a user utterance input is received after a session is terminated, the device starts a new session and provides a conversation service without providing information on the history of a previous conversation. Thus, when the user asks a question related to the details of a conversation that he or she had with the device during a previous session after a new session is started, the device can not be able to provide an accurate response to the user question.

[0050] To solve the above-described problem, the present disclosure provides a method performed by an electronic device that stores a conversation history of each user by automatically recognizing the user without requiring the user to register his or her account, and creates a response using the conversation history.

[0051] Figure 1 is a diagram illustrating an example in which a user 10 visited a store a month ago and purchased an air conditioner, a television (TV), and a refrigerator through an electronic device 200. According to an embodiment of the present disclosure, when the user 10 simply approaches the electronic device 200, the electronic device 200 can automatically perform face-based user authentication without requiring input of user information for user authentication. The electronic device 200 can initiate a conversation service via the face-based user authentication without a separate command of the user.

[0052] According to an embodiment of the present disclosure, the electronic device 200 can check that a face ID of the user 10 has been stored by recognizing the face of the user, and determine that the user 10 has used the conversation service before. The electronic device 200 can output a voice message "You are back again" based on the service usage history of the user.

[0053] According to an embodiment of the present disclosure, when no face ID matching the face of the user is found, the electronic device 200 can determine that the user 10 has never used the conversation service. When it is determined that the user 10 is using the conversation service for the first time, the electronic device 200 can output a voice message "Welcome to the first visit". According to an embodiment of the present disclosure, the electronic device 200 can output a user response message suitable for a situation based on the service usage history of the user.

[0054] As Figure 1 shown, the user 10 can make a request "Do you know what I purchased last time" related to products that the user 10 previously (e.g., a month ago) purchased through the electronic device 200.

[0055] Because the related art electronic device that manages a dialogue based on each session does not store the content of past dialogues with the user 10, the electronic device can not ensure continuity of the dialogue or provide an appropriate response message even when the user 10 makes an utterance related to the content of his or her past utterance.

[0056] On the other hand, the electronic device 200 according to an exemplary embodiment of the disclosure can interpret the utterance of the user 10 based on the past dialogue history of the user 10, thereby outputting a response message based on the past dialogue history.

[0057] For example, as shown in Figure 1 , the electronic device 200 can output a response message "You purchased an air conditioner, a television, and a refrigerator one month ago. Which product do you mean?" based on a dialogue history related to a product purchased by the user 10 through the electronic device 200 one month ago. According to an embodiment of the disclosure, when it is determined that the past dialogue history of the user is needed to interpret the utterance of the user, the electronic device 200 can interpret the utterance of the user based on the past dialogue history of the user matching the user face identified to be stored, and generate a response message based on the interpretation result.

[0058] Figure 2a FIG. 1 is a diagram illustrating an exemplary system for providing a dialogue service according to an embodiment of the disclosure, and Figure 2b FIG. 2 is a diagram illustrating an exemplary system for providing a dialogue service according to an embodiment of the disclosure.

[0059] As shown in Figure 2a , according to an embodiment of the disclosure, the electronic device 200 can be used alone to provide a dialogue service to the user 10. Examples of the electronic device 200 can include, but are not limited to, a home appliance (e.g., a television, a refrigerator, a washing machine, etc.), a smart phone, a PC, a wearable device, a PDA, a media player, a micro server, a global positioning system (GPS), an electronic book terminal, a digital broadcasting terminal, a navigation device, a kiosk, an MP3 player, a digital camera, other mobile or non-mobile computing devices, etc. The electronic device 200 can provide a dialogue service by, for example, executing a chatbot application or a dialogue agent application, etc.

[0060] According to an embodiment of the disclosure, the electronic device 200 can receive an utterance input of the user 10, and generate and output a response message to the received utterance input.

[0061] According to an embodiment of the disclosure, the electronic device 200 can provide a method of storing a dialogue history for each user and generating a response using the dialogue history by automatically identifying the user 10 without the user 10 registering an account.

[0062] For example, according to an embodiment of the present disclosure, the electronic device 200 can identify the user's face via the camera and search for a face ID matching the identified face from among the stored face IDs. The electronic device 200 can retrieve the conversation history and the conversation service usage history mapped to the found face ID. The electronic device 200 can provide the conversation service to the user 10 based on the conversation history and update the conversation history after ending the conversation with the user 10.

[0063] According to an embodiment of the present disclosure, when there is no face ID matching the identified face among the stored face IDs, the electronic device 200 can check with the user 10 whether to store information related to the user's face. When the user 10 agrees to store his or her face ID, the electronic device 200 can map the conversation history and the service usage history to his or her face ID so as to be stored after ending the conversation with the user 10.

[0064] According to an embodiment of the present disclosure, in managing the stored face IDs and the conversation history, the electronic device 200 can designate a maximum storage period in which the face IDs and the conversation history can be stored based on the fact that the storage capacity is limited and personal information needs to be protected. When the maximum storage period elapses, the electronic device 200 can delete the stored face IDs and the conversation history. However, when the user 10 is re-identified before the maximum storage period elapses (for example, when the user 10 re-visits the store), the electronic device 200 can extend and flexibly manage the storage period. The electronic device 200 can designate different information storage periods for each user according to the time interval and the number of times the user 10 uses the conversation service.

[0065] According to an embodiment of the present disclosure, when it is determined that the past conversation history and the service usage history are needed to interpret the user's question, the electronic device 200 can generate and output a response message based on the conversation history of the user 10.

[0066] The electronic device 200 can receive a user utterance input and determine a context related to a time included in the utterance input. The context related to the time included in the utterance input may, for example, relate to information about a time needed to generate a response according to the user's intention included in the utterance input. The electronic device 200 can determine which conversation history information to use from among information about the conversation history accumulated over a first period and information about the conversation history accumulated over a second period based on a result of the determination of the context. Accordingly, the electronic device 200 can reduce the amount of time needed to interpret the user's question and provide an appropriate response by identifying the context of the conversation based on only the selected conversation history information.

[0067] In addition, as Figure 2bAs illustrated, according to an embodiment of the disclosure, the electronic device 200 can provide a conversation service in connection with another electronic device 300 and / or a server 201. The electronic device 200, the other electronic device 300, and the server 201 can be connected to each other by wired or wireless means.

[0068] The other electronic device 300 or the server 201 can share data, resources, and services with the electronic device 200, perform control or file management of the electronic device 200, or monitor the entire network. For example, the other electronic device 300 can be a mobile or non-mobile computing device.

[0069] The electronic device 200 can generate and output a response message to a user utterance input through communication with the other electronic device 300 and / or the server 201.

[0070] As Figure 2a and Figure 2b As illustrated, according to an embodiment of the disclosure, a system for providing a conversation service can include at least one electronic device and / or a server. For ease of description, a method of providing a conversation service performed by an "electronic device" will be described below. However, some or all operations of the electronic device to be described below can also be performed by another electronic device and / or a server connected to the electronic device, and can be partially performed by a plurality of electronic devices.

[0071] Figure 3 is a flowchart illustrating an example method of providing a conversation service performed by the electronic device 200 according to an embodiment of the disclosure.

[0072] According to an embodiment of the disclosure, the electronic device 200 can receive a user utterance input (operation S310).

[0073] According to an embodiment of the disclosure, the electronic device 200 can initiate a conversation service and receive a user utterance input. According to an embodiment of the disclosure, the electronic device 200 can initiate a conversation service when a user approaches the electronic device 200 within a certain distance from the electronic device 200, when a voice signal of a predetermined intensity or higher is received, and when a voice signal for uttering a predetermined enabling word is received.

[0074] According to an embodiment of the disclosure, when a user approaches the electronic device 200 within a certain distance, the electronic device 200 can obtain a face image of the user and determine whether a face ID matching the obtained face image of the user is stored by searching a database. It should be understood that any suitable means for recognizing a user can be used, and the disclosure is not limited to face ID recognition. For example, but not limited to, voice recognition, biometric recognition, input of a user interface, etc. can be used, and the use of a face ID is for ease of explanation and illustration. The electronic device 200 can initiate a conversation service based on the determination result.

[0075] According to an embodiment of the disclosure, when a face ID matching the obtained face image of the user is stored in the database, the electronic device 200 can update the stored service usage history mapped to the face ID. Otherwise, when a face ID matching the obtained face image of the user is not stored in the database, the electronic device 200 can generate a new face ID and a service usage history mapped to the new face ID.

[0076] According to an embodiment of the disclosure, the electronic device 200 can receive and store an audio signal including a user utterance input via a microphone. For example, the electronic device 200 can detect the presence and absence of a person's voice in units of sentences by using voice activity detection (VAD), endpoint detection (EPD), etc., thereby receiving and storing the utterance input.

[0077] According to an embodiment of the disclosure, the electronic device 200 can identify a time expression representing a time obtained from a user utterance input (e.g., text) (operation S320).

[0078] The electronic device 200 can obtain text by performing voice recognition on the user utterance input, and determine an entity representing at least one of a time point, a duration, or a time period included in the text as a time expression. However, it should be understood that determining a time expression is not limited to obtaining and analyzing text.

[0079] The entity can include at least one of a word, a phrase, or a morpheme included in the text having a specific meaning. The electronic device 200 can identify at least one entity in the text, and determine which domain includes each entity according to the meaning of the at least one entity. For example, the electronic device 200 can determine whether the identified entity in the text is an entity representing, for example, but not limited to, a person, an object, a geographical area, a time, a date, etc.

[0080] According to an embodiment of the disclosure, the electronic device 200 can determine, as a time expression, a time point representing an operation or a state indicated by, for example, but not limited to, text, or adverbs, adjectives, nouns, words, phrases, etc. included in the text representing a time point, a time, a time period, etc.

[0081] The electronic device 200 can perform an embedding for mapping a text to a plurality of vectors. The electronic device 200 can assign a beginning-inside-outside (BIO) tag to at least one morpheme included in the text that represents at least one of a time point, a duration, or a time period by applying a bidirectional long short-term memory (LSTM) model to the mapped vectors. The electronic device 200 can identify an entity representing a time in the text based on the BIO tag. The electronic device 200 can determine the identified entity as a temporal expression.

[0082] According to an embodiment of the disclosure, the electronic device 200 can determine a time point related to the user utterance input based on the temporal expression (operation S330).

[0083] The time point related to the user utterance input can be, for example, a time point when information necessary for the electronic device 200 to generate a response according to an intent in the user utterance, a time point when a past utterance input of the user including the information is received, or a time point when the information is stored in a dialogue history. The time point related to the user utterance input can include a time point when a past utterance of the user including information for specifying content of the user utterance input is received or stored. For example, the time point related to the user utterance input can include a time point when a past utterance related to the user purchasing, mentioning, inquiring about a product mentioned in the user utterance input, and a service request for the product is received or stored.

[0084] The electronic device 200 can predict a probability value, for example, a probability that the temporal expression indicates each of a plurality of time points. The electronic device 200 can determine a time point corresponding to a highest probability value among the predicted probability values as the time point related to the user utterance input.

[0085] According to an embodiment of the disclosure, the electronic device 200 can select a database corresponding to the time point related to the user utterance input from among a plurality of databases storing information related to a dialogue history of a user using a dialogue service (operation S340). On the other hand, a single database can be used, and the disclosure is not limited to a plurality of databases.

[0086] The plurality of databases can include a first database for storing information about a user conversation history accumulated before a preset time point and a second database for storing information about a user conversation history accumulated after the preset time point. The electronic device 200 can select the first database when a time point related to the user utterance input is before the preset time point. The electronic device 200 can select the second database when the time point related to the user utterance input is after the preset time point. Further, the first database included in the database can be stored in an external server, and the second database can be stored in the electronic device 200.

[0087] According to an embodiment of the disclosure, the preset time point can be one of: a time point when at least some of the information about the user's conversation history included in the second database is transmitted to the first database, a time point when a face image of the user is obtained, and a time point when the conversation service starts.

[0088] According to an embodiment of the disclosure, the electronic device 200 can interpret the text based on the information about the user's conversation history acquired from the selected database (operation S350).

[0089] The electronic device 200 can determine an entity included in the text and requiring detailed explanation. The electronic device 200 can acquire detailed explanation information for detailed explanation of the determined entity by retrieving the information about the user's conversation history acquired from the selected database. The electronic device 200 can interpret the text and the detailed explanation information using, for example, but not limited to, a natural language understanding (NLU) model or the like.

[0090] The information about the user's conversation history acquired from the database can include, for example, but not limited to, past utterance inputs received from the user, past response messages provided to the user, information related to the past utterance inputs and the past response messages, and the like. For example, the information related to the past utterance inputs and the past response messages can include an entity included in the past utterance input, content included in the past utterance input, a category of the past utterance input, a time point when the past utterance input is received, an entity included in the past response message, content included in the past response message, a time point when the past response message is output, information about a situation before and after the past utterance input is received, and a user's interest product, mood, payment information, voice characteristics, and the like, which are determined based on the past utterance input and the past response message.

[0091] According to an embodiment of the disclosure, the electronic device 200 can generate a response message to the received user utterance input based on the interpretation result (operation S360).

[0092] The electronic device 200 can determine a type of a response message, for example, and without limitation, by applying a dialog manager (DM) model to the interpretation result. The electronic device 200 can generate a response message of the determined type using, for example, and without limitation, a natural language generation (NLG) model.

[0093] According to an embodiment of the disclosure, the electronic device 200 can output the generated response message (operation S370). For example, the electronic device 200 can output the response message in the form of at least one of voice, text, or an image.

[0094] According to an embodiment of the disclosure, the electronic device 200 can share at least one of a user face ID, a service usage history, or a dialog history with another electronic device. For example, after a dialog service provided to the user ends, the electronic device 200 can transmit a face ID of the user to another electronic device. When the user wishes to receive a dialog service through another electronic device, the other electronic device can identify the user, and based on determining that the identified user corresponds to the received face ID of the user, request information about a dialog history of the user from the electronic device 200. In response to the request received from the other electronic device, the electronic device 200 can transmit information about the dialog history of the user included in the second database stored in the database to the other electronic device.

[0095] Figure 4a FIG. 2 is a diagram illustrating an example in which the electronic device 200 provides a dialog service based on a dialog history according to an embodiment of the disclosure, Figure 4b FIG. 2 is a diagram illustrating an example in which the electronic device 200 provides a dialog service based on a dialog history according to an embodiment of the disclosure.

[0096] Figure 4a FIG. 2 is a diagram illustrating an example in which the electronic device 200 provides a dialog service based on a dialog history according to an embodiment of the disclosure. Figure 4a On May 5, 2019, the electronic device 200 can receive an utterance input of the user 10 inquiring about a price of the air conditioner A. The electronic device 200 can output a response message informing the price of the air conditioner A in response to the user utterance input.

[0097] On May 15, 2019, the electronic device 200 can receive a user 10's utterance input of re-visiting a store. The electronic device 200 can obtain text of saying "I asked you the price of the air conditioner last time?" through voice recognition of the user's utterance input. The electronic device 200 can identify a temporal expression in the obtained text (e.g., "last time," "asked," and "price of the air conditioner"). The electronic device 200 can determine that the user's dialogue history is needed to interpret the obtained text based on the identified temporal expression.

[0098] The electronic device 200 can determine an entity in the text that needs to be specified in detail, and acquire specification information for specifying the entity based on the user's dialogue history. The electronic device 200 can interpret the text in which the entity is specified using an NLU model.

[0099] According to an embodiment of the disclosure, the electronic device 200 can determine an entity in the text that represents a category of a product, and specify the product based on dialogue history information. For example, as shown in Figure 4a The electronic device 200 can determine "air conditioner" as an entity that represents a category of a product entity in the user's utterance "I asked you the price of the air conditioner last time?" The electronic device 200 can determine that the product of which the user wants to know the price is "air conditioner A" based on the dialogue history of the date of May 5. The electronic device 200 can output a response message informing the price of air conditioner A in response to the user's question.

[0100] Figure 4b is a diagram illustrating an example in which the electronic device 200 is a self-checkout counter in a restaurant according to an embodiment of the disclosure. Referring to Figure 4b On May 10, 2019, the electronic device 200 can receive a user 10's utterance input of ordering a salad. The electronic device 200 can output a response message informing confirmation of ordering a salad in response to the user's utterance input.

[0101] On May 15, 2019, the electronic device 200 can receive a user 10's utterance input of re-visiting a restaurant. The electronic device 200 can obtain text of stating "order that one I always eat" through voice recognition of the user's utterance input. The electronic device 200 can identify a temporal expression in the obtained text (e.g., "always" and "eat"). The electronic device 200 can determine that the user's dialogue history is needed to interpret the obtained text based on the identified temporal expression.

[0102] According to an embodiment of the disclosure, the electronic device 200 can determine a noun in the text that needs to be specified in detail as an entity that needs to be specified in detail, and specify an object indicated by the noun based on dialogue history information. For example, as shown in Figure 4bAs illustrated, the electronic device 200 can determine "that" in the user's utterance "order that which I usually eat" as an entity that needs to be specified in detail. The electronic device 200 can determine that the food that the user 10 wants to order is "salad" based on the conversation history for the date of May 10. The electronic device 200 can output a response message requesting confirmation of whether it is correct to order the salad in response to the user utterance input.

[0103] In addition, according to an embodiment of the disclosure, in order to perform a smooth conversation with the user 10, the electronic device 200 can need to shorten the time required to generate a response message based on a conversation history. Accordingly, according to an embodiment of the disclosure, the electronic device 200 can shorten the time required to retrieve conversation history information by searching only a database selected from among a plurality of databases that store information related to a conversation history.

[0104] Figure 5 FIG. 2B is a diagram illustrating an example of a process of providing a conversation service by the electronic device 200 according to an embodiment of the disclosure.

[0105] Figure 5 FIG. 2C illustrates an example in which the electronic device 200 according to an embodiment of the disclosure is a self-service kiosk in a store selling electronic products. According to an embodiment of the disclosure, the electronic device 200 can use a first database 501 for storing conversation history information accumulated before a current session for providing a conversation service is started and a second database 502 for storing information related to a conversation performed during the current session.

[0106] Figure 6 FIG. 3 is a diagram illustrating an example of a service usage history and a conversation history for each user stored in the first database 501 according to an embodiment of the disclosure. As Figure 6 illustrated, the electronic device 200 can map the service usage history 620 and the conversation history 630 to the face ID of the user 10 for storage.

[0107] The service usage history 620 can include, for example, and without limitation, the number of visits made by the user 10, a visit period, a last visit date 621, and a date on which the service usage history 620 is scheduled to be deleted. The conversation history 630 can include information about products in which the user is interested and categories of past questions of the user (which are determined based on past questions of the user), whether the user 10 purchased the product in which the user is interested, a time point at which the user's question was received.

[0108] The electronic device 200 can identify the user 10 by performing face recognition on a face image of the user 10 obtained via a camera (operation S510). For example, the electronic device 200 can retrieve a conversation service use history corresponding to a face ID of the identified user 10 by searching the first database 501.

[0109] The electronic device 200 can update the conversation service use history of the identified user 10 by adding information indicating that the user 10 is currently using the conversation service (operation S520).

[0110] For example, as shown in Figure 6 The electronic device 200 can update the conversation service use history 620 of the user 10 in such a way that information 621 indicating that the user 10 used the conversation service on May 16, 2019 is added.

[0111] According to an embodiment of the disclosure, the electronic device 200 can delay scheduling a date for deleting information related to the user 10 based on at least one of the number of visits or the visit period of the user 10 recorded in the user's use history. For example, as the number of visits of the user 10 increases and the visit period becomes shorter, the electronic device 200 can extend the storage period of information related to the user 10.

[0112] The electronic device 200 can initiate the conversation service and receive a user utterance input (operation S530). The electronic device 200 can obtain a text stating "I have a question about the product I purchased last time" from the user utterance input. The electronic device 200 can identify "last time" as a temporal expression representing time in the obtained text.

[0113] The electronic device 200 can determine that information related to the user conversation history accumulated before a time point at which the current session starts is needed to interpret the obtained text based on the temporal expression "last time".

[0114] The electronic device 200 can retrieve information about the user conversation history from the first database 501 based on the determination that information related to the user conversation history accumulated before the start of the current session is needed (operation S540). The electronic device 200 can determine "product" included in the user's utterance "I have a question about the product I purchased last time" as an entity that needs to be specified in detail, and interpret the text based on information related to the user's conversation history obtained from the first database 501.

[0115] For example, the electronic device 200 can determine that the product the user 10 wants to refer to is "Computer B" having a model name 19COMR1 based on the conversation history 631 for the date of May 10.

[0116] In response to the user utterance input, the electronic device 200 can output a response message confirming whether the user 10 visited the store due to a problem with the computer B purchased on May 10 (operation S550).

[0117] After ending the conversation service, the electronic device 200 can update information about a conversation history of the user in the first database 510 in a manner of adding information about a history of a conversation performed during the conversation (operation S560).

[0118] Figure 7a FIG. 6 is a flowchart illustrating an example method of providing a conversation service by the electronic device 200 according to an embodiment of the disclosure, Figure 7b FIG. 6 is a flowchart illustrating an example method of providing a conversation service by the electronic device 200 according to an embodiment of the disclosure.

[0119] The user can start using the conversation service by approaching the electronic device 200.

[0120] According to an embodiment of the disclosure, the electronic device 200 can recognize the face of the user through the camera (operation S701). According to an embodiment of the disclosure, the electronic device 200 can search a database to find a stored face ID corresponding to the recognized user (operation S702). According to an embodiment of the disclosure, the electronic device 200 can determine whether the face ID corresponding to the recognized user is stored in the database (operation S703).

[0121] When the face ID of the user is stored in the database (YES in operation S703), the electronic device 200 can retrieve the service usage history of the user. According to an embodiment of the disclosure, the electronic device 200 can update the service usage history of the user (operation S705). For example, the electronic device 200 can update information about the date of the last visit of the user included in the service usage history of the user.

[0122] According to an embodiment of the disclosure, when the face ID of the user is not stored in the database (NO in operation S703), the electronic device 200 can ask the user whether he or she agrees to store the face ID and the conversation history of the user in the future (operation S704). According to an embodiment of the disclosure, when the user agrees to store his or her face ID and the conversation history (YES in operation S704), the electronic device 200 can update the service usage history of the user in operation S705. According to an embodiment of the disclosure, when the user does not agree to store his or her face ID and the conversation history, the electronic device 200 can perform a conversation with the user on a per-conversation basis.

[0123] Referring to Figure 7bAccording to embodiments of the disclosure, the electronic device 200 can receive a user's utterance input (operation S710).

[0124] According to embodiments of the disclosure, the electronic device 200 can determine whether the dialogue history information is needed to interpret the user's utterance input (operation S721). This will be described in greater detail below with reference to FIG. 7B. Figure 8 Operation S721 is described in greater detail.

[0125] When it is determined that the dialogue history information is not needed for interpretation (NO in operation S721), according to embodiments of the disclosure, the electronic device 200 can generate and output a general response without using the dialogue history information (operation S731).

[0126] When it is determined that the dialogue history information is needed for interpretation (YES in operation S721), according to embodiments of the disclosure, the electronic device 200 can determine whether the dialogue history information included in the first database is needed (operation S723).

[0127] For example, according to embodiments of the disclosure, the electronic device 200 can determine whether the dialogue history information of the user accumulated and stored in the first database before a preset time point is needed, or whether the dialogue history information of the user accumulated and stored in the second database after the preset time point is needed, based on the preset time point. The first database can store dialogue history information accumulated over a relatively long period of time from when the dialogue history of the user was first stored to the preset time point. The second database can store dialogue history information accumulated over a short period of time from the preset time point to the current time point. This will be described in greater detail below with reference to FIG. 7C. Figure 9 Operation S723 is described in greater detail.

[0128] When it is determined that the dialogue history information included in the first database is needed to interpret the user's utterance input (YES in operation S723), according to embodiments of the disclosure, the electronic device 200 can generate a response message based on the dialogue history information acquired from the first database (operation S733). When it is determined that the dialogue history information included in the first database is not needed to interpret the user's utterance input (NO in operation S723), according to embodiments of the disclosure, the electronic device 200 can generate a response message based on the dialogue history information acquired from the second database (operation S735).

[0129] According to embodiments of the disclosure, the electronic device 200 can output the generated response message (operation S740).

[0130] According to an embodiment of the disclosure, the electronic device 200 can determine whether the conversation has ended (operation S750). For example, when the distance of the user from the electronic device 200 is greater than or equal to a threshold distance, when no user utterance input is received for more than a threshold time, or when it is determined that the user deviates from any space (e.g., a store or a restaurant) in which the electronic device 200 is located, the electronic device 200 can determine that the conversation has ended.

[0131] When it is determined that the conversation has ended (YES in operation S750), according to an embodiment of the disclosure, the electronic device 200 can also store the conversation history information related to the current conversation in the stored conversation history mapped to the user face ID (operation S760). Otherwise, when it is determined that the conversation has not ended (NO in operation S750), according to an embodiment of the disclosure, the electronic device 200 can return to operation S710 and repeat the process of receiving a user utterance input and generating a response message to the user utterance input.

[0132] Figure 8 FIG. 8 is a flowchart illustrating an example method of determining whether the electronic device 200 will generate a response based on conversation history information according to an embodiment of the disclosure.

[0133] For example, Figure 7b Operation S721 of FIG. 7A can be subdivided into operations S810, S820, and S830 of FIG. 8. Figure 8 Operation S810, S820, and S830 of FIG. 8.

[0134] According to an embodiment of the disclosure, the electronic device 200 can receive a user utterance input (operation S710). According to an embodiment of the disclosure, the electronic device 200 can obtain text by performing speech recognition (e.g., automatic speech recognition (ASR)) on the received user utterance input (operation S810).

[0135] According to an embodiment of the disclosure, the electronic device 200 can extract a temporal expression from the obtained text (operation S820). According to an embodiment of the disclosure, the electronic device 200 can extract a temporal expression by applying a pre-trained temporal expression extraction model to the obtained text.

[0136] According to an embodiment of the disclosure, the electronic device 200 can determine whether the extracted temporal expression represents a past time point, a time period, or a duration (operation S830). When the extracted temporal expression is not a past temporal expression (NO in operation S830), the electronic device 200 can generate a response to the user utterance input based on general NLU that does not consider the dialogue history (operation S841). Otherwise, when the extracted temporal expression is a past temporal expression (YES in operation S830), the electronic device 200 can determine that a response needs to be generated based on the dialogue history information (operation S843).

[0137] Figure 9 FIG. 8 is a flowchart illustrating an example method of selecting a database based on a user utterance input, according to an embodiment of the disclosure.

[0138] For example, the operation S723 of FIG. 7 can be subdivided into operations S910, S920, S930, and S940 of FIG. 9. Figure 7b Figure 9

[0139] According to an embodiment of the disclosure, the electronic device 200 can determine that the dialogue history information is needed to interpret the user utterance input at operation S843.

[0140] According to an embodiment of the disclosure, the electronic device 200 can extract a temporal expression from text obtained based on the user utterance input (operation S910). According to an embodiment of the disclosure, the electronic device 200 can extract an expression representing the past from the extracted temporal expression (operation S920). Because the operation S910 of FIG. 9 corresponds to the operation S820 of FIG. 8, an embodiment of the disclosure can not perform the operation S910. Figure 9 Figure 8 Figure 9 When the operation S910 of FIG. 9 is not performed, the electronic device 200 can use the temporal expression extracted and stored in operation S820.

[0141] According to an embodiment of the disclosure, the electronic device 200 can predict a time point related to the user utterance input based on the temporal expression representing the past (operation S930). According to an embodiment of the disclosure, the electronic device 200 can determine a time point related to the extracted past temporal expression by applying a pre-trained time point prediction model to the extracted past temporal expression.

[0142] As described above, the electronic device 200 can determine a time point related to the past temporal expression by applying a pre-trained time point prediction model to the past temporal expression. Figure 10 ​​​​As illustrated, the electronic device 200 can predict probability values, for example, representing probabilities of the past time schedule indicating each of a plurality of time points, and generate a graph 1000 representing the predicted probability values. The electronic device 200 can determine a time point corresponding to a highest probability value 1001 among the predicted probability values as a time point related to the user utterance input. In the graph 1000, an x-axis and a y-axis can represent time and probability values, respectively. A zero point on a time axis in the graph 1000 represents a preset time point serving as a reference point for selecting a database.

[0143] According to an embodiment of the disclosure, the electronic device 200 can determine whether the predicted time point is before the preset time point (operation S940).

[0144] When the predicted time point is before the preset time point (YES in operation S940), the electronic device 200 can generate a response to the user utterance input based on the dialogue history information acquired from the first database (operation S733). When the predicted time point is at or after the preset time point (NO in operation S940), the electronic device 200 can generate a response to the user utterance input based on the dialogue history information acquired from the second database (operation S735).

[0145] According to an embodiment of the disclosure, the electronic device 200 can manage a plurality of databases according to a period in which dialogue history is accumulated, thereby reducing a time required to retrieve dialogue history. According to an embodiment of the disclosure, the electronic device 200 can switch between databases such that at least some information about dialogue history of a user stored in one database is stored in another database.

[0146] In the disclosure, although Figure 10 An example in which the electronic device 200 uses two databases is illustrated, but embodiments of the disclosure are not limited thereto. The database used by the electronic device 200 can include three or more databases. For ease of description, in the disclosure, a case in which the database includes the first and second databases is described as an example.

[0147] Figure 11 is a graph illustrating an example method of switching a database in which dialogue history of a user is stored, performed by the electronic device 200, according to an embodiment of the disclosure.

[0148] According to an embodiment of the disclosure, the first database 1101 can store information about a user conversation history accumulated before a preset time point, and the second database can store information about a user conversation history accumulated after the preset time point. For example, the preset time point can be one of a time point at which at least some information about a user conversation history included in the second database 1102 is transmitted to the first database 1101, a time point at which a user face image is obtained, a time point at which a conversation service is started, and a time point occurring a predetermined time before a current time point, but the disclosure is not limited thereto.

[0149] According to an embodiment of the disclosure, the first database 1101 can store conversation history information accumulated for a relatively long period of time from when a user's conversation history is first stored to a preset time point. The second database 1102 can store conversation history information accumulated for a short period of time from the preset time point to a current time point.

[0150] For example, the first database 1101 can be included in an external server, and the second database 1102 can be included in the electronic device 200. The first database 1101 can also store a service usage history of the user.

[0151] According to an embodiment of the disclosure, the electronic device 200 can switch at least some information about a user's conversation history stored in the second database 1102 to the first database 1101.

[0152] According to an embodiment of the disclosure, the electronic device 200 can switch a database in which a user's conversation history information is stored periodically or after starting or ending a specific operation or when a storage space of the database is insufficient.

[0153] For example, the electronic device 200 can transmit information about a user's conversation history stored in the second database 1102 to the first database 1101 according to a predetermined period of time (for example, but not limited to, 6 hours, 1 day, 1 month, etc.), and delete information about a user's conversation history from the second database 1102.

[0154] As another example, when a conversation service ends, the electronic device 200 can transmit information about a user's conversation history accumulated in the second database 1102 to the first database 1101 while providing the conversation service, and delete information about a user's conversation history from the second database 1102.

[0155] According to an embodiment of the disclosure, when switching between databases, the electronic device 200 can abstract information excluding sensitive information of the user, thereby mitigating the risk of user personal information leakage and reducing memory usage.

[0156] Raw data, which is unprocessed data, can be stored in the second database 1102. The second database 1102 can store the raw data as the user's conversation history information as it is input to the electronic device 200.

[0157] For example, the user can not want to store detailed information (e.g., specific conversation content, a captured user image, a user's voice, a user's billing information, a user's location, etc.) related to the user's personal information in the electronic device 200 for a long time. Accordingly, according to an embodiment of the disclosure, the electronic device 200 can manage the conversation history information including information sensitive to the user such that the conversation history information is stored in the second database 1102, which stores the conversation history information for only a short time.

[0158] Processed data can be stored in the first database 1101. The first database 1101 can store data summarized by excluding information sensitive to the user from the raw data stored in the second database 1102 as the user's conversation history information.

[0159] As Figure 11 indicated, the raw content of a conversation between the user and the electronic device 200 stored in the second database 1102 can be summarized as data about a conversation category, content, and a product of interest, and stored in the first database 1101. The captured image frames of the user and the user's voice stored in the second database 1102 can be summarized as the user's mood at a point in time when the conversation service is provided, and stored in the first database 1101. In addition, the user's payment information stored in the second database 1102 can be summarized as a product purchased by the user and a purchase price, and stored in the first database 1101.

[0160] According to an embodiment of the disclosure, the electronic device 200 can share at least one of a user face ID, a service usage history, or a conversation history with other electronic devices. Figure 12a is a diagram illustrating an exemplary process according to an embodiment of the disclosure, in which a plurality of electronic devices 200-a, 200-b, and 200-c share a user's conversation history with each other. For example, the electronic devices 200-a, 200-b, and 200-c can be self-service kiosks located in different spaces (e.g., different floors) of a store. The user 10 can receive guidance about a product or help to purchase a product based on a conversation service provided by the electronic devices 200-a, 200-b, and 200-c.

[0161] Referring to Figure 12a , the electronic device 200-c can provide a conversation service to the user 10. The electronic device 200-c can receive an utterance input of the user 10 and generate and output a response message to the utterance input.

[0162] Figure 12b is a diagram illustrating an exemplary process in which a plurality of electronic devices 200-a, 200-b, and 200-c share a user's conversation history with each other, according to an embodiment of the disclosure.

[0163] Referring to Figure 12b After completing the negotiation with the electronic device 200-c, the user 10 can move away from the electronic device 200-c by a distance greater than or equal to a predetermined distance. The electronic device 200-c can identify that the conversation is paused based on the distance from the user 10. The electronic device 200-c can store information about the history of the conversation with the user 10 performed during the current session in the database. For example, the electronic device 200-c can store information about the conversation history with the user 10 performed during the current session in the second database included in the electronic device 200-c.

[0164] Figure 12c is a diagram illustrating an exemplary process in which a plurality of electronic devices 200-a, 200-b, and 200-c share a user's conversation history with each other, according to an embodiment of the disclosure.

[0165] Referring to Figure 12c The electronic device 200-c can share or broadcast the face ID of the user 10 with which the negotiation has been completed, for example, and without limitation, to the other electronic devices 200-a and 200-b in the store.

[0166] Figure 12d is a diagram illustrating an exemplary process in which a plurality of electronic devices 200-a, 200-b, and 200-c share a user's conversation history with each other, according to an embodiment of the disclosure.

[0167] Referring to Figure 12d After viewing the second floor of the store, the user 10 can go down to the first floor and approach the electronic device 200-a. The electronic device 200-a can identify the user 10 to provide a conversation service to the user 10. When it is determined that the identified user 10 corresponds to the face ID shared by the electronic device 200-c, the electronic device 200-a can request the electronic device 200-c to share the database in which information related to the shared face ID is stored.

[0168] The electronic device 200-c can share a database in which a conversation history corresponding to the face ID of the user 10 is stored with the electronic device 200-a. The electronic device 200-a can interpret the user utterance input based on the conversation history stored in the shared database. Accordingly, even when the electronic device 200-a receives an utterance input related to a conversation with the electronic device 200-c from the user 10, the electronic device 200-a can output a response message that guarantees continuity of the conversation.

[0169] The configuration of the electronic device 200 according to an embodiment of the disclosure will now be described in greater detail. Each component of the electronic device 200 to be described below can perform each operation of the method of providing a conversation service performed by the electronic device 200 as described above.

[0170] Figure 13a is a block diagram illustrating an exemplary configuration of an exemplary electronic device 200 according to an embodiment of the disclosure.

[0171] The electronic device 200 for providing a conversation service can include a processor (e.g., including processing circuitry) 250 that provides a conversation service to a user by executing one or more instructions stored in a memory. Although Figure 13a The electronic device 200 is illustrated as including one processor 250, but embodiments of the disclosure are not limited thereto. The electronic device 200 can include a plurality of processors. When the electronic device 200 includes a plurality of processors, the operations and functions of the processor 250 to be described below can be partially performed by the processors.

[0172] The inputter 220 of the electronic device 200 can include various input circuitry and receive a user utterance input.

[0173] According to an embodiment of the disclosure, the processor 250 can identify a time expression representing a time in the text obtained from the user utterance input.

[0174] The processor 250 can obtain a text by performing speech recognition on the user utterance input and perform embedding for mapping the text to a plurality of vectors. For example, by applying a bidirectional LSTM model to the mapping vectors, the processor 250 can assign a BIO tag to at least one morpheme including at least one of a time point, a duration, or a time period representing a time in the text. The processor 250 can determine an entity including at least one of a time point, a duration, or a time period representing a time as a time expression based on the BIO tag.

[0175] According to an embodiment of the disclosure, the processor 250 can determine a time point related to the user utterance input based on the time expression.

[0176] The processor 250 can predict probability values, for example, the recognized time expression indicating the probability of each of a plurality of time points, and determine the time point corresponding to the highest probability value among the predicted probability values ​​as the time point related to the user's utterance input.

[0177] According to embodiments of this disclosure, processor 250 can select from a plurality of databases used to store information related to the dialogue history of users using the dialogue service the database corresponding to the time point associated with the user's utterance input.

[0178] Multiple databases may include a first database for storing information about user dialogue history accumulated before a preset time point and a second database for storing information about user dialogue history accumulated after the preset time point. When the time point associated with the user's speech input is before the preset time point, the processor 250 can select the first database from the databases. When the time point associated with the user's speech input is after the preset time point, the processor 250 can select the second database from the databases.

[0179] Furthermore, the first database can be stored on an external server, while the second database can be stored in the electronic device 200. The preset time point t used as a reference point for selecting the database can be one of the following: when at least some information about the user's conversation history included in the second database is switched to be included in the first database; when the user's facial image is obtained; and when the conversation service is started.

[0180] According to embodiments of this disclosure, processor 250 can interpret text based on information related to a user's conversation history obtained from a selected database.

[0181] Processor 250 can identify entities included in the text that require detailed description. Processor 250 can obtain detailed description information for the identified entities by retrieving information about the user's dialogue history from a selected database. Processor 250 can use an NLU model to interpret the text and detailed description information. Processor 250 can determine the type of response message by applying a DM model to the interpretation results and use an NLG model to generate a response message of the determined type.

[0182] The processor 250 can generate a response message to the received user speech input based on the interpretation results. The output device 230 of the electronic device 200 may include various output circuits and output the generated response message.

[0183] The configuration of the electronic device 200 according to various embodiments of the present disclosure is not limited to... Figure 13a The configuration shown in the block diagram. For example, Figure 13bis a block diagram illustrating an exemplary configuration of an exemplary electronic device 200 according to another embodiment of the disclosure.

[0184] Referring to Figure 13b , the electronic device 200 according to another embodiment of the disclosure can include a communicator 210 that can have various communication circuitry, and receive a user utterance input via an external device and transmit a response message to the user utterance input to the external device. The processor 250 can select a database based on a time point related to the user utterance input, and generate the response message based on a user conversation history stored in the selected database. The description already provided above with respect to Figure 13a is omitted.

[0185] Figure 14 is a block diagram illustrating an exemplary electronic device 200 according to an embodiment of the disclosure.

[0186] As Figure 14 indicated, the inputter 220 of the electronic device 200 can include various input circuitry and receive a user input for controlling the electronic device 200. According to an embodiment of the disclosure, the inputter 220 can include a user input device including a touch panel for receiving a user's touch, a button for receiving a user's pressing operation, a wheel for receiving a user's rotating operation, a keyboard, a dome switch, etc., but is not limited thereto. For example, the inputter 220 can include at least one of, for example, and without limitation, a camera 221 for recognizing a user's face, a microphone 223 for receiving a user utterance input, or a payment device 225 for receiving a user's payment information.

[0187] According to an embodiment of the disclosure, the outputter 230 of the electronic device 200 can include various output circuitry and output information that is received from the outside, processed by the processor 250, or stored in the memory 270 or at least one database 260 in the form of at least one of, for example, and without limitation, light, sound, an image, or vibration. For example, the outputter 230 can include at least one of a display 231 for outputting a response message to a user utterance input or a speaker 233.

[0188] According to an embodiment of the disclosure, the electronic device 200 can further include at least one database 260 for storing a user's conversation history. According to an embodiment of the disclosure, the database 260 included in the electronic device 200 can store a user's conversation history information accumulated before a preset time point.

[0189] According to an embodiment of the disclosure, the electronic device 200 can further include a memory 270. The memory 270 can include at least one of data used by the processor 250, a result processed by the processor 250, a command executed by the processor 250, or an artificial intelligence (AI) model used by the processor 250.

[0190] The memory 270 can include at least one type of storage medium, such as a flash memory type storage medium, a hard disk type storage medium, a multimedia card micro type storage medium, a card type storage medium (e.g., an SD card or an XD memory), a random access memory (RAM), a static RAM (SRAM), a read-only memory (ROM), an electrically erasable programmable ROM (EEPROM), a PROM, a magnetic storage medium, a magnetic disk, or an optical disk.

[0191] Although Figure 14 Although the database 260 and the memory 270 are illustrated as separate components, embodiments of the disclosure are not limited thereto. For example, the database 260 can be included in the memory 270.

[0192] According to an embodiment of the disclosure, the communicator 210 can include various communication circuitry and communicate with an external electronic device or a server using a wireless or wired communication method. For example, the communicator 210 can include a short-range wireless communication module, a wired communication module, a mobile communication module, and a broadcast receiving module.

[0193] According to an embodiment of the disclosure, the electronic device 200 can share, for example, and without limitation, at least one of a user's face ID, a service use history, or a conversation history with another electronic device via the communicator 210. For example, after a conversation service provided to a user ends, the electronic device 200 can transmit the user's face ID to another electronic device. When the user wishes to receive a conversation service through another electronic device, the other electronic device can identify the user, and based on determining that the identified user corresponds to the received user's face ID, request information about the user's conversation history from the electronic device 200. In response to the request received from the other electronic device, the electronic device 200 can transmit information about the user's conversation history stored in the database 260 to the other electronic device.

[0194] Figure 15a is a block diagram illustrating an exemplary processor 250 included in the electronic device 200 according to an embodiment of the disclosure, Figure 15b is a block diagram illustrating another exemplary processor 250 according to an embodiment of the disclosure.

[0195] According to an embodiment of the disclosure, operations and functions performed by the processor 250 included in the electronic device 200 can be performed by Figure 15aThe various modules shown in FIG. 21 can be implemented using various numbers of hardware and / or software components that perform certain functions.

[0196] The face recognition module 1510 can include various processing circuitry and / or executable program elements, and is a module for recognizing a face in an image captured through a camera (221) of the electronic device 200. Figure 14

[0197] The service management module 1520 can include various processing circuitry and / or executable program elements, and is a module for managing a use history of the electronic device 200 by a user, and can manage an activity history such as purchasing a product and / or searching for product information via the electronic device 200.

[0198] The voice recognition module 1530 can include various processing circuitry and / or executable program elements, and obtains text from a user utterance input and generates a response message to the user utterance input based on a result of interpreting the text.

[0199] The database management module 1540 can include various processing circuitry and / or executable program elements, and selects at least one database for obtaining conversation history information from a plurality of databases, and manages a period during which information stored in the database is deleted.

[0200] Referring to Figure 15b , Figure 15a The face recognition module 1510 of the electronic device 200 can include a face detection module including various processing circuitry and / or executable program elements for detecting a face in an image, and a face search module including various processing circuitry and / or executable program elements for searching a database with respect to the detected face.

[0201] In addition, referring to Figure 15b , Figure 15a ​The voice recognition module 1530 can include at least one of an automatic speech recognition (ASR) module including various processing circuitry and / or executable program elements for converting a voice signal into a text signal, an NLU module including various processing circuitry and / or executable program elements for interpreting a meaning of the text, an entity extraction module including various processing circuitry and / or executable program elements for extracting an entity included in the text, a classification module including various processing circuitry and / or executable program elements for classifying the text according to a category of the text, a context management module including various processing circuitry and / or executable program elements for managing a dialogue history, a temporal context detection module including various processing circuitry and / or executable program elements for detecting a temporal expression in a user utterance input, or a NLG module including various processing circuitry and / or executable program elements for generating a response message corresponding to a result of interpreting the text and the temporal expression.

[0202] Further, referring to Figure 15b , Figure 15a The database management module 1540 can include a deletion period management module including various processing circuitry and / or executable program elements for managing a period during which information stored in the first database 1561 or the second database 1562 is deleted, and a database selection module including various processing circuitry and / or executable program elements for selecting at least one database from the first and second databases 1561 and 1562 to acquire and store information.

[0203] Figure 16 is a block diagram illustrating an exemplary configuration of a voice recognition module 1530 according to an embodiment of the disclosure.

[0204] Referring to Figure 16 , according to an embodiment of the disclosure, the voice recognition module 1530 included in the processor 250 of the electronic device 200 can include an ASR module 1610, an NLU module 1620, a DM module 1630, a NLG module 1640, and a text-to-speech (TTS) module 1650, each of which can include various processing circuitry and / or executable program elements.

[0205] The ASR module 1610 can convert a voice signal into text. The NLU module 1620 can interpret a meaning of the text. The DM module 1630 can guide a dialogue by managing context information including a dialogue history, determining a category of a question, and generating a response to the question. The NLG module 1640 can convert a response written in a computer language into a natural language that a human can understand. The TTS module 1650 can convert text into a voice signal.

[0206] According to an embodiment of the disclosure,Figure 16 The NLU module 1620 can interpret the text obtained from the user utterance by performing pre-processing (1621), performing embedding (1623), applying a temporal expression extraction model (1625), and applying a time point prediction model (1627).

[0207] In operation 1621 of performing pre-processing, the speech recognition module 1530 can remove special characters included in the text, unify synonyms into a single word, and perform morphological analysis through a part-of-speech (POS) tag. In operation 1623 of performing embedding, the speech recognition module 1530 can perform embedding on the pre-processed text. In operation 1623, the speech recognition module 1530 can map the pre-processed text to a plurality of vectors.

[0208] In operation 1625 of applying the temporal expression extraction model, the speech recognition module 1530 can extract a temporal expression included in the text obtained from the user utterance input based on the result of the embedding. In operation 1627 of applying the time point prediction model, the speech recognition module 1530 can predict the degree to which the extracted temporal expression represents the past with respect to the current time point.

[0209] Operations 1625 and 1627 of applying the temporal expression extraction model and the time point prediction model will be described below in more detail with reference to Figure 17 and Figure 18

[0210] Figure 17 is a diagram illustrating an exemplary temporal expression extraction model (e.g., including various processing circuitry and / or executable program elements) according to an embodiment of the disclosure.

[0211] For example, the application of the temporal expression extraction model 1625 can include an AI model that processes input data as illustrated in FIG. 17. Figure 17

[0212] According to an embodiment of the disclosure, the speech recognition module 1530 can receive text obtained by converting an utterance input into an input sentence. The speech recognition module 1530 can embed the input sentence in units of words and / or characters. The speech recognition module 1530 can perform cascaded embedding in order to use a word embedding result and a character embedding result together. The speech recognition module 1530 can generate a conditional random field (CRF) model by applying a bidirectional LSTM model to the text mapped to a plurality of vectors.

[0213] ​​The speech recognition module 1530 can generate a CRF by applying a probability-based tagging model under a predetermined condition. According to an embodiment of the disclosure, the speech recognition module 1530 can pre-learn a condition in which a word, a stem, or a morpheme is more likely to be a time expression and a tagged part having a probability value greater than or equal to a threshold value based on the pre-learned condition. For example, the speech recognition module 1530 can extract a time expression through a BIO tagging.

[0214] Figure 18 FIG. 18 is a graph illustrating an exemplary time point prediction model according to an embodiment of the disclosure.

[0215] For example, applying the time point prediction model 1627 can include determining a time point related to a time expression based on the time expression identified in the text by the time expression extraction model according to an embodiment of the disclosure. Figure 18 The AI model processes input data as illustrated.

[0216] According to an embodiment of the disclosure, the speech recognition module 1530 can determine a time point related to a time expression based on the time expression identified in the text by the time expression extraction model. The speech recognition module 1530 can predict a probability value regarding which time points are represented by the identified time expression, and determine a time point having the highest probability value as a time point represented by the time expression.

[0217] The speech recognition module 1530 can pre-learn time points indicated by various time expressions. The speech recognition module 1530 can derive a graph 1810 including a probability value (e.g., a probability that the identified time expression represents a plurality of time points) by applying a pre-trained model to the identified time expression as input. In the graph 1810, the x-axis and the y-axis can represent a time point and a probability value, respectively. The graph 1810 can represent a probability value regarding an arbitrarily designated time point at a specific time interval, or can display a probability value at a time point at which a past utterance is made or at a time point at which a conversational service is used.

[0218] When the time point related to the time expression is determined, the speech recognition module 1530 can perform a binary classification 1820 for determining whether the determined time point is before or after a preset time point.

[0219] The speech recognition module 1530 can select the first database 261 when it is determined based on a determination result via the binary classification 1820 that the time point related to the time expression is before the preset time point. The speech recognition module 1530 can select the second database 263 when it is determined based on a determination result via the binary classification 1820 that the time point related to the time expression is after the preset time point.

[0220] According to an embodiment of the disclosure, the speech recognition module 1530 can interpret the text based on a user conversation history acquired from the selected database. Although Figure 16The NLU module 1620 of the speech recognition module 1530, which is only illustrated, determines a time point in the text related to the utterance input and selects a database based on the time point, but the speech recognition module 1530 can again perform an NLU process to interpret the text based on information about the user's dialog history acquired from the selected database. The NLU module 1620 can disambiguate at least one entity included in the text and interpret the disambiguated text based on information about the user's dialog history acquired from the selected database.

[0221] The DM module 1630 can receive, as input, a result of interpreting the disambiguated text via the NLU module 1620 and output a list of instructions for the NLG module 1640, taking into account state variables such as a dialog history. The NLG module 1640 can generate a response message to the user's utterance input based on the received list of instructions.

[0222] According to various embodiments of the present disclosure, the electronic device 200 can use AI technology in the entire process for providing a dialog service to a user. AI-related functions according to the present disclosure are operated by a processor and a memory. The processor can include one or more processors. In this case, the one or more processors can be a general-purpose processor such as, for example, and without limitation, a central processing unit (CPU), an application processor (AP), or a digital signal processor (DSP); a special-purpose graphic processor such as, for example, a graphic processing unit (GPU) or a visual processing unit (VPU); or a special-purpose AI processor such as, for example, a neural processing unit (NPU). The one or more processors can control input data to be processed according to a predetermined operation rule or an AI model stored in the memory. When the one or more processors are a special-purpose AI processor, the special-purpose AI processor can be designed to have a hardware structure specialized for processing a specific AI model.

[0223] The predetermined operation rule or the AI model can be created through a training process. For example, this can refer to a predetermined operation rule or an AI model designed to perform a desired characteristic (or purpose) created by training a basic AI model based on a learning algorithm utilizing a large amount of training data. The training process can be performed by a device on which AI is implemented according to an embodiment of the present disclosure or a separate server and / or system. Examples of the learning algorithm can include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.

[0224] The AI model can include a plurality of neural network layers. Each of the neural network layers can have a plurality of weight values, and can perform a neural network calculation via an arithmetic operation on a result of a calculation in a previous layer and the plurality of weight values in the current layer. The plurality of weights in each of the neural network layers can be optimized by a result of training the AI model. For example, the plurality of weights can be updated to reduce or minimize a loss or cost value acquired by the AI model during a training process. The artificial neural network can include a deep neural network (DNN), and can include, for example, and without limitation, a convolutional neural network (CNN), a DNN, a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent DNN (BRDNN), a deep Q-network (DQN), etc., but is not limited thereto.

[0225] Embodiments of the present disclosure can be implemented as a software program including instructions stored in a computer-readable storage medium.

[0226] A computer can refer to, for example, a device configured to retrieve instructions stored in a computer-readable storage medium and operate in response to the retrieved instructions, and can include a terminal device and a remote control device according to embodiments of the present disclosure.

[0227] A computer-readable storage medium can be provided in the form of a non-transitory storage medium. In this regard, the "non-transitory" storage medium can not include a signal and be tangible, and the term does not distinguish between data stored semi-permanently and data temporarily stored in the storage medium.

[0228] In addition, the electronic device and method according to embodiments of the present disclosure can be provided in the form of a computer program product. The computer program product can be traded between a seller and a buyer as a product.

[0229] The computer program product can include a software program and a computer-readable storage medium having the software program stored therein. For example, the computer program product can include a product (e.g., a downloadable application) in the form of a software program that is electronically distributed by a manufacturer of an electronic device or through an electronic market (e.g., Google Play Store and App Store). For such electronic distribution, at least a portion of the software program can be stored on a storage medium or can be temporarily generated. The storage medium can be a storage medium of a server of the manufacturer, a server of the electronic market, or a relay server for temporarily storing the software program.

[0230] In a system including a server and a terminal (e.g., a terminal device or a remote control device), the computer program product can include a storage medium of the server or a storage medium of the terminal. In the presence of a third device (e.g., a smartphone) that communicates with the server or the terminal, the computer program product can include a storage medium of the third device. The computer program product can include a software program that is transmitted from the server to the terminal or the third device or from the third device to the terminal.

[0231] In this case, one of the server, the terminal, and the third device can execute the computer program product, thereby performing the method according to an embodiment of the disclosure. At least two of the server, the terminal, and the third device can execute the computer program product, thereby performing the method according to an embodiment of the disclosure in a distributed manner.

[0232] For example, the server (e.g., a cloud server, an AI server, etc.) can execute the computer program product stored in the server, and can control the terminal that communicates with the server to perform the method according to an embodiment of the disclosure.

[0233] As another example, the third device can execute the computer program product, and can control the terminal that communicates with the third device to perform the method according to an embodiment of the disclosure. As a specific example, the third device can remotely control a terminal device or a remote control device to transmit or receive a packed image.

[0234] In the case where the third device executes the computer program product, the third device can download the computer program product from the server, and can execute the downloaded computer program product. The third device can execute the computer program product that is preloaded therein, and can perform the method according to an embodiment of the disclosure.

[0235] While the disclosure has been illustrated and described with reference to various exemplary embodiments, it should be understood that the various exemplary embodiments are intended to be illustrative, not restrictive. Those of ordinary skill in the art will understand that various changes in form and detail can be made without departing from the scope of the disclosure, including the appended claims and their equivalents.

Claims

1. A method of providing a conversation service, performed by an electronic device, the method comprising: receiving an utterance input; identifying a time expression representing a time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from among a plurality of databases storing information related to a conversation history of a user using the conversation service; interpreting the text based on information related to the conversation history of the user obtained from the selected database; generating a response message to the utterance input based on a result of the interpreting; and outputting the generated response message. The identifying the time expression comprises:

2. The method of claim 1, wherein, obtaining the text by performing speech recognition on the utterance input; and determining an entity representing at least one of a time point, a duration, or a time period included in the text as the time expression. The determining the entity comprises:

3. The method of claim 2, wherein, performing embedding on the text to map the text to a plurality of vectors; assigning a BIO tag to at least one morpheme representing at least one of the time point, the duration, or the time period included in the text by applying a bidirectional long short-term memory (LSTM) model to the plurality of vectors; and identifying the entity in the text based on the BIO tag. The determining the time point related to the utterance input comprises:

4. The method of claim 1, wherein, predicting a probability value including a probability that the time expression will represent each of a plurality of time points; and determining a time point corresponding to a highest probability value among the predicted probability values as the time point related to the utterance input. The plurality of databases include a first database storing information related to a conversation history of the user accumulated before a preset time point and a second database storing information related to a conversation history of the user accumulated after the preset time point, and wherein the selecting the database comprises:

5. The method of claim 1, wherein, selecting the first database from among the plurality of databases based on the time point related to the utterance input being before the preset time point; and selecting the second database from among the plurality of databases based on the time point related to the utterance input being after the preset time point. The first database is stored in an external server and the second database is stored in the electronic device, and 6. The method of claim 5, wherein, wherein the preset time point includes at least one of a time point based on at least some of information related to the conversation history of the user included in the second database being transmitted to the first database, a time point based on a face image of the user being obtained, and a time point based on the conversation service being initiated. The interpreting the text comprises:

7. The method of claim 1, wherein, determining an entity included in the text that needs elaboration; obtaining elaboration information for elaborating the determined entity by retrieving information related to the conversation history of the user obtained from the selected database; and interpreting the text and the elaboration information using a natural language understanding (NLU) model. ​ 8. The method of claim 1, wherein, generating the response message includes: determining a type of the response message by applying a dialog manager (DM) model to a result of the interpreting; and generating the response message of the determined type using a natural language generation (NLG) model. 9.The method of claim 1, further comprising: obtaining a face image of the user; determining whether a face ID corresponding to the obtained face image is stored by searching a first database included in the plurality of databases; and initiating the dialog service based on a result of the determining.

10. The method of claim 9, wherein, initiating the dialog service includes: updating a stored service usage history mapped to the face ID based on the face ID corresponding to the obtained face image being stored in the first database; and generating a new face ID and a service usage history mapped to the new face ID based on the face ID corresponding to the obtained face image not being stored in the first database. 11.The method of claim 9, further comprising: transmitting the face ID to another electronic device after the dialog service ends; and transmitting information related to a dialog history of the user, which is stored in a second database included in the plurality of databases, to the other electronic device in response to a request received from the other electronic device. 12.An electronic device configured to provide a dialog service, the electronic device comprising: a memory storing one or more instructions; and at least one processor configured to execute the one or more instructions to provide the dialog service to the user, wherein the at least one processor is further configured to execute the one or more instructions to control the electronic device to: receive an utterance input; identify a time expression representing a time in text obtained from the utterance input; determine a time point related to the utterance input based on the time expression; select a database corresponding to the determined time point from among a plurality of databases storing information related to dialog histories of users using the dialog service; interpret the text based on information related to a dialog history of the user obtained from the selected database; generate a response message to the utterance input based on a result of the interpreting; and output the generated response message. 13.The electronic device of claim 12, further comprising: a camera configured to obtain a face image of the user; and a microphone configured to receive the utterance input, wherein, before initiating the dialog service, the at least one processor is further configured to execute the one or more instructions to control the electronic device to: determine whether a face ID corresponding to the obtained face image is stored by searching a first database included in the plurality of databases; update a stored service usage history mapped to the face ID based on the face ID corresponding to the obtained face image being stored in the first database; and generate a new face ID and a service usage history mapped to the new face ID based on the face ID corresponding to the obtained face image not being stored in the first database. ​ ​ ​ 14. The electronic device of claim 13, further comprising: a communication interface including communication circuitry configured to: transmit the face ID to another electronic device after the conversation service ends, and in response to a request received from the other electronic device, transmit information about a conversation history of the user, which is stored in a second database included in the plurality of databases, to the other electronic device. 15.A computer-readable recording medium, wherein a program is stored, which, when executed, causes an electronic device to perform operations for providing a conversation service, the operations comprising: receiving an utterance input; identifying a time expression representing a time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from among a plurality of databases storing information about a conversation history of a user using the conversation service; interpreting the text based on information about the conversation history of the user obtained from the selected database; generating a response message to the utterance input based on a result of the interpretation; and outputting the generated response message. ​