Electronic device and method for providing a conversational service
By using facial recognition and dialogue history management, the problem of providing personalized dialogue services in existing technologies has been solved, achieving continuity and accuracy of dialogue between devices and reducing the inconvenience of user registration.
Patent Information
- Application Number
- CN202080062807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-26
- Filing Date
- 2020-03-02
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2040-03-02
AI Technical Summary
Existing technologies in devices that provide dialogue services to multiple users cannot provide accurate personalized responses when users utter words related to their past conversation history.
By automatically recognizing users' facial IDs, storing and managing their conversation history, and using this history to generate response messages, the continuity and personalization of conversations can be achieved.
It provides personalized conversational services without requiring users to register an account, ensuring the continuity and accuracy of the conversation and reducing user inconvenience.
Smart Images

Figure CN114391143B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to electronic devices and methods for providing dialogue services, and, for example, to methods and devices for interpreting user utterance input and outputting response messages based on a user's dialogue history. Background Technology
[0002] With the recent development of electronic devices (such as smartphones) that perform various functions in complex ways, electronic devices equipped with voice recognition capabilities have emerged to improve operability. Voice recognition technology can be applied to conversational user interfaces to output response messages to questions input by the user in everyday natural language, providing a user-friendly conversational service.
[0003] A conversational user interface (QA) is an intelligent user interface that operates when the user is speaking in their own language. QA systems can be used to output answers to user questions. Unlike information retrieval technologies that simply retrieve and present lists of information related to user questions, QA systems search for and provide answers to user questions.
[0004] For example, personal electronic devices such as smartphones, computers, personal digital assistants (PDAs), portable multimedia players (PMPs), smart home appliances, navigation devices, and wearable devices can provide conversational services by connecting to servers or running applications.
[0005] As another example, public electronic devices installed in shops or public institutions, such as unmanned information terminals, unmanned kiosks, and unmanned checkout counters, can also provide conversational services. Public electronic devices installed in public places need to store and use the conversation history for each user in order to accurately analyze the user's verbal input and provide them with personalized responses. Summary of the Invention
[0006] Technical issues
[0007] When using devices that provide conversational services to multiple users, there is a need for a method that can receive accurate, personalized responses from the device even when a user utters utterances related to their past conversation history.
[0008] Technical solution
[0009] Embodiments of this disclosure provide a method and apparatus for providing conversation services by performing a process of retrieving stored conversation history associated with a user account in a more user-friendly manner.
[0010] Other aspects will be set forth in part in the following description, and will be apparent in part from the description.
[0011] According to an exemplary embodiment of this disclosure, a method for providing a dialogue service performed by an electronic device includes: receiving utterance input; identifying a time expression representing time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from a plurality of databases, wherein the plurality of databases store dialogue history information of a user using the dialogue service; interpreting the text based on the user's dialogue history information obtained from the selected database; generating a response message to the utterance input based on the interpretation result; and outputting the generated response message. Attached Figure Description
[0012] The above and other aspects, features, and advantages of certain embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0013] Figure 1 This is a diagram illustrating an example of an electronic device providing dialogue services based on dialogue history according to an embodiment of the present disclosure;
[0014] Figure 2a This is a diagram illustrating an exemplary system for providing a dialogue service according to embodiments of the present disclosure;
[0015] Figure 2b This is a diagram illustrating an exemplary system for providing a dialogue service according to embodiments of the present disclosure;
[0016] Figure 3 This is a flowchart illustrating an exemplary method for providing a dialogue service performed by an electronic device according to an embodiment of the present disclosure;
[0017] Figure 4a This is a diagram illustrating an example of an electronic device providing dialogue services based on dialogue history according to an embodiment of the present disclosure;
[0018] Figure 4b This is a diagram illustrating an example of an electronic device providing dialogue services based on dialogue history according to an embodiment of the present disclosure;
[0019] Figure 5 This is a diagram illustrating an exemplary process performed by an electronic device to provide a dialogue service according to an embodiment of the present disclosure;
[0020] Figure 6 This is a diagram illustrating an example of stored conversation history information according to an embodiment of this disclosure;
[0021] Figure 7a This is a flowchart illustrating an exemplary method performed by an electronic device to provide a dialogue service according to an embodiment of the present disclosure;
[0022] Figure 7bThis is a flowchart illustrating an exemplary method performed by an electronic device to provide a dialogue service according to an embodiment of the present disclosure;
[0023] Figure 8 This is a flowchart illustrating an exemplary method for determining whether an electronic device will use conversation history information to generate a response, according to embodiments of the present disclosure;
[0024] Figure 9 This is a flowchart illustrating an exemplary method performed by an electronic device, based on user verbal input, according to an embodiment of the present disclosure;
[0025] Figure 10 This is an exemplary probability map of time points related to user speech input determined by an electronic device according to embodiments of this disclosure;
[0026] Figure 11 This is a diagram illustrating an exemplary method performed by an electronic device, according to an embodiment of the present disclosure, for switching a database in which a user's conversation history is stored;
[0027] Figure 12a This is a diagram illustrating an exemplary process in which multiple electronic devices share a user's conversation history with each other, according to embodiments of this disclosure;
[0028] Figure 12b This is a diagram illustrating an exemplary process in which multiple electronic devices share a user's conversation history with each other, according to embodiments of this disclosure;
[0029] Figure 12c This is a diagram illustrating an exemplary process in which multiple electronic devices share a user's conversation history with each other, according to embodiments of this disclosure;
[0030] Figure 12d This is a diagram illustrating an exemplary process in which multiple electronic devices share a user's conversation history with each other, according to embodiments of this disclosure;
[0031] Figure 13a This is a block diagram illustrating an exemplary configuration of an exemplary electronic device according to embodiments of the present disclosure;
[0032] Figure 13b This is a block diagram illustrating an exemplary configuration of an exemplary electronic device according to another embodiment of the present disclosure;
[0033] Figure 14 This is a block diagram illustrating an exemplary electronic device according to an embodiment of the present disclosure;
[0034] Figure 15a This is a block diagram illustrating an exemplary processor included in an exemplary electronic device according to an embodiment of the present disclosure;
[0035] Figure 15b This is a block diagram illustrating an exemplary processor included in an exemplary electronic device according to an embodiment of the present disclosure;
[0036] Figure 16 This is a block diagram illustrating an exemplary speech recognition module according to an embodiment of the present disclosure;
[0037] Figure 17 This is a diagram illustrating an exemplary time expression extraction model according to embodiments of the present disclosure; and
[0038] Figure 18 This is a diagram illustrating an exemplary time-point prediction model according to an embodiment of the present disclosure. Detailed Implementation
[0039] According to an exemplary embodiment of this disclosure, a method for providing a dialogue service performed by an electronic device includes: receiving utterance input; identifying a time expression representing time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from a plurality of databases, the plurality of databases storing dialogue history information of users using the dialogue service; interpreting the text based on the user's dialogue history information obtained from the selected database; generating a response message to the utterance input based on the interpretation result; and outputting the generated response message.
[0040] According to another exemplary embodiment of this disclosure, an electronic device configured to provide a dialogue service includes: a memory storing one or more instructions; and at least one processor configured to execute one or more instructions to provide a dialogue service to a user, wherein the at least one processor is further configured to execute one or more instructions to control the electronic device to: receive speech input; identify a time expression representing time in text obtained from the speech input; determine a time point related to the speech input based on the time expression; select a database corresponding to the determined time point from a plurality of databases, the plurality of databases storing information about the dialogue history of a user using the dialogue service; interpret the text based on the user's dialogue history information obtained from the selected database; generate a response message to the speech input based on the interpretation result; and output the generated response message.
[0041] According to another exemplary embodiment of this disclosure, one or more non-transitory computer-readable recording media store a program for performing a method of providing a dialogue service, the method comprising: receiving utterance input; identifying a time expression representing time in text obtained from the utterance input; determining a time point related to the utterance input based on the time expression; selecting a database corresponding to the determined time point from a plurality of databases, the plurality of databases storing dialogue history information of users using the dialogue service; interpreting the text based on the user's dialogue history information obtained from the selected database; generating a response message to the utterance input based on the interpretation result; and outputting the generated response message.
[0042] Various exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. However, embodiments of the present disclosure may take different forms and should not be construed as limited to the various exemplary embodiments described herein. Furthermore, parts unrelated to the present disclosure may be omitted to make the description clear, and the same reference numerals in the drawings always denote the same elements.
[0043] Throughout the disclosure, the expression "at least one of a, b, or c" means only a, only b, only c, both a and b, both a and c, both b and c, all of a, b, and c, or variations thereof.
[0044] Some embodiments of this disclosure can be described based on functional block components and various processing steps. All or some of the functional blocks can be implemented using any number of hardware and / or software components configured to perform a specific function. For example, the functional blocks of this disclosure can be implemented by one or more microprocessors or circuit components for performing a predetermined function. Furthermore, for example, the functional blocks of this disclosure can be implemented using various programming or scripting languages. The functional blocks can be implemented in algorithms running on one or more processors. Additionally, this disclosure can employ techniques from related fields for electronic configuration, signal processing, and / or data processing.
[0045] Furthermore, the connecting lines or connectors between the components shown in the figures are intended only to illustrate exemplary functional relationships and / or physical or logical connections between the components. It should be noted that many alternative or additional functional relationships, physical connections, or logical connections may exist in actual devices.
[0046] Various exemplary embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings.
[0047] Typically, devices used to provide conversational services to multiple users can manage conversations based on each user account or on each session without identifying the user. A session can refer to, for example, the time period from the start to the end of a conversational service that performs speech recognition on a user's query and outputs a response to it.
[0048] Typically, in scenarios where the device manages conversations based on each user account, the user needs to register his or her account by entering user information such as a user ID or name. Furthermore, each time a user uses the device, they may suffer the inconvenience of having to enter user information and retrieve their account.
[0049] To reduce user inconvenience, devices that manage conversations on a per-session basis can be used without recognizing the user. In this scenario, when a device manages conversations on a per-session basis, it starts a new session and provides conversation services upon receiving user utterance input after a previous session has ended, without providing information about the history of the previous conversation. Therefore, when a user asks questions related to the details of their conversation with the device during a previous session after a new session has started, the device may not be able to provide accurate responses to the user's questions.
[0050] To address the aforementioned problems, this disclosure provides a method executed by an electronic device that stores each user's conversation history by automatically identifying the user without requiring the user to register his or her account, and utilizes that conversation history to create responses.
[0051] Figure 1 This diagram illustrates an example of a user 10 visiting a store and purchasing an air conditioner, a television (TV), and a refrigerator via electronic device 200 a month ago. According to embodiments of this disclosure, when user 10 simply approaches electronic device 200, electronic device 200 can automatically perform facial-based user authentication without requiring the input of user information for authentication. Electronic device 200 can initiate a conversational service via facial-based user authentication without a separate command from the user.
[0052] According to embodiments of this disclosure, electronic device 200 can detect that user 10's facial ID has been stored by recognizing the user's face, and determine that user 10 has used the chat service. Electronic device 200 can output the voice message "You're back again" based on the user's service usage history.
[0053] According to embodiments of this disclosure, when no facial ID matching the user's face is found, the electronic device 200 can determine that the user 10 has never used the chat service before. When it is determined that the user 10 is using the chat service for the first time, the electronic device 200 can output the voice message "Welcome to our first visit". According to embodiments of this disclosure, the electronic device 200 can output context-appropriate user response messages based on the user's service usage history.
[0054] like Figure 1 As shown, user 10 can send a request related to a product that user 10 previously (e.g., a month ago) purchased through electronic device 200: “Do you know what I bought last time?”
[0055] Because electronic devices that manage conversations based on each session do not store the content of past conversations with user 10, the electronic device may not be able to ensure the continuity of the conversation or provide appropriate response messages even when user 10 is making utterances related to the content of his or her past utterances.
[0056] On the other hand, the electronic device 200 according to an exemplary embodiment of the present disclosure can interpret the user 10's utterances based on the user 10's past dialogue history, thereby outputting a response message based on the past dialogue history.
[0057] For example, such as Figure 1 As shown, electronic device 200 can output a response message “You purchased an air conditioner, a television, and a refrigerator a month ago. Which product are you referring to?” based on the conversation history related to products purchased by user 10 through electronic device 200 a month ago. According to embodiments of this disclosure, when it is determined that past conversation history of the user is needed to interpret the user's utterance, electronic device 200 can interpret the user's utterance based on the past conversation history of the user matched with the user's face identified for storage, and generate a response message based on the interpretation result.
[0058] Figure 2a This is a diagram illustrating an exemplary system for providing a dialogue service according to embodiments of the present disclosure, and Figure 2b This is a diagram illustrating an exemplary system for providing a dialogue service according to embodiments of the present disclosure.
[0059] like Figure 2a As shown, according to embodiments of this disclosure, electronic device 200 can be used independently to provide conversational services to user 10. Examples of electronic device 200 may include, but are not limited to: home appliances (e.g., televisions, refrigerators, washing machines, etc.), smartphones, PCs, wearable devices, PDAs, media players, microservers, Global Positioning System (GPS), e-book terminals, digital broadcasting terminals, navigation devices, kiosks, MP3 players, digital cameras, and other mobile or non-mobile computing devices. Electronic device 200 can provide conversational services by, for example, executing chatbot applications or conversational proxy applications.
[0060] According to embodiments of this disclosure, electronic device 200 can receive verbal input from user 10 and generate and output a response message to the received verbal input.
[0061] According to embodiments of this disclosure, electronic device 200 may provide a method for storing conversation history for each user and using the conversation history to generate responses by automatically identifying user 10 without requiring user 10 to register an account.
[0062] For example, according to embodiments of this disclosure, electronic device 200 can recognize a user's face via a camera and search for a face ID matching the recognized face from a stored collection of face IDs. Electronic device 200 can retrieve the dialogue history and dialogue service usage history mapped to the found face ID. Electronic device 200 can provide dialogue services to user 10 based on the dialogue history and update the dialogue history after ending the dialogue with user 10.
[0063] According to embodiments of this disclosure, when no matching face ID is found in the stored face IDs, the electronic device 200 may check with the user 10 to see if information related to the user's face can be stored. When the user 10 agrees to store his or her face ID, the electronic device 200 may map conversation history and service usage history to his or her face ID for storage after the conversation with the user 10 ends.
[0064] According to embodiments of this disclosure, when managing stored facial IDs and conversation history, electronic device 200 can specify a maximum storage period for storing facial IDs and conversation history based on the fact that storage capacity is limited and personal information needs to be protected. When the maximum storage period expires, electronic device 200 can delete the stored facial IDs and conversation history. However, when user 10 is re-identified before the maximum storage period expires (e.g., when user 10 revisits a store), electronic device 200 can extend and flexibly manage the storage period. Depending on the time interval and frequency of user 10's use of the conversation service, electronic device 200 can specify different information storage periods for each user.
[0065] According to embodiments of this disclosure, when it is determined that past conversation history and service usage history are needed to explain a user's problem, electronic device 200 can generate and output a response message based on the user 10's conversation history.
[0066] Electronic device 200 can receive user utterance input and determine the context related to the time included in the utterance input. The context related to the time included in the utterance input may, for example, involve information about the time required to generate a response based on the user's intent included in the utterance input. Based on the context determination, electronic device 200 can determine which dialogue history information to use from information accumulated in a first time period and information accumulated in a second time period. Therefore, electronic device 200 can reduce the amount of time required to interpret the user's question and provide an appropriate response by identifying the context of the dialogue solely based on the selected dialogue history information.
[0067] In addition, such as Figure 2bAs shown, according to embodiments of this disclosure, electronic device 200 can provide dialogue services in conjunction with another electronic device 300 and / or server 201. Electronic device 200, another electronic device 300, and server 201 can be connected to each other via wired or wireless means.
[0068] Another electronic device 300 or server 201 can share data, resources and services with electronic device 200, perform control or file management of electronic device 200, or monitor the entire network. For example, the other electronic device 300 can be a mobile or non-mobile computing device.
[0069] Electronic device 200 can generate and output response messages to user speech input through communication with another electronic device 300 and / or server 201.
[0070] like Figure 2a and Figure 2b As shown, according to embodiments of this disclosure, a system for providing a dialogue service may include at least one electronic device and / or a server. For ease of description, a method for providing a dialogue service performed by an "electronic device" will be described below. However, some or all of the operations of the electronic device described below may also be performed by another electronic device and / or a server connected to that electronic device, and may be performed in part by multiple electronic devices.
[0071] Figure 3 This is a flowchart illustrating an exemplary method of providing a dialogue service performed by an electronic device 200 according to an embodiment of the present disclosure.
[0072] According to embodiments of this disclosure, electronic device 200 can receive user voice input (operation S310).
[0073] According to embodiments of this disclosure, electronic device 200 can initiate a dialogue service and receive user speech input. According to embodiments of this disclosure, when a user approaches electronic device 200 at a certain distance, and when a voice signal of predetermined or higher intensity is received, and when a voice signal for issuing a predetermined activation word is received, electronic device 200 can initiate a dialogue service.
[0074] According to embodiments of this disclosure, when a user approaches the electronic device 200 at a certain distance, the electronic device 200 can obtain a facial image of the user and determine, by searching a database, whether a facial ID matching the obtained facial image of the user is stored. It should be understood that any suitable device for identifying a user can be used, and this disclosure is not limited to facial ID recognition. For example, but not limited to, voice recognition, biometrics, user interface input, etc., can be used, and facial IDs are used for ease of explanation and illustrative purposes. The electronic device 200 can initiate a dialogue service based on the determination result.
[0075] According to embodiments of this disclosure, when a facial ID matching the obtained user's facial image is stored in a database, the electronic device 200 can update the stored service usage history mapped to the facial ID. Otherwise, when a facial ID matching the obtained user's facial image is not stored in the database, the electronic device 200 can generate a new facial ID and a service usage history mapped to the new facial ID.
[0076] According to embodiments of this disclosure, electronic device 200 can receive and store audio signals including user speech input via a microphone. For example, electronic device 200 can detect the presence or absence of human speech by using speech activity detection (VAD), endpoint detection (EPD), etc., thereby receiving and storing speech input on a sentence-by-sentence basis.
[0077] According to embodiments of this disclosure, electronic device 200 can recognize a time expression representing time obtained from user speech input (e.g., text) (operation S320).
[0078] Electronic device 200 can obtain text by performing speech recognition on user speech input and identify entities in the text that represent at least one point in time, duration, or time period as time expressions. However, it should be understood that identifying time expressions is not limited to obtaining and analyzing text.
[0079] An entity may include at least one of a word, phrase, or morpheme that has a specific meaning and is contained in the text. Electronic device 200 can identify at least one entity in the text and determine which field includes each entity based on the meaning of that at least one entity. For example, electronic device 200 can determine whether an entity identified in the text is an entity representing, but not limited to, a person, an object, a geographical region, time, date, etc.
[0080] According to embodiments of this disclosure, electronic device 200 can determine time expressions, such as, but not limited to, time points indicating operations or states indicated by text, or adverbs, adjectives, nouns, words, phrases, etc., included in text indicating time points, times, time periods, etc.
[0081] Electronic device 200 can perform embeddings to map text to multiple vectors. By applying a bidirectional long short-term memory (LSTM) model to the mapped vectors, electronic device 200 can assign a beginning-inside-outside (BIO) label to at least one morpheme included in the text representing at least one point in time, duration, or time period. Electronic device 200 can identify entities representing time in the text based on the BIO labels. Electronic device 200 can determine the identified entities as time expressions.
[0082] According to embodiments of this disclosure, electronic device 200 can determine the time point related to user speech input based on a time expression (operation S330).
[0083] The point in time associated with user utterance input can be, for example, the point in time when the electronic device 200 generates the information necessary for a response based on the intent in the user's utterance, the point in time when it receives past utterance input from the user including that information, or the point in time when it stores that information in the conversation history. The point in time associated with user utterance input can include the point in time when past utterances by the user are received or stored, which include information detailing the content of the user's utterance input. For example, the point in time associated with user utterance input can include the point in time when past utterances are received or stored that relate to the user purchasing, mentioning, or querying a product mentioned in the user's utterance input, or a service request for that product.
[0084] Electronic device 200 can predict probability values; for example, a time expression indicating the probability of each of a plurality of time points. Electronic device 200 can determine the time point corresponding to the highest probability value among the predicted probability values as the time point associated with the user's verbal input.
[0085] According to embodiments of this disclosure, electronic device 200 can select a database corresponding to a time point related to the user's utterance input from a plurality of databases storing information related to the dialogue history of users using the dialogue service (operation S340). Alternatively, a single database can be used, and this disclosure is not limited to multiple databases.
[0086] Multiple databases may include a first database for storing information about user dialogue history accumulated before a preset time point and a second database for storing information about user dialogue history accumulated after the preset time point. When the time point associated with the user's speech input is before the preset time point, the electronic device 200 may select the first database. When the time point associated with the user's speech input is after the preset time point, the electronic device 200 may select the second database. Furthermore, the first database may be stored on an external server, while the second database may be stored within the electronic device 200.
[0087] According to embodiments of this disclosure, the preset time point can be one of the following: the time point when at least some of the information related to the user's conversation history included in the second database is sent to the first database, the time point when the user's facial image is obtained, and the time point when the conversation service starts.
[0088] According to embodiments of this disclosure, electronic device 200 can interpret text based on information about a user's conversation history obtained from a selected database (operation S350).
[0089] Electronic device 200 can identify entities included in text that require detailed description. Electronic device 200 can obtain detailed description information for the identified entities by retrieving information about the user's dialogue history from a selected database. Electronic device 200 can use, for example, but not limited to, Natural Language Understanding (NLU) models to interpret the text and detailed description information.
[0090] Information about a user's conversation history retrieved from a database may include, for example, but not limited to, past verbal input received from the user, past response messages provided to the user, and information related to past verbal input and past response messages. For example, information related to past verbal input and past response messages may include entities included in the past verbal input, the content included in the past verbal input, the category of the past verbal input, the time point when the past verbal input was received, entities included in the past response messages, the content included in the past response messages, the time point when the past response messages were output, information related to the situation before and after the past verbal input was received, and the user's interested products, mood, payment information, voice characteristics, etc., which are determined based on past verbal input and past response messages.
[0091] According to embodiments of this disclosure, electronic device 200 can generate a response message to the received user speech input based on the interpretation result (operation S360).
[0092] Electronic device 200 may determine the type of response message, for example, but not limited to, by applying a dialogue manager (DM) model to the interpretation results. Electronic device 200 may use, for example, but not limited to, a natural language generation (NLG) model to generate the determined type of response message.
[0093] According to embodiments of this disclosure, electronic device 200 can output the generated response message (operation S370). For example, electronic device 200 can output the response message in the form of at least one of voice, text, or image.
[0094] According to embodiments of this disclosure, electronic device 200 may share at least one of a user's facial ID, service usage history, or conversation history with another electronic device. For example, after a conversation service provided to a user ends, electronic device 200 may send the user's facial ID to another electronic device. When a user wishes to receive conversation services through another electronic device, the other electronic device may identify the user and, based on determining that the identified user corresponds to the facial ID of the receiving user, request information related to the user's conversation history from electronic device 200. In response to the request received from the other electronic device, electronic device 200 may send the information related to the user's conversation history stored in a second database included in the database to the other electronic device.
[0095] Figure 4a This is a diagram illustrating an example of an electronic device 200 providing dialogue services based on dialogue history according to an embodiment of the present disclosure. Figure 4b This is a diagram illustrating an example of an electronic device 200 providing dialogue services based on dialogue history according to an embodiment of the present disclosure.
[0096] Figure 4a This is a diagram illustrating an example of an electronic device 200 according to an embodiment of the present disclosure, which is an unmanned vending kiosk in a store selling electronic products. A vending kiosk can, for example, refer to an unmanned information terminal installed in a public place. (See also...) Figure 4a On May 5, 2019, electronic device 200 can receive voice input from user 10, who inquires about the price of air conditioner A. Electronic device 200 can respond to the user's voice input by outputting a response message informing the user of the price of air conditioner A.
[0097] On May 15, 2019, electronic device 200 can receive speech input from user 10 who is revisiting the store. Electronic device 200 can obtain the text saying "I asked you how much your air conditioner cost last time?" through speech recognition of the user's speech input. Electronic device 200 can recognize temporal expressions in the obtained text (e.g., "last time," "asked," and "how much the price"). Electronic device 200 can determine the user's dialogue history needed to interpret the obtained text based on the recognized temporal expressions.
[0098] Electronic device 200 can identify entities in text that require detailed explanation and obtain detailed explanation information for those entities based on the user's dialogue history. Electronic device 200 can use an NLU model to interpret the text that details the entities.
[0099] According to embodiments of this disclosure, electronic device 200 can identify entities in text representing product categories and describe the product in detail based on conversation history information. For example, such as Figure 4a As shown, electronic device 200 can determine that "air conditioner" is the entity representing the product entity category in the user's utterance "I asked you last time how much your air conditioner cost?". Electronic device 200 can determine, based on the dialogue history dated May 5th, that the product the user wants to know the price of is "Air Conditioner A". Electronic device 200 can respond to the user's question by outputting a response message informing them of the price of Air Conditioner A.
[0100] Figure 4b This is a diagram illustrating an example of an electronic device 200, according to an embodiment of the present disclosure, which is a self-service checkout counter in a restaurant. (Refer to...) Figure 4b On May 10, 2019, electronic device 200 was able to receive verbal input from user 10 who ordered a salad. Electronic device 200 could respond to the user's verbal input by outputting a confirmation message regarding the salad order.
[0101] On May 15, 2019, electronic device 200 can receive speech input from user 10 who is revisiting the restaurant. Electronic device 200 can obtain text stating "order the one I usually eat" through speech recognition of the user's speech input. Electronic device 200 can recognize temporal expressions in the obtained text (e.g., "always" and "eat"). Electronic device 200 can determine the user's dialogue history needed to interpret the obtained text based on the recognized temporal expressions.
[0102] According to embodiments of this disclosure, electronic device 200 can identify nouns in text that require detailed explanation as entities requiring detailed explanation, and can provide detailed explanations of the objects indicated by the nouns based on dialogue history information. For example, such as Figure 4bAs shown, electronic device 200 can identify "that" in the user's utterance "order the one I usually eat" as an entity requiring further details. Electronic device 200 can determine, based on the conversation history dated May 10th, that the food the user 10 wishes to order is "salad." In response to the user's utterance input, electronic device 200 can output a response message requesting confirmation that the salad order is correct.
[0103] Furthermore, according to embodiments of this disclosure, in order to conduct a smooth dialogue with user 10, electronic device 200 may need to reduce the time required to generate response messages based on dialogue history. Therefore, according to embodiments of this disclosure, electronic device 200 can reduce the time required to retrieve dialogue history information by searching only a database selected from multiple databases storing information related to dialogue history.
[0104] Figure 5 This is a diagram illustrating an exemplary process performed by an electronic device 200 to provide a dialogue service according to an embodiment of the present disclosure.
[0105] Figure 5 The electronic device 200 shown according to an embodiment of the present disclosure is an example of an unmanned vending machine in a store selling electronic products. According to an embodiment of the present disclosure, the electronic device 200 may use a first database 501 and a second database 502, the first database 501 for storing dialogue history information accumulated before the start of the current session for providing dialogue services, and the second database 502 for storing information related to dialogues performed during the current session.
[0106] Figure 6 This is a diagram illustrating an example of service usage history and conversation history stored in a first database 501 for each user, according to an embodiment of this disclosure. Figure 6 As shown, electronic device 200 can map service usage history 620 and conversation history 630 to user 10's facial ID for storage.
[0107] Service usage history 620 may include, for example, but not limited to, the number of visits made by user 10, the frequency of visits, the last visit date 621, and the date on which service usage history 620 was scheduled to be deleted. Conversation history 630 may include information about: products the user is interested in and the categories of past questions the user has asked (these are determined based on past questions the user has asked), whether user 10 has purchased any products of interest, and the time when the user's questions were received.
[0108] Electronic device 200 can identify user 10 by performing facial recognition on a facial image of user 10 obtained via a camera (operation S510). For example, electronic device 200 can retrieve the conversation service usage history corresponding to the facial ID of the identified user 10 by searching the first database 501.
[0109] Electronic device 200 can update the conversation service usage history of the identified user 10 by adding information indicating that user 10 is currently using the conversation service (operation S520).
[0110] For example, such as Figure 6 As shown, electronic device 200 can update user 10's conversation service usage history 620 in such a way as to add information 621 indicating that user 10 used the conversation service on May 16, 2019.
[0111] According to embodiments of this disclosure, electronic device 200 may postpone the date for deleting information related to user 10 based on at least one of the number of visits by user 10 or the visit cycle recorded in the user's usage history. For example, as the number of visits by user 10 increases and the visit cycle shortens, electronic device 200 may extend the storage period of information related to user 10.
[0112] Electronic device 200 can initiate a dialogue service and receive user speech input (operation S530). Electronic device 200 can obtain text from the user speech input stating "I have a problem with the product I purchased last time". Electronic device 200 can recognize "last time" as a time expression representing time in the obtained text.
[0113] Electronic device 200 can determine, based on the time expression "last time," the information needed to interpret the obtained text in relation to the user's conversation history accumulated before the start of the current session.
[0114] Electronic device 200 can retrieve information about user dialogue history from first database 501 based on determining that information related to user dialogue history accumulated before the start of the current session is needed (operation S540). Electronic device 200 can identify "product" as a noun included in the user's statement "I have a problem with the product I bought last time" as an entity that needs to be detailed, and interpret the text based on the information related to the user's dialogue history obtained from first database 501.
[0115] For example, electronic device 200 can determine, based on conversation history 631 dated May 10, that the product user 10 wants to reference is “Computer B” with model name 19COMR1.
[0116] In response to user voice input, electronic device 200 can output a response message confirming whether user 10 visited the store due to a problem with computer B purchased on May 10 (operation S550).
[0117] After the conversation service ends, the electronic device 200 can update the information related to the user's conversation history in the first database 510 by adding information about the history of conversations performed during the session (operation S560).
[0118] Figure 7a This is a flowchart illustrating an exemplary method of providing a dialogue service performed by an electronic device 200 according to an embodiment of the present disclosure. Figure 7b This is a flowchart illustrating an exemplary method of providing a dialogue service performed by an electronic device 200 according to an embodiment of the present disclosure.
[0119] Users can begin using the conversation service by approaching the electronic device 200.
[0120] According to embodiments of this disclosure, electronic device 200 can recognize a user's face using a camera (operation S701). According to embodiments of this disclosure, electronic device 200 can search a database to find a stored face ID corresponding to the recognized user (operation S702). According to embodiments of this disclosure, electronic device 200 can determine whether the face ID corresponding to the recognized user is stored in the database (operation S703).
[0121] When a user's facial ID is stored in the database ("Yes" in operation S703), the electronic device 200 can retrieve the user's service usage history. According to embodiments of this disclosure, the electronic device 200 can update the user's service usage history (operation S705). For example, the electronic device 200 can update information included in the user's service usage history related to the date of the user's last visit.
[0122] According to embodiments of this disclosure, when a user's facial ID is not stored in the database ("No" in operation S703), the electronic device 200 can ask the user whether he or she agrees to store the user's facial ID and conversation history in the future (operation S704). According to embodiments of this disclosure, when the user agrees to store his or her facial ID and conversation history ("Yes" in operation S704), the electronic device 200 can update the user's service usage history in operation S705. According to embodiments of this disclosure, when the user does not agree to store his or her facial ID and conversation history, the electronic device 200 can perform conversations with the user on a per-session basis.
[0123] Reference Figure 7bAccording to embodiments of this disclosure, electronic device 200 can receive user voice input (operation S710).
[0124] According to embodiments of this disclosure, electronic device 200 can determine whether dialogue history information is needed to interpret the user's utterance input (operation S721). Reference will be made below. Figure 8 The operation of S721 is described in more detail.
[0125] When it is determined that dialogue history information is not needed for interpretation (No in operation S721), according to embodiments of this disclosure, electronic device 200 may generate and output a generic response without using dialogue history information (operation S731).
[0126] When it is determined that dialogue history information is needed for interpretation ("Yes" in operation S721), according to an embodiment of this disclosure, the electronic device 200 may determine whether dialogue history information needs to be included in the first database (operation S723).
[0127] For example, according to embodiments of this disclosure, electronic device 200 can determine, based on a preset time point, whether user dialogue history information needs to be accumulated and stored in a first database before the preset time point, or whether user dialogue history information needs to be accumulated and stored in a second database after the preset time point. The first database can store dialogue history information accumulated over a relatively long period from the first time the user's dialogue history was stored at the preset time point. The second database can store dialogue history information accumulated over a short period from the preset time point to the current time point. Reference will be made below. Figure 9 The operation of S723 is described in more detail.
[0128] When it is determined that dialogue history information from the first database is needed to interpret user utterance input ("Yes" in operation S723), according to an embodiment of this disclosure, the electronic device 200 may generate a response message based on the dialogue history information obtained from the first database (operation S733). When it is determined that dialogue history information from the first database is not needed to interpret user utterance input ("No" in operation S723), according to an embodiment of this disclosure, the electronic device 200 may generate a response message based on the dialogue history information obtained from the second database (operation S735).
[0129] According to embodiments of this disclosure, electronic device 200 can output the generated response message (operation S740).
[0130] According to embodiments of this disclosure, electronic device 200 can determine whether a conversation has ended (operation S750). For example, electronic device 200 can determine that the conversation has ended when the user moves away from electronic device 200 at a distance greater than or equal to a threshold distance, when no user speech input is received for a threshold time, or when it is determined that the user has moved away from any space (e.g., a store or restaurant) where electronic device 200 is located.
[0131] When it is determined that the dialogue has ended ("Yes" in operation S750), according to an embodiment of this disclosure, the electronic device 200 may also store the dialogue history information related to the current dialogue in the existing dialogue history mapped to the user's face ID (operation S760). Otherwise, when it is determined that the dialogue has not ended ("No" in operation S750), according to an embodiment of this disclosure, the electronic device 200 may return to operation S710 and repeat the process of receiving user speech input and generating response messages to user speech input.
[0132] Figure 8 This is a flowchart illustrating an exemplary method for determining whether an electronic device 200 will generate a response based on dialogue history information, according to an embodiment of the present disclosure.
[0133] For example, Figure 7b The operation of S721 can be subdivided into Figure 8 Operations of S810, S820, and S830.
[0134] According to embodiments of this disclosure, electronic device 200 can receive user speech input (operation S710). According to embodiments of this disclosure, electronic device 200 can obtain text by performing speech recognition (e.g., automatic speech recognition (ASR)) on the received user speech input (operation S810).
[0135] According to embodiments of this disclosure, electronic device 200 can extract time expressions from acquired text (operation S820). According to embodiments of this disclosure, electronic device 200 can extract time expressions by applying a pre-trained time expression extraction model to the acquired text.
[0136] According to embodiments of this disclosure, electronic device 200 can determine whether the extracted time expression represents a past point in time, a time period, or a duration (operation S830). When the extracted time expression does not represent a past time expression ("No" in operation S830), electronic device 200 can generate a response to the user's utterance input based on general NLU without considering dialogue history (operation S841). Otherwise, when the extracted time expression represents a past time expression ("Yes" in operation S830), electronic device 200 can determine that a response needs to be generated based on dialogue history information (operation S843).
[0137] Figure 9 This is a flowchart illustrating an exemplary method of selecting a database based on user verbal input, performed by an electronic device 200 according to an embodiment of the present disclosure.
[0138] For example, it can be Figure 7b The operation of S723 is subdivided into Figure 9 Operations S910, S920, S930, and S940.
[0139] According to an embodiment of this disclosure, in operation S843, the electronic device 200 can determine that dialogue history information is needed to interpret user speech input.
[0140] According to embodiments of this disclosure, electronic device 200 can extract a time expression from text obtained based on user speech input (operation S910). According to embodiments of this disclosure, electronic device 200 can extract an expression representing the past from the extracted time expression (operation S920). Because Figure 9 Operation S910 corresponds to Figure 8 Since operation S820 is performed, operation S910 may not be performed according to embodiments of this disclosure. When not performed... Figure 9 During operation S910, electronic device 200 can use the time expression extracted and stored in operation S820.
[0141] According to embodiments of this disclosure, electronic device 200 can predict time points related to user utterance input based on time expressions representing the past (operation S930). According to embodiments of this disclosure, electronic device 200 can determine time points related to extracted past time expressions by applying a pre-trained time point prediction model to the extracted past time expressions.
[0142] like Figure 10As shown, electronic device 200 can predict probability values, for example, by representing the probability of each of multiple time points using a past time expression, and generate a graph 1000 representing the predicted probability values. Electronic device 200 can determine the time point corresponding to the highest probability value 1001 among the predicted probability values as the time point relevant to the user's verbal input. In graph 1000, the x-axis and y-axis can represent time and probability values, respectively. The zero point on the time axis in graph 1000 represents a preset time point used as a reference point for selecting a database.
[0143] According to embodiments of this disclosure, electronic device 200 can determine whether the predicted time point is before a preset time point (operation S940).
[0144] When the predicted time point is before the preset time point ("Yes" in operation S940), the electronic device 200 can generate a response to the user's utterance input based on the dialogue history information obtained from the first database (operation S733). When the predicted time point is at or after the preset time point ("No" in operation S940), the electronic device 200 can generate a response to the user's utterance input based on the dialogue history information obtained from the second database (operation S735).
[0145] According to embodiments of this disclosure, electronic device 200 can manage multiple databases based on the time period during which dialogue history is accumulated, thereby reducing the time required to retrieve dialogue history. According to embodiments of this disclosure, electronic device 200 can switch between databases, such that at least some information about a user's dialogue history stored in one database is stored in another database.
[0146] In this disclosure, although Figure 10 An example is shown in which the electronic device 200 uses two databases, but embodiments of this disclosure are not limited thereto. The databases used by the electronic device 200 may include three or more databases. For ease of description, the case in this disclosure where the databases include a first and a second database is described as an example.
[0147] Figure 11 This is a diagram illustrating an exemplary method performed by an electronic device 200 according to an embodiment of the present disclosure for switching a database in which a user's conversation history is stored.
[0148] According to embodiments of this disclosure, a first database 1101 may store information about user conversation history accumulated before a preset time point, and a second database may store information about user conversation history accumulated after the preset time point. For example, the preset time point may be one of the time when at least some information about user conversation history included in the second database 1102 is sent to the first database 1101, the time when a user's facial image is obtained, the time when the conversation service begins, or a time point that occurs at a predetermined time before the current time point; however, this disclosure is not limited to these.
[0149] According to embodiments of this disclosure, the first database 1101 can store dialogue history information accumulated over a relatively long period from when the user's dialogue history is first stored to a preset time point. The second database 1102 can store dialogue history information accumulated over a short period from the preset time point to the current time point.
[0150] For example, the first database 1101 may be included in an external server, while the second database 1102 may be included in the electronic device 200. The first database 1101 may also store the user's service usage history.
[0151] According to embodiments of this disclosure, electronic device 200 can switch at least some information about a user's conversation history stored in a second database 1102 to a first database 1101.
[0152] According to embodiments of this disclosure, electronic device 200 may switch between databases that periodically store user conversation history information or after starting or ending a specific operation or when the database storage space is insufficient.
[0153] For example, electronic device 200 can send information related to the user's conversation history stored in second database 1102 to first database 1101 according to a predetermined time period (e.g., but not limited to 6 hours, 1 day, 1 month, etc.), and delete information related to the user's conversation history from second database 1102.
[0154] As another example, when the dialogue service ends, the electronic device 200 can send information related to the user's dialogue history accumulated in the second database 1102 to the first database 1101 while providing the dialogue service, and delete the information related to the user's dialogue history from the second database 1102.
[0155] According to embodiments of this disclosure, when switching between databases, electronic device 200 can summarize information other than sensitive user information, thereby mitigating the risk of leakage of user personal information and reducing memory usage.
[0156] The raw data, as unprocessed data, can be stored in the second database 1102. The second database 1102 can store the raw data as if it were input into the electronic device 200 as user conversation history information.
[0157] For example, users may not want to store detailed information related to their personal information (e.g., specific conversation content, captured user images, user voice, user billing information, user location, etc.) in electronic device 200 for an extended period. Therefore, according to embodiments of this disclosure, electronic device 200 can manage conversation history information, including user-sensitive information, such that the conversation history information is stored in a second database 1102, which stores the conversation history information only for a short period of time.
[0158] The processed data can be stored in a first database 1101. The first database 1101 can store data summarized from the original data stored in the second database 1102, excluding user-sensitive information, as the user's conversation history information.
[0159] like Figure 11 As shown, the original content of the dialogue between the user and the electronic device 200, stored in the second database 1102, can be summarized as data about the dialogue category, content, and products of interest, and is stored in the first database 1101. The captured image frames and user voice stored in the second database 1102 can be summarized as the user's mood at the time the dialogue service was provided, and are stored in the first database 1101. Furthermore, the user's payment information stored in the second database 1102 can be summarized as the products purchased by the user and the purchase price, and is stored in the first database 1101.
[0160] According to embodiments of this disclosure, electronic device 200 may share at least one of user face ID, service usage history, or conversation history with other electronic devices. Figure 12a This is a diagram illustrating an exemplary process according to an embodiment of the present disclosure, wherein multiple electronic devices 200-a, 200-b, and 200-c share a user's conversation history with each other. For example, electronic devices 200-a, 200-b, and 200-c may be unmanned kiosks located in different spaces (e.g., different floors) of a store. User 10 can receive guidance about products or assistance in purchasing products based on the conversation services provided by electronic devices 200-a, 200-b, and 200-c.
[0161] refer to Figure 12a Electronic device 200-c can provide dialogue services to user 10. Electronic device 200-c can receive speech input from user 10 and generate and output response messages to the speech input.
[0162] Figure 12b This is a diagram illustrating an exemplary process according to an embodiment of the present disclosure, wherein multiple electronic devices 200-a, 200-b and 200-c share a user’s conversation history with each other.
[0163] Reference Figure 12b After completing negotiation with electronic device 200-c, user 10 can move away from electronic device 200-c by a distance greater than or equal to a predetermined distance. Electronic device 200-c can recognize that the conversation has been paused based on the distance from user 10. Electronic device 200-c can store information in a database about the history of conversations with user 10 that occurred during the current session. For example, electronic device 200-c can store information about the history of conversations with user 10 that occurred during the current session in a second database included in electronic device 200-c.
[0164] Figure 12c This is a diagram illustrating an exemplary process according to an embodiment of the present disclosure, wherein multiple electronic devices 200-a, 200-b and 200-c share a user’s conversation history with each other.
[0165] refer to Figure 12c Electronic device 200-c may, for example but not limited to, share or broadcast the facial ID of user 10 with which it has already negotiated to other electronic devices 200-a and 200-b in the store.
[0166] Figure 12d This is a diagram illustrating an exemplary process according to an embodiment of the present disclosure, wherein multiple electronic devices 200-a, 200-b and 200-c share a user’s conversation history with each other.
[0167] Reference Figure 12d After viewing the second floor of the store, user 10 can descend to the first floor and approach electronic device 200-a. Electronic device 200-a can identify user 10 to provide conversational services. When it is determined that the identified user 10 corresponds to a facial ID shared by electronic device 200-c, electronic device 200-a can request electronic device 200-c to share a database in which information related to the shared facial ID is stored.
[0168] Electronic device 200-c can share a database with electronic device 200-a, which stores the dialogue history corresponding to the facial ID of user 10. Electronic device 200-a can interpret user speech input based on the dialogue history stored in the shared database. Therefore, even when electronic device 200-a receives speech input from user 10 related to the dialogue with electronic device 200-c, electronic device 200-a can output a response message that ensures the continuity of the dialogue.
[0169] The configuration of the electronic device 200 according to embodiments of the present disclosure will now be described in more detail. Each component of the electronic device 200 described below can perform each operation of the method for providing dialogue services as described above, performed by the electronic device 200.
[0170] Figure 13a This is a block diagram illustrating an exemplary configuration of an exemplary electronic device 200 according to an embodiment of the present disclosure.
[0171] Electronic device 200 for providing conversational services may include a processor (e.g., including processing circuitry) 250, which provides conversational services to a user by executing one or more instructions stored in memory. Although Figure 13a The illustrated electronic device 200 includes a processor 250, but embodiments of this disclosure are not limited thereto. The electronic device 200 may include multiple processors. When the electronic device 200 includes multiple processors, the operation and functions of the processor 250, which will be described below, may be performed in part by the processors.
[0172] The input device 220 of the electronic device 200 may include various input circuits and receive user voice input.
[0173] According to embodiments of this disclosure, processor 250 can recognize time expressions representing time in text obtained from user speech input.
[0174] Processor 250 can obtain text by performing speech recognition on user utterance input and perform nested operations to map the text to multiple vectors. For example, by applying a bidirectional LSTM model to the mapped vectors, processor 250 can assign BIO tags to at least one morpheme included in the text that represents at least one of a point in time, duration, or time period. Processor 250 can determine entities included in the text that represent at least one of a point in time, duration, or time period as a time expression based on the BIO tags.
[0175] According to embodiments of this disclosure, processor 250 can determine time points related to user speech input based on time expressions.
[0176] The processor 250 can predict probability values, for example, the recognized time expression indicating the probability of each of a plurality of time points, and determine the time point corresponding to the highest probability value among the predicted probability values as the time point related to the user's utterance input.
[0177] According to embodiments of this disclosure, processor 250 can select from a plurality of databases used to store information related to the dialogue history of users using the dialogue service the database corresponding to the time point associated with the user's utterance input.
[0178] Multiple databases may include a first database for storing information about user dialogue history accumulated before a preset time point and a second database for storing information about user dialogue history accumulated after the preset time point. When the time point associated with the user's speech input is before the preset time point, the processor 250 can select the first database from the databases. When the time point associated with the user's speech input is after the preset time point, the processor 250 can select the second database from the databases.
[0179] Furthermore, the first database can be stored on an external server, while the second database can be stored in the electronic device 200. The preset time point t used as a reference point for selecting the database can be one of the following: when at least some information about the user's conversation history included in the second database is switched to be included in the first database; when the user's facial image is obtained; and when the conversation service is started.
[0180] According to embodiments of this disclosure, processor 250 can interpret text based on information related to a user's conversation history obtained from a selected database.
[0181] Processor 250 can identify entities included in the text that require detailed description. Processor 250 can obtain detailed description information for the identified entities by retrieving information about the user's dialogue history from a selected database. Processor 250 can use an NLU model to interpret the text and detailed description information. Processor 250 can determine the type of response message by applying a DM model to the interpretation results and use an NLG model to generate a response message of the determined type.
[0182] The processor 250 can generate a response message to the received user speech input based on the interpretation results. The output device 230 of the electronic device 200 may include various output circuits and output the generated response message.
[0183] The configuration of the electronic device 200 according to various embodiments of the present disclosure is not limited to... Figure 13a The configuration shown in the block diagram. For example, Figure 13bThis is a block diagram illustrating an exemplary configuration of an exemplary electronic device 200 according to another embodiment of the present disclosure.
[0184] Reference Figure 13b According to another embodiment of this disclosure, an electronic device 200 may include a communicator 210, which may have various communication circuits, and receives user speech input via an external device and sends response messages to the user speech input to the external device. A processor 250 may select a database based on time points associated with the user speech input and generate response messages based on user dialogue history stored in the selected database. The above relates to... Figure 13a The descriptions already provided have been omitted.
[0185] Figure 14 This is a block diagram illustrating an exemplary electronic device 200 according to an embodiment of the present disclosure.
[0186] like Figure 14 As shown, the input device 220 of the electronic device 200 may include various input circuits and receive user input for controlling the electronic device 200. According to embodiments of this disclosure, the input device 220 may include a user input device, such as a touch panel for receiving user touches, buttons for receiving user press operations, a wheel for receiving user rotation operations, a keyboard, a dome switch, etc., but not limited to these. For example, the input device 220 may include at least one of, but not limited to, a camera 221 for recognizing a user's face, a microphone 223 for receiving user speech input, or a payment device 225 for receiving user payment information.
[0187] According to embodiments of this disclosure, the output device 230 of the electronic device 200 may include various output circuits and output information, which are received from the outside, processed by the processor 250, or stored in the memory 270 or at least the database 260 in at least one form, such as, but not limited to, light, sound, image, or vibration. For example, the output device 230 may include at least one of a display 231 or a speaker 233 for inputting and outputting response messages to user speech.
[0188] According to embodiments of this disclosure, the electronic device 200 may further include at least one database 260 for storing a user's conversation history. According to embodiments of this disclosure, the database 260 included in the electronic device 200 may store user conversation history information accumulated up to a preset time point.
[0189] According to embodiments of this disclosure, the electronic device 200 may further include a memory 270. The memory 270 may include at least one of data used by the processor 250, results processed by the processor 250, commands executed by the processor 250, or artificial intelligence (AI) models used by the processor 250.
[0190] The memory 270 may include at least one type of storage medium, such as flash memory, hard disk memory, multimedia card micro-memory, card-type memory (e.g., SD card or XD memory), random access memory (RAM), static RAM (SRAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), PROM, magnetic storage, magnetic disk, or optical disk.
[0191] although Figure 14 The database 260 and memory 270 are shown as separate components, but embodiments of this disclosure are not limited thereto. For example, the database 260 may be included in the memory 270.
[0192] According to embodiments of this disclosure, the communicator 210 may include various communication circuits and communicate with external electronic devices or servers using wireless or wired communication methods. For example, the communicator 210 may include a short-range wireless communication module, a wired communication module, a mobile communication module, and a broadcast receiving module.
[0193] According to embodiments of this disclosure, electronic device 200 can share at least one of, but not limited to, user facial ID, service usage history, or conversation history with another electronic device via communicator 210. For example, after a conversation service provided to a user ends, electronic device 200 can send the user's facial ID to another electronic device. When a user wishes to receive a conversation service through another electronic device, the other electronic device can identify the user and, based on determining that the identified user corresponds to the facial ID of the receiving user, request information related to the user's conversation history from electronic device 200. In response to the request received from the other electronic device, electronic device 200 can send the information about the user's conversation history stored in database 260 to the other electronic device.
[0194] Figure 15a This is a block diagram illustrating an exemplary processor 250 included in an electronic device 200 according to an embodiment of the present disclosure. Figure 15b This is a block diagram illustrating another exemplary processor 250 according to an embodiment of the present disclosure.
[0195] According to embodiments of this disclosure, the operations and functions performed by the processor 250 included in the electronic device 200 can be... Figure 15aThe various modules shown are used to represent this. Some or all of the modules can be implemented using a variety of hardware and / or software components that perform specific functions.
[0196] The face recognition module 1510 may include various processing circuits and / or executable program elements, and is used for recognizing faces through a camera ( Figure 14 (221) The module of the face in the captured image.
[0197] The service management module 1520 may include various processing circuits and / or executable program elements, and is a module for managing the user's usage history of the electronic device 200, and can manage activity history, such as purchasing products and / or searching product information via the electronic device 200.
[0198] The speech recognition module 1530 may include various processing circuits and / or executable program elements, and obtains text from user speech input and generates a response message to the user speech input based on the result of interpreting the text.
[0199] The database management module 1540 may include various processing circuits and / or executable program elements, and selects at least one database from multiple databases for retrieving dialogue history information, and manages the time periods during which information stored in the database is deleted.
[0200] Reference Figure 15b , Figure 15a The face recognition module 1510 may include a face detection module and a face search module. The face detection module includes various processing circuits and / or executable program elements for detecting faces in an image, and the face search module includes various processing circuits and / or executable program elements for searching a database of detected faces.
[0201] In addition, refer to Figure 15b , Figure 15aThe speech recognition module 1530 may include at least one of the following: an automatic speech recognition (ASR) module including various processing circuits and / or executable program elements for converting speech signals into text signals; an NLU module including various processing circuits and / or executable program elements for interpreting the meaning of text; an entity extraction module including various processing circuits and / or executable program elements for extracting entities included in text; a classification module including various processing circuits and / or executable program elements for classifying text according to text categories; a context management module including various processing circuits and / or executable program elements for managing dialogue history; a time context detection module including various processing circuits and / or executable program elements for detecting time expressions in user utterance input; or an NLG module including various processing circuits and / or executable program elements for generating response messages corresponding to the results of interpreting text and time expressions.
[0202] In addition, refer to Figure 15b , Figure 15a The database management module 1540 may include a deletion period management module and a database selection module. The deletion period management module includes various processing circuits and / or executable program elements for managing the period during which information stored in the first database 1561 or the second database 1562 is deleted. The database selection module includes various processing circuits and / or executable program elements for selecting at least one database from the first and second databases 1561 and 1562 to obtain and store information.
[0203] Figure 16 This is a block diagram illustrating an exemplary configuration of a speech recognition module 1530 according to an embodiment of the present disclosure.
[0204] Reference Figure 16 According to embodiments of the present disclosure, the speech recognition module 1530 included in the processor 250 of the electronic device 200 may include an ASR module 1610, an NLU module 1620, a DM module 1630, an NLG module 1640, and a text-to-speech (TTS) module 1650, each module may include various processing circuits and / or executable program elements.
[0205] The ASR module 1610 can convert speech signals into text. The NLU module 1620 can interpret the meaning of the text. The DM module 1630 can guide the dialogue by managing contextual information, including dialogue history, determining the category of questions, and generating responses to those questions. The NLG module 1640 can convert responses written in computer languages into natural language that humans can understand. The TTS module 1650 can convert text into speech signals.
[0206] According to embodiments of this disclosure, Figure 16 The NLU module 1620 can interpret text obtained from user speech by performing preprocessing (1621), performing nesting (1623), applying a time expression extraction model (1625), and applying a time point prediction model (1627).
[0207] In preprocessing operation 1621, speech recognition module 1530 can remove special characters included in the text, unify synonyms into single words, and perform morphological analysis using speech part (POS) tags. In nesting operation 1623, speech recognition module 1530 can perform nesting on the preprocessed text. In operation 1623, speech recognition module 1530 can map the preprocessed text to multiple vectors.
[0208] In operation 1625, which applies the time expression extraction model, the speech recognition module 1530 can extract the time expression included in the text obtained from the user's speech input based on the nested results. In operation 1627, which applies the time point prediction model, the speech recognition module 1530 can predict how much the extracted time expression represents the past relative to the current time point.
[0209] The following will refer to Figure 17 and Figure 18 The operations of applying time expression extraction models and applying time point prediction models are described in more detail in 1625 and 1627.
[0210] Figure 17 This is a diagram illustrating an exemplary time expression extraction model (e.g., including various processing circuits and / or executable program elements) according to embodiments of the present disclosure.
[0211] For example, applying time expression extraction model 1625 can include based on Figure 17 The AI model shown processes the input data.
[0212] According to embodiments of this disclosure, the speech recognition module 1530 can receive text obtained by converting utterance input into an input sentence. The speech recognition module 1530 can nest the input sentence at the word and / or character level. The speech recognition module 1530 can perform cascading nesting to use word nesting results and character nesting results together. The speech recognition module 1530 can generate a Conditional Random Field (CRF) by applying a bidirectional LSTM model to the text mapped to multiple vectors.
[0213] The speech recognition module 1530 can generate a CRF by applying a probability-based tagging model under predetermined conditions. According to embodiments of this disclosure, the speech recognition module 1530 can pre-learn conditions in which words, stems, or morphemes are more likely to be temporal expressions, and label portions based on these pre-learned conditions that have a probability value greater than or equal to a threshold. For example, the speech recognition module 1530 can extract temporal expressions using BIO tagging.
[0214] Figure 18 This is a diagram illustrating an exemplary time-point prediction model according to an embodiment of the present disclosure.
[0215] For example, applying the point-in-time prediction model 1627 may include based on Figure 18 The AI model shown processes the input data.
[0216] According to embodiments of this disclosure, the speech recognition module 1530 can determine time points associated with time expressions based on time expressions identified in text by a time expression extraction model. The speech recognition module 1530 can predict the probability values of which time points are represented by the identified time expressions and determine the time points with the highest probability values as the time points represented by the time expressions.
[0217] The speech recognition module 1530 can pre-learn time points indicated by various time expressions. The speech recognition module 1530 can derive a graph 1810 including probability values (e.g., the recognized time expression represents the probability of multiple time points) by applying the pre-trained model to the recognized time expression (as input). In the graph 1810, the x-axis and y-axis can represent time points and probability values, respectively. The graph 1810 can represent probability values for any specified time point within a specific time interval, or it can display probability values at the time point when a past utterance was made or at the time point when the dialogue service was used.
[0218] When a time point related to a time expression is determined, the speech recognition module 1530 can perform a binary classification 1820 to determine whether the determined time point is before or after a preset time point.
[0219] The speech recognition module 1530 may select a first database 261 when the time point related to the time expression is determined to be before a preset time point based on the determination result via binary classification 1820. The speech recognition module 1530 may select a second database 263 when the time point related to the time expression is determined to be after the preset time point based on the determination result via binary classification 1820.
[0220] According to embodiments of this disclosure, the speech recognition module 1530 can interpret text based on user dialogue history obtained from a selected database. Although Figure 16This section only shows the process by which the NLU module 1620 of the speech recognition module 1530 determines time points in the text related to the utterance input and selects a database based on those time points. However, the speech recognition module 1530 can perform the NLU process again to interpret the text based on information obtained from the selected database that relates to the user's dialogue history. The NLU module 1620 can specify at least one entity included in the text and interpret the specified text based on information obtained from the selected database that relates to the user's dialogue history.
[0221] The DM module 1630 can receive the result of the NLU module 1620 interpreting the detailed text as input, and take into account state variables such as dialogue history to output a list of instructions for the NLG module 1640. The NLG module 1640 can generate a response message to the user's utterance input based on the received list of instructions.
[0222] According to various embodiments of this disclosure, electronic device 200 can utilize AI technology throughout the process of providing a dialogue service to a user. The AI-related functions according to this disclosure are operated by a processor and memory. The processor may include one or more processors. In this case, the one or more processors may be general-purpose processors, such as, but not limited to, a central processing unit (CPU), application processor (AP), or digital signal processor (DSP); a dedicated graphics processor, such as a graphics processing unit (GPU) or vision processing unit (VPU); or a dedicated AI processor, such as a neural processing unit (NPU). The one or more processors can control the input data to be processed according to predetermined operating rules or AI models stored in memory. When the one or more processors are dedicated AI processors, the dedicated AI processors can be designed with hardware architectures specifically designed for processing particular AI models.
[0223] Predefined operating rules or AI models can be created through a training process. For example, this can refer to predetermined operating rules or AI models designed to perform desired characteristics (or purposes) by training a basic AI model using a learning algorithm that utilizes a large amount of training data. The training process can be performed by a device implementing AI according to embodiments of this disclosure, or by a separate server and / or system. Examples of learning algorithms can include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, and reinforcement learning.
[0224] AI models can include multiple neural network layers. Each neural network layer can have multiple weight values, and neural network computations can be performed via arithmetic operations on the computation results of the previous layer and the multiple weight values of the current layer. The multiple weights in each neural network layer can be optimized by training the AI model. For example, multiple weights can be updated to reduce or minimize the loss or cost values acquired by the AI model during the training process. Artificial neural networks can include deep neural networks (DNNs), and can include, for example, but not limited to, convolutional neural networks (CNNs), DNNs, recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent DNNs (BRDNNs), deep Q-networks (DQNs), etc., but are not limited to these.
[0225] Embodiments of this disclosure can be implemented as software programs including instructions stored in a computer-readable storage medium.
[0226] A computer may refer, for example, to a device configured to retrieve instructions stored in a computer-readable storage medium and operate in response to the retrieved instructions, and may include terminal devices and remote control devices according to embodiments of the present disclosure.
[0227] Computer-readable storage media may be provided in the form of non-transitory storage media. In this respect, "non-transitory" storage media may not include signals and may be tangible, and the term does not distinguish between data that is stored semi-permanently and data that is temporarily stored in a storage medium.
[0228] Furthermore, the electronic devices and methods according to embodiments of this disclosure can be provided in the form of computer program products. These computer program products can be traded as products between sellers and buyers.
[0229] Computer program products may include software programs and computer-readable storage media in which the software programs are stored. For example, a computer program product may include a product in the form of a software program (e.g., a downloadable application) distributed electronically by a manufacturer of an electronic device or through an electronic marketplace (e.g., the Google Play Store and the App Store). For such electronic distribution, at least a portion of the software program may be stored on the storage medium or may be temporarily generated. The storage medium may be a server of the manufacturer, a server of the electronic marketplace, or a relay server used for temporary storage of the software program.
[0230] In a system including a server and a terminal (e.g., a terminal device or a remote control device), the computer program product may include the storage medium of the server or the storage medium of the terminal. Where a third device (e.g., a smartphone) communicates with the server or terminal, the computer program product may include the storage medium of the third device. The computer program product may include software programs sent from the server to the terminal or the third device, or from the third device to the terminal.
[0231] In this configuration, one of the server, terminal, and third device may execute a computer program product to perform the method according to embodiments of the present disclosure. At least two of the server, terminal, and third device may execute a computer program product to perform the method according to embodiments of the present disclosure in a distributed manner.
[0232] For example, a server (e.g., a cloud server, an AI server, etc.) can execute computer program products stored on the server and can control terminals communicating with the server to execute methods according to embodiments of this disclosure.
[0233] As another example, the third device can execute a computer program product and can control a terminal communicating with the third device to perform the method according to embodiments of this disclosure. As a specific example, the third device can remotely control a terminal device or a remote control device to send or receive packaged images.
[0234] When a third device executes a computer program product, the third device can download the computer program product from the server and execute the downloaded computer program product. The third device can execute a computer program product pre-loaded therein and can perform a method according to an embodiment of this disclosure.
[0235] While this disclosure has been illustrated and described with reference to various exemplary embodiments, it should be understood that these exemplary embodiments are intended to be illustrative and not restrictive. Those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure (including the appended claims and their equivalents).
Claims
1. A method for providing a dialogue service performed by an electronic device, the method comprising: Receive speech input; Identify time expressions representing time in the text obtained from the utterance input; The time points related to the utterance input are determined based on the time expression; Select the database corresponding to the determined time point from multiple databases that store information related to the conversation history of users using the conversation service; The text is interpreted based on information related to the user's conversation history obtained from a selected database; Based on the interpretation results, a response message is generated for the utterance input; as well as Output the generated response message. The plurality of databases include a first database and a second database. The first database stores information related to the user's accumulated dialogue history before a preset time point, and the second database stores information related to the user's accumulated dialogue history after the preset time point.
2. The method according to claim 1, wherein, Identifying the time expression includes: The text is obtained by performing speech recognition on the utterance input; and Identify an entity in the text that represents at least one of a point in time, duration, or time period as the time expression.
3. The method according to claim 2, wherein, Determining the entity includes: Nesting is performed on the text to map the text to multiple vectors; By applying a bidirectional long short-term memory (LSTM) model to the plurality of vectors, BIO tags are assigned to at least one morpheme included in the text that represents at least one of the time point, the duration, or the time period; and The entity in the text is identified based on the BIO tag.
4. The method according to claim 1, wherein, Determining the time point associated with the utterance input includes: The prediction includes the probability value of the time expression representing the probability of each of a plurality of time points; and The time point corresponding to the highest probability value among the predicted probability values is determined as the time point related to the utterance input.
5. The method according to claim 1, wherein, Selecting the database includes: Based on the time point associated with the utterance input, prior to the preset time point, the first database is selected from the plurality of databases; and Based on the time point associated with the utterance input, after the preset time point, the second database is selected from the plurality of databases.
6. The method according to claim 5, wherein, The first database is stored on an external server, and the second database is stored in the electronic device. The preset time point includes at least one of the following: the time point when at least some of the information related to the user's dialogue history included in the second database is sent to the first database, the time point when the user's facial image is obtained, and the time point when the dialogue service is started.
7. The method according to claim 1, wherein, The interpretation of the text includes: Identify the entities in the text that require detailed description; By retrieving information related to the user's conversation history from a selected database, detailed description information for specifying the identified entities is obtained; and The text and detailed description information are interpreted using a Natural Language Understanding (NLU) model.
8. The method according to claim 1, wherein, Generating the response message includes: The type of the response message is determined by applying the Dialogue Manager (DM) model to the result of the interpretation; and The response message of the determined type is generated using a Natural Language Generation (NLG) model.
9. The method according to claim 1, further comprising: Obtain the user's facial image; By searching the multiple databases, it is determined whether a facial ID corresponding to the obtained facial image is stored. as well as The dialogue service is initiated based on the determined result.
10. The method according to claim 9, wherein, Initiating the aforementioned dialogue service includes: Based on the facial ID corresponding to the obtained facial image stored in the multiple databases, the stored service usage history mapped to the facial ID is updated; and Based on the fact that the facial ID corresponding to the obtained facial image is not stored in the multiple databases, a new facial ID and a service usage history mapped to the new facial ID are generated.
11. The method of claim 9, further comprising: After the dialogue service ends, the facial ID is sent to another electronic device; as well as In response to a request received from the other electronic device, information related to the user's conversation history stored in the plurality of databases is sent to the other electronic device.
12. An electronic device configured to provide a dialogue service, the electronic device comprising: Memory, which stores one or more instructions; as well as At least one processor is configured to execute one or more instructions to provide the dialogue service to the user. The at least one processor is further configured to execute the one or more instructions to control the electronic device: Receive speech input; Identify time expressions representing time in the text obtained from the utterance input; The time points related to the utterance input are determined based on the time expression; Select the database corresponding to the determined time point from multiple databases that store information related to the conversation history of users using the conversation service; The text is interpreted based on information related to the user's conversation history obtained from a selected database; Based on the interpretation, a response message is generated for the utterance input; and Output the generated response message. The plurality of databases include a first database and a second database. The first database stores information related to the user's accumulated dialogue history before a preset time point, and the second database stores information related to the user's accumulated dialogue history after the preset time point.
13. The electronic device according to claim 12, further comprising: The camera is configured to acquire an image of the user's face. as well as A microphone, configured to receive the speech input, wherein, Prior to initiating the dialogue service, the at least one processor is also configured to execute one or more instructions to control the electronic device: By searching the multiple databases, it is determined whether a face ID corresponding to the obtained face image is stored. Based on the facial ID corresponding to the obtained facial image stored in the multiple databases, the service usage history mapped to the facial ID is updated; as well as Based on the fact that the facial ID corresponding to the obtained facial image is not stored in the multiple databases, a new facial ID and a service usage history mapped to the new facial ID are generated.
14. The electronic device according to claim 13, further comprising: A communication interface, comprising a communication circuit, wherein the communication circuit is configured to: After the dialogue service ends, the facial ID is sent to another electronic device, and In response to a request received from the other electronic device, information related to the user's conversation history stored in the plurality of databases is sent to the other electronic device.
15. A computer-readable recording medium storing a program that, when executed, causes an electronic device to perform operations for providing a conversational service, the operations including: Receive speech input; Identify time expressions representing time in the text obtained from the utterance input; The time points related to the utterance input are determined based on the time expression; Select the database corresponding to the determined time point from multiple databases that store information related to the conversation history of users using the conversation service; The text is interpreted based on information related to the user's conversation history obtained from a selected database; Based on the interpretation results, a response message is generated for the utterance input; as well as Output the generated response message. The plurality of databases include a first database and a second database. The first database stores information related to the user's accumulated dialogue history before a preset time point, and the second database stores information related to the user's accumulated dialogue history after the preset time point.
Citation Information
Patent Citations
System and method for a cooperative conversational voice user interface
US20080091406A1
Reception apparatus, reception system, reception method, and storage medium
US20190095750A1