Emotionally intelligent responses to information-seeking questions

The method and system enhance digital assistant interactions by identifying users' emotional states and generating responses that acknowledge and provide relevant information, addressing emotional and informational needs effectively.

JP7795648B2Active Publication Date: 2026-01-07GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024555199
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-03-18
Filing Date
2023-03-08
Publication Date
2026-01-07
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

Existing voice-enabled digital assistants fail to adequately address users' emotional needs alongside informational queries, lacking the ability to provide emotionally intelligent responses that acknowledge and empathize with users' emotional states.

Method used

A method and system that utilize speech recognition, natural language understanding, and text-to-speech technology to identify users' emotional states and generate responses that include both emotional acknowledgment and relevant information, using prosodic embeddings to adjust tone and content based on the user's emotional state.

Benefits of technology

Enhances user interaction by providing emotionally intelligent responses that address both emotional and informational needs, improving user satisfaction and engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007795648000001
    Figure 0007795648000001
  • Figure 0007795648000002
    Figure 0007795648000002
  • Figure 0007795648000003
    Figure 0007795648000003
Patent Text Reader

Abstract

A method (600) for generating an emotionally intelligent response to an information-seeking question includes receiving audio data (202) corresponding to a query (106) spoken by a user and captured by an assistant-enabled device associated with the user, and processing the audio data using a speech recognition model (211) to determine a transcription (204) of the query. The method also includes performing a query interpretation on the query transcription to identify an emotional state (318) of the user who spoke the query and an action (218) to perform. The method also includes obtaining a response preamble (324) based on the emotional state of the user, and performing the identified action to obtain information responsive to the query. The method further includes generating a response (402) including the obtained response preamble followed by information responsive to the query.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to emotionally intelligent responses to information-seeking questions. [Background technology]

[0002] Voice-enabled environments allow users to utter queries, and digital assistants perform actions to answer the queries. In particular, when interacting with an assistant-enabled device via voice, users may seek emotional connection or acknowledgment from the assistant-enabled device. Therefore, it may be advantageous for an assistant-enabled device to identify emotional needs based on a voice query. In some instances, identifying an emotional need requires determining that one or more words in the voice query indicate the user's emotional need. As a result, a digital assistant receiving a query must have some way of identifying the emotional needs of the user who uttered the query. Furthermore, the digital assistant needs to identify an emotionally intelligent response to the query that meets the user's emotional and informational needs. Summary of the Invention

[0003] One aspect of the present disclosure provides a method for generating emotionally intelligent responses to information-seeking questions. The method includes receiving audio data corresponding to a query spoken by a user and captured by an assistant-enabled device associated with the user, and processing the audio data using a speech recognition model to determine a transcription of the query. The method also includes performing query interpretation on the query transcription to identify an emotional state of the user who spoke the query and an intended action to perform. The method further includes obtaining a response preamble based on the emotional state of the user, performing the identified action to obtain information responsive to the query, and generating a response including the obtained response preamble followed by the information responsive to the query.

[0004] Implementations of the present disclosure may include one or more of the following optional features: In some implementations, performing the identified action to obtain information responsive to the query further includes querying a search engine using one or more terms in the transcription to obtain information responsive to the query. In some examples, the method further includes obtaining a prosodic embedding based on the identified emotional state of the user who spoke the query, and converting, using a text-to-speech (TTS) system, the text representation of the emotionally intelligent response into synthetic speech having a target prosody specified by the prosodic embedding. Wherein performing query interpretation on the query transcription further includes identifying a severity of the user's emotional state, and where obtaining the prosodic embedding is further based on the severity of the user's emotional state.

[0005] In some implementations, obtaining a response preamble based on the user's emotional state further includes querying a preamble data store containing a set of different preambles using the user's identified emotional state, where each preamble in the set of different preambles is mapped to a different emotional state. In some examples, obtaining a response preamble based on the user's emotional state further includes generating a preamble mapped to the user's emotional state using a preamble generator configured to receive the user's emotional state as input.

[0006] In some embodiments, obtaining a response preamble based on the user's emotional state further includes determining whether the user's emotional state indicates an emotional need. In these embodiments, determining whether the user's emotional state includes an emotional need is based on content of the query. Additionally or alternatively, determining whether the user's emotional state includes an emotional need further includes determining whether the user's emotional state is associated with an emotional category. In some embodiments, the method further includes generating a response without obtaining a response preamble when the identified emotional state of the user does not indicate an emotional need.

[0007] Another aspect of the present disclosure provides a system for generating emotionally intelligent responses to information-seeking questions. The system includes data processing hardware and memory hardware in communication with the data processing hardware. The memory hardware stores instructions that, when executed by the data processing hardware, cause the data processing hardware to perform operations including receiving audio data corresponding to a query spoken by a user and captured by an assistant-enabled device associated with the user, and processing the audio data using a speech recognition model to determine a transcription of the query. The operations also include performing query interpretation on the query transcription to identify an emotional state of the user who spoke the query and an action to perform. The operations further include obtaining a response preamble based on the emotional state of the user, performing the identified action to obtain information responsive to the query, and generating a response including the obtained response preamble followed by the information responsive to the query.

[0008] This aspect may include one or more of the following optional features: In some implementations, performing the identified action to obtain information responsive to the query further includes querying a search engine using one or more terms in the transcription to obtain information responsive to the query. In some examples, the operations further include obtaining a prosodic embedding based on the identified emotional state of the user who uttered the query, and converting, using a text-to-speech (TTS) system, the text representation of the emotionally intelligent response into synthetic speech having a target prosody specified by the prosodic embedding. Wherein performing query interpretation on the query transcription further includes identifying a severity of the user's emotional state, and where obtaining the prosodic embedding is further based on the severity of the user's emotional state.

[0009] In some implementations, obtaining a response preamble based on the user's emotional state further includes querying a preamble data store containing a set of different preambles using the user's identified emotional state, where each preamble in the set of different preambles is mapped to a different emotional state. In some examples, obtaining a response preamble based on the user's emotional state further includes generating a preamble mapped to the user's emotional state using a preamble generator configured to receive the user's emotional state as input.

[0010] In some embodiments, obtaining a response preamble based on the user's emotional state further includes determining whether the user's emotional state indicates an emotional need. In these embodiments, determining whether the user's emotional state includes an emotional need is based on content of the query. Additionally or alternatively, determining whether the user's emotional state includes an emotional need further includes determining whether the user's emotional state is associated with an emotional category. In some embodiments, the operations further include generating a response without obtaining a response preamble when the identified emotional state of the user does not indicate an emotional need.

[0011] The details of one or more embodiments of the disclosure are set forth in the accompanying drawings and the description below. Other aspects, features, and advantages will become apparent from the description and drawings, and from the claims. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram of an exemplary system including a digital assistant that generates emotionally intelligent responses to information-seeking questions. [Figure 2] FIG. 1 is a schematic diagram of exemplary components of a digital assistant. [Figure 3] FIG. 1 is a schematic diagram of the intent detection process. [Figure 4] FIG. 1 is a schematic diagram of a response generator process. [Figure 5] FIG. 1 is a schematic diagram of an exemplary training process for promoting an intent model to learn consistent, emotionally intelligent responses to information-seeking questions. [Figure 6] 1 is a flowchart of an exemplary arrangement of operations of a method for generating emotionally intelligent responses to information-seeking questions. [Figure 7] FIG. 1 is a schematic diagram of an exemplary computing device that may be used to implement the systems and methods described herein. DETAILED DESCRIPTION OF THE INVENTION

[0013] Like reference symbols in the various drawings refer to like elements.

[0014] A user's manner of interacting with an Assistant-enabled device is designed to be primarily, but not exclusively, via voice input. However, a user's emotional expectations when using voice input may be higher than when using text input. In particular, when interacting with an Assistant-enabled device via voice, a user may seek emotional connection or affirmation from a digital assistant accessible through the Assistant-enabled device. For example, a user may experience emotional needs, such as anxiety, when interacting with a digital assistant. Due to the personal nature of voice input, a user may expect an Assistant-enabled device to recognize the user's emotional needs and, when providing a response, to include both competent and warm responses. For example, a user may benefit from an answer to a query that includes a preamble that acknowledges and empathizes with the emotions experienced by the user, in addition to an informative answer to the query.

[0015] In some scenarios, the same query from a user represents multiple needs of the user. For example, a user might ask an assistant-enabled device, "I'm looking forward to going for a bike ride today. What's the weather forecast for this afternoon?" Here, the query can address both an emotional need (e.g., connection) and an informational need (e.g., the weather forecast for the user's area). By identifying both of these needs, the assistant-enabled device can generate an emotionally intelligent response to the query that addresses both the user's emotional need and the user's informational need. For example, the assistant-enabled device might generate a response to the user saying, "That sounds exciting. It's supposed to be 72 degrees and sunny this afternoon."

[0016] FIG. 1 illustrates an exemplary system 100 including an assistant-enabled device (AED) 104 and / or a remote system 120 that communicates with the AED 104 via a network 132. The AED 104 and / or the remote system 120 execute a digital assistant 200 with which a user 102 may interact via voice, enabling the digital assistant 200 to generate emotionally intelligent responses to information-seeking questions received from the user 102. In the illustrated example, the AED 104 corresponds to a smart speaker. However, the AED 104 may include other computing devices, such as, but not limited to, a smartphone, a tablet, a smart display, a desktop / laptop, a smartwatch, a smart appliance, headphones, smart glasses / headsets, or a vehicle infotainment device. The AED 104 includes data processing hardware 10 and memory hardware 12 that stores instructions that, when executed on the data processing hardware 10, cause the data processing hardware 10 to perform operations. The remote system 120 (e.g., a server, cloud computing environment) also includes data processing hardware 123 and memory hardware 125 that stores instructions that, when executed on the data processing hardware 123, cause the data processing hardware 123 to perform operations. As described in more detail below, the AED 104 and / or the digital assistant 200 executing on the remote system 120 can execute a speech recognizer 210, a response generator 400, and a text-to-speech (TTS) system 410, and can access one or more information sources 212 and sets of emotional preambles 320 stored in the memory hardware 125.

[0017] The AED 104 includes an array of one or more microphones 16 configured to capture sounds, such as voices, directed at the AED 104. The AED 104 may also include or communicate with an audio output device (e.g., a speaker) 18 that may output audio, such as music and / or synthesized voice 122, from the digital assistant 200. The remote system 120 may be a single computer, multiple computers, or a distributed system (e.g., a cloud environment) with scalable / elastic computing resources 123 (e.g., data processing hardware) and / or storage resources 125 (e.g., memory hardware).

[0018] The AED 104 may include a hot word detector 107 configured to detect the presence of hot words in the streaming audio without performing semantic analysis or speech recognition processing on the streaming audio. The AED 104 may also include an acoustic feature extractor (not shown), which may be implemented as part of the hot word detector or as a separate component for extracting audio data 202 ( FIG. 2 ) from the query 106. For example, with reference to FIGS. 1 and 2 , the acoustic feature extractor may receive streaming audio captured by one or more microphones 16 of the AED 104 corresponding to the query 106 spoken by the user 102 and extract the audio data 202. The audio data 202 may include acoustic features such as Mel-Frequency Cepstrum Coefficients (MFCCs) or filter bank energies calculated over a window of the audio signal. In the illustrated example, the query 106 spoken by the user 102 includes, "Google, I fell down the stairs and got hurt. How far is the nearest hospital?"

[0019] The hotword detector 107 may receive the audio data 202 and determine whether the query 106 includes a particular hotword (e.g., Google) spoken by the user 102. That is, the hotword detector 107 is trained to detect the presence of a hotword (e.g., Google) or one or more variants of the hotword (e.g., Hey Google) in the audio data 202, and may wake the AED 104 from a sleep or hibernation state and trigger the speech recognizer 210 to perform speech recognition on the hotword and / or one or more other terms that follow the hotword, such as a voiced query that follows the hotword and specifies an action to perform.

[0020] Continuing with reference to system 100 of FIG. 1 and digital assistant 200 of FIG. 2, speech recognizer 210 executes an automatic speech recognition (ASR) model 211 (e.g., speech recognition model 211), which may receive audio data 202 as input and use speech recognition model 211 to generate / predict a corresponding transcription 204 for query 106. In the illustrated example, one or more words following the hot words in query 106 and captured in the streaming audio include, "I fell down the stairs and got hurt, how far is the nearest hospital?", which specifies the emotional state 318 of user 102 (i.e., fear, pain) and an action 218 for digital assistant 200 to perform to obtain information 228 responsive to query 106. In response, the response generator 400 generates an emotionally intelligent response 402 including a response preamble 324 followed by the retrieved information 228 in response to the query 106 requesting playback of an audible output from the speaker 18 about the nearest hospital to the user 102. The response generator 400 may generate the emotionally intelligent response 402 as a text representation and convert the text representation of the emotionally intelligent response 402 into synthesized speech 122 using a TTS system 410. In the illustrated example, the digital assistant 200 generates the synthesized speech 122 for audible output from the speaker 18 of the AED 104, saying, "Please remain calm. Providence Hospital is 4.3 miles away. Please call your emergency contact for help." As described in more detail below, the synthesized speech 122, "Please remain calm," corresponds to the response preamble 324, and the synthesized speech 122, "Providence Hospital is 4.3 miles away," corresponds to the information 228 in response to the query 106. In some examples, when the AED 104 includes or is in communication with a display screen, the digital assistant 200 instructs the AED 104 to cause a text representation of the emotionally intelligent response 402 to be displayed on the display screen for the user to read, in addition to or instead of generating a synthesized voice 122 representation of the emotionally intelligent response 402.

[0021] As shown, digital assistant 200 may further process a follow-up action, such as calling user 102's emergency contact, without input from user 102. Specifically, synthesized speech 122 includes "Calling emergency contact for help." Here, digital assistant 200 performs the action of identifying user 102's emergency contact and initiating a call to the emergency contact. Additionally or alternatively, digital assistant 200 may output emotionally intelligent response 402 as a graphical response in addition to outputting emotionally intelligent response synthesized speech 122. For example, digital assistant 200 may generate a text representation of emotionally intelligent response 402 for display on screen 50 while also generating synthesized speech 122 for audible output from AED 104. In another example, AED 104 first seeks user 102's approval before performing the follow-up action. Here, the digital assistant 200 generates a synthesized voice saying, "May I call your emergency contact for help?" and waits for the user 102 to provide authorization to call the emergency contact.

[0022] 2 , digital assistant 200 further includes a natural language understanding (NLU) module 220 configured to perform query interpretation on corresponding transcription 204 to identify the emotional state 318 of user 102 who uttered query 106 and the action 218 specified by query 106 for digital assistant 200 to perform. Specifically, NLU module 220 receives as input the corresponding transcription 204 generated by speech recognizer 210 and performs semantic interpretation on corresponding transcription 204 to identify the emotional state 318 and the action 218. That is, NLU module 220 determines the meaning behind corresponding transcription 204 based on one or more words in corresponding transcription 204 for use by response generator 400 when generating emotionally intelligent response 402.

[0023] The NLU module 220 may include an intent model 310 and an action identifier model 224. The intent model 310 may be configured to identify an emotional state 318 of the user 102 and obtain a response preamble 324 based on the emotional state 318 of the user 102 that addresses the emotional needs of the user 102, while the action identifier model 224 may be configured to identify an action 218 for the digital assistant to perform to obtain information 228 responsive to the query 106. The NLU module 220 may further interpret the corresponding transcription 204 to derive the context of the corresponding transcription 204 to determine the meaning behind the corresponding transcription 204, as well as other information about the environment of the user 102, which may be used by the action identifier model 224 to obtain the information 228 responsive to the query 106. For example, if the query 106 processed by the speech recognizer 210 includes the corresponding transcription 204, “Rainy days tend to make me feel down, what's the weather forecast for tomorrow?”, the NLU module 220 can perform query interpretation on the corresponding transcription 204 to identify that the user 102 has an emotional state 318 corresponding to a sad and / or depressed emotional state 318 and that the user 102 is seeking an action 218, such as checking local weather information for the next day. Additionally, the NLU module 220 can receive context indicating the location of the user 102 and determine the correct region for obtaining a weather forecast. The NLU module 220 may analyze and tag the corresponding transcription 204 as part of its processing. For example, for the text “Rainy days tend to make me feel down,” “depressed” may be tagged as the emotional state 318 (i.e., sadness, indicating an emotional need) and “weather forecast” may be tagged as the action 218 to be performed by the AED 104 (i.e., query a search engine).

[0024] In some implementations, the action identifier model 224 of the NLU module 220 performs the identified action 218 and obtains information 228 responsive to the query 106 by querying the information source 212. The information source 212 may include a data store 216 and / or a search engine 214. The data store 216 may include a plurality of question-answer pairs, where one or more of the questions may correspond to one or more terms in the corresponding transcription 204. In these examples, when the information source 212 identifies a question-answer pair that corresponds to one or more terms in the corresponding transcription 204, the information source 212 returns the answer to the action identifier model 224 as information 228 responsive to the query 106. The data store 216 may further include a respective set of resources associated with the user 102. For example, the data store 216 may include user settings for contact information for the user 102's contacts, the user's 102's personal calendar, the user's 102's email accounts, the user's 102's music collection, and / or other resources associated with the user 102.

[0025] In some examples, performing the identified action 218 includes querying the search engine 214 using one or more terms in the corresponding transcription 204 to obtain information 228 responsive to the query 106. For example, the identified action 218 to obtain information 228 related to the user 102's "how far is the nearest hospital?" may include querying the search engine 214 using the term "nearest hospital" to obtain information 228 related to the nearest hospital to the user 102 in response to the query 106. In some implementations, the action identifier model 224 includes the location of the user 102 at the time the search engine 214 is queried to obtain the information 228 sought by the user 102. The user's location may only be included if the user explicitly consents to sharing their location, which may be revoked by the user 102 at any time. Here, the search engine 214 may return a list of hospitals closest to the user 102, but may only return the nearest hospital (i.e., Providence Hospital is 4.3 miles away) as information 228 responsive to the query 106.

[0026] 3 includes an example intent detection process 300 for identifying an emotional state 318 of a user 102 who uttered a query 106 and obtaining a response preamble 324 based on the emotional state 318 of the user 102. The intent model 310 may include an emotion detector 312 configured to detect the emotional state 318 based on one or more words in the corresponding transcription 204, a severity determiner 314 configured to process the detected emotional state 318 and determine a severity level of the emotional state 318, and a preamble generator 316 configured to receive the emotional state 318 of the user 102 and generate a response preamble 324 based on the emotional state 318 of the user 102. In some implementations, the preamble generator 316 queries an emotion preamble data store 320 with the emotional state 318 of the user 102, and the emotion preamble data store 320 returns the response preamble 324 to the intent model 310.

[0027] The emotion preamble data store 320 includes different sets of response preambles 324 for one or more intent level categories 322, 322a-n. That is, each intent level category 322 may include a respective set of response preambles 324 associated with the intent category 322. Some response preambles 324 may be shared among two or more of the intent level categories 322. In some implementations, the emotion detector 312 determines that the emotion state 318 of the user 102 includes an emotional need by determining that the emotion state 318 is associated with an emotion category corresponding to an intent level category 322 in the emotion preamble data store 320. Here, each intent level category 322 may correspond to a different emotion category (e.g., happy, sad, fearful, surprised, angry, anxious) and may include a respective set of response preambles 324 corresponding to the emotion category. For example, the emotional state of loneliness 318 may be included in the intent level category 322 corresponding to the emotional category of sadness, which includes the response preambles 324 "Let's find someone to talk to," "That's too bad," and "It's going to be okay." Here, the emotional preamble data store 320 maps the response preambles 324 to the emotional state of loneliness 318.

[0028] In some examples, the emotion detector 312 determines whether the emotional state 318 of the user 102 indicates an emotional need before obtaining the response preamble 324. In other words, the emotion detector 312 may function as a filter that determines whether the emotionally intelligent response preamble 324 is included in the response 402. The emotion detector 312 may receive the corresponding transcription 204 as an input and identify the emotional state 318 of the user 102 as an output. The emotional state 318 may be further defined as either a neutral emotional state 318 (e.g., calm, relaxed, bored) or a non-neutral emotional state 318 (e.g., excited, fearful, anxious). When the emotion detector 312 identifies the emotional state 318 of the user 102 as a non-neutral emotional state 318, the user 102 can benefit from an emotionally intelligent response preamble 324 that addresses the emotional needs of the user 102. Therefore, the preamble generator 316 generates the response preamble 324 based on the non-neutral emotional state 318 of the user 102. Conversely, when the emotion detector 312 identifies the emotional state 318 of the user 102 as a neutral emotional state 318, the user 102 cannot benefit from the emotionally intelligent response preamble 324. Here, the response generator 400 generates the response 402 without obtaining the response preamble 324.

[0029] After the emotion detector 312 determines the emotional state 318 of the user 102, the severity determiner 314 may further process the corresponding transcription 204 to determine the severity of the emotional state 318 of the user 102. For example, the sad emotional state 318 may generally include one or more types of sadness, such as low sadness (e.g., feeling depressed), moderate sadness (e.g., feeling sad), and high sadness (e.g., feeling blue). In some examples, the severity may be associated with different prosody for adjusting the synthesized speech 122 generated by the digital assistant 200. Specifically, the TTS system 410 may use the prosody embedding 324 associated with the severity of the emotional state 318 of the user 102 to generate the synthesized speech 122 having a target prosody specified by the prosody embedding 324 that is suitable for addressing the severity of the emotional state 318 of the user 102. Generally, when audibly outputting an emotion-related response 402, the TTS system 410 uses prosodic embedding 324 to adjust / modify prosodic features such as fundamental frequency, duration, and / or amplitude of the synthesized speech 122 to reflect the emotional state 318 of the user 102.

[0030] 2 and 4 , the NLU module 220 generates as output a response preamble 324 based on the emotional state 318 of the user 102 and the information 228 responsive to the query 106. The response generator 400 receives the response preamble 324 and the information 228 as input and combines the response preamble 324 and the information 228 to create a text representation of the emotionally intelligent response 402. Specifically, the text representation includes a response preamble transcription 404 "Stay calm," which addresses the emotional state 318 of the user 102, followed by an information transcription 406 "Providence Hospital is 4.3 miles away," which provides the information 228 responsive to the query 106.

[0031] In some implementations, the TTS system 410 converts the text representation of the emotionally intelligent response 402 into corresponding synthesized speech 122 that can be audibly output from the speaker of the AED 104. When the intent model 310 generates / selects a prosody embedding 326 (e.g., the severity determiner 314 determines a severity level that requires a gentle emotionally intelligent response), the TTS system 410 uses the prosody embedding 326 when converting the text transcription to generate synthesized speech 122 with the target prosody specified by the prosody embedding. For example, when the emotional state 318 is sad (especially when the severity is high), the resulting synthesized speech can have a gentle, soothing prosody that conveys to the user 102 that the digital assistant 200 is aware of the user's 102 emotional state. In some implementations, the TTS system 410 resides on the remote system 120 and transmits audio data packets representing the time-domain speech waveform of the synthesized speech 122 to the AED 104 for audible output from the speaker 18. In another embodiment, the TTS system 410 resides in the AED 104 and receives the textual representation of the emotionally intelligent response 402 (and prosodic embedding 326 ) for conversion into synthesized speech 122 .

[0032] FIG. 5 illustrates an example training process 500 for training the intent model 310 to generate a response preamble 324 based on the identified emotional state 318 of the user 102. The training process 500 may be executed on the remote system 120 of FIG. 1. As shown, the training process 500 obtains one or more training datasets 510 stored in a data store 501 and trains the intent model 310 with the training datasets 510. The data store 501 may reside in the memory hardware 125 of the remote system 120. Each training dataset 510 includes multiple training examples 520, 520a-n, and each training example 520 may include an emotional state transcription 521 paired with a corresponding response preamble 522. As shown, the training example 520 may include the emotional state transcription 521, "I had a stressful day, are there any anxiety support groups nearby?" and the corresponding response preamble 522, "That's a shame." Briefly, the training process 500 trains the intent model 310 to learn to predict response preambles 522 to emotional state transcriptions 521.

[0033] In some implementations, each of the training examples 520 is labeled with an emotional state category that corresponds to an intent-level category 322 of the emotional state transcription 521, such that the intention model 310 learns, via the training process 500, to generate an emotionally intelligent response preamble 324 in response to detecting an emotional state 318 associated with the labeled emotional category. In other implementations, the training examples 520 are not labeled with an emotional category. Instead, the intention model 310 predicts the emotional category of the emotional state transcription 521 by identifying one or more words in the emotional state transcription 521 that correspond to the emotional category. In these implementations, the intention model 310 may identify the words stressful, anxious, and support in the emotional state transcription 521, determine that the identified words indicate the emotional category of anxious, and generate the emotionally intelligent response preamble 324 of "That's a shame."

[0034] In the illustrated example, the intent model 310 receives training examples 520 as input and outputs predictions y r Generate the output prediction y r contains predicted response preambles that are tested for their accuracy in addressing the emotional state 318 of the user 102. At each time step, or batch of time steps, of the training process 500, the intent model 310 can be trained using a loss function 550 based on the output predictions yr of the training examples 520 and the ground truth response preambles 324.

[0035] 6 is a flowchart of an example arrangement of operations for a method 600 for generating emotionally intelligent responses to information-seeking questions. The method 600 includes, at operation 602, receiving audio data 202 corresponding to a query 106 spoken by a user 102 and captured by an assistant-enabled device 104 associated with the user 102. At operation 604, the method 600 includes processing the audio data 202 to determine a transcription 204 of the query 106 using a speech recognition model 211.

[0036] At operation 606, the method 600 also includes performing query interpretation on the transcription 204 of the query 106 to identify an emotional state 318 of the user 102 who uttered the query 106 and an action 218 to perform. The method 600 further includes, at operation 608, obtaining a response preamble 324 based on the emotional state 318 of the user 102. At operation 610, the method 600 also includes performing the identified action 218 to obtain information 228 responsive to the query 106. The method 600 also includes, at operation 612, generating a response 402 including the obtained response preamble 324 followed by the information 228 responsive to the query 106.

[0037] 5 is a schematic diagram of an exemplary computing device 700 that may be used to implement the systems and methods described herein. Computing device 700 is intended to represent various types of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The connections and relationships of components and their functions shown here are intended to be illustrative only and are not intended to limit the implementation of the invention as described and / or claimed herein.

[0038] Computing device 700 includes a processor 710, memory 720, a storage device 730, a high-speed interface / controller 740 connecting to memory 720 and a high-speed expansion port 750, and a low-speed interface / controller 760 connecting to a low-speed bus 770 and storage device 730. Each of the components 710, 720, 730, 740, 750, and 760 are interconnected using various buses and may be mounted on a common motherboard or in other manners as desired. Processor 710 (e.g., data processing hardware 10 of FIG. 1 ) processes instructions for execution in computing device 700, including instructions stored in memory 720 or storage device 730, and can display graphical information for a graphical user interface (GUI) on an external input / output device, such as a display 780 connected to high-speed interface 740. In other implementations, multiple processors and / or multiple buses may be used as desired, along with multiple memories and types of memory. Additionally, multiple computing devices 700 may be connected, each providing a portion of the required operations (eg, as a bank of servers, a group of blade servers, or a multi-processor system).

[0039] Memory 720 (i.e., memory hardware 12 of FIG. 1) stores information non-transiently in computing device 700. Memory 720 may be a computer-readable medium, volatile memory unit(s), or non-volatile memory unit(s). Non-transient memory 720 may be a physical device used to temporarily or permanently store programs (e.g., instruction sequences) or data (e.g., program state information) for use by computing device 700. Examples of non-volatile memory include, but are not limited to, flash memory and read-only memory (ROM) / programmable read-only memory (PROM) / erasable programmable read-only memory (EPROM) / electrically erasable programmable read-only memory (EEPROM) (e.g., typically used for firmware such as boot programs). Examples of volatile memory include, but are not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), phase-change memory (PCM), and disk or tape.

[0040] The storage device 730 can provide mass storage for the computing device 700. In some implementations, the storage device 730 is a computer-readable medium. In various different implementations, the storage device 730 can be a series of devices including a floppy disk drive, a hard disk drive, an optical disk drive, or a tape drive, a flash memory or other similar solid-state memory device, or a storage area network or other configuration of devices. In additional implementations, a computer program product is tangibly embodied on an information carrier. The computer program product includes instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 720, the storage device 730, or memory in the processor 710.

[0041] The high-speed controller 740 manages bandwidth-intensive operations of the computing device 700, and the low-speed controller 760 manages less bandwidth-intensive operations. This role assignment is merely exemplary. In some implementations, the high-speed controller 740 is coupled to memory 720, a display 780 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 750 that can accept various expansion cards (not shown). In some implementations, the low-speed controller 760 is coupled to a storage device 730 and a low-speed expansion port 790. The low-speed expansion port 790 may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) and may be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, etc., or to a network device such as a switch or router, for example, via a network adapter.

[0042] As shown, computing device 700 can be implemented in a number of different forms. For example, it may be implemented as a standard server 700a, or multiple times within a group of such servers 700a, as a laptop computer 700b, or as part of a rack server system 700c.

[0043] Various implementations of the systems and techniques described herein may be realized in digital electronic and / or optical circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementation in one or more computer programs executable and / or interpretable by a programmable system including at least one programmable processor, which may be specialized or general-purpose, coupled to receive data and instructions from and transmit data and instructions to the storage system, at least one input device, and at least one output device.

[0044] A software application (i.e., a software resource) may refer to computer software that causes a computing device to perform a task. In some examples, a software application may be referred to as an "application," "app," or "program." Exemplary applications include, but are not limited to, system diagnostic applications, system management applications, system maintenance applications, word processing applications, spreadsheet applications, messaging applications, media streaming applications, social networking applications, and gaming applications.

[0045] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, non-transitory computer-readable medium, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic circuit (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0046] The processes and logic flows described herein may be implemented by one or more programmable processors, also referred to as data processing hardware, executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be implemented by special-purpose logic circuitry, such as an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose processors, as well as any one or more processors of any kind of digital computer. Generally, a processor receives instructions and data from a read-only memory, a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices, such as magnetic, magneto-optical, or optical disks, for storing data, or is operably connected to receive data from or transmit data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0047] To interact with a user, one or more aspects of the present disclosure can be implemented in a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, or a touch screen, for displaying information to the user, and optionally a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user. For example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input from the user can be received in any form, such as acoustic input, voice input, or tactile input. Additionally, the computer can interact with the user by sending and receiving documents to a device used by the user, for example, by sending a web page to a web browser on the user's client device in response to a request received from the web browser.

[0048] Several embodiments have been described. Of course, it will be understood that various modifications can be made without departing from the spirit and scope of the disclosure. Accordingly, other embodiments are within the scope of the following claims.

Claims

1. A computer-implemented method (600), when executed by data processing hardware (123), causing the data processing hardware (123) to: receiving audio data (202) corresponding to a query (106) spoken by a user and captured by an assistant-enabled device associated with the user; processing the audio data (202) using a speech recognition model (211) to determine a transcription (204) of the query (106); performing query interpretation on the transcription (204) of the query (106); the emotional state (318) of the user who uttered the query (106); and Identifying an action to perform (218); obtaining a response preamble (324) based on the emotional state (318) of the user, wherein obtaining the response preamble (324) based on the emotional state (318) of the user includes determining whether the emotional state (318) of the user indicates an emotional need based on content of the query (106); performing the identified action (218) to obtain information (228) responsive to the query (106); generating a response (402) including the obtained response preamble (324) followed by the information (228) responsive to the query (106); The computer-implemented method (600).

2. 2. The computer-implemented method of claim 1, wherein performing the identified action to obtain the information responsive to the query further comprises querying a search engine using one or more terms in the transcription to obtain the information responsive to the query.

3. The operation is obtaining a prosodic embedding (326) based on the identified emotional state (318) of the user who uttered the query (106); converting the textual representation of the response (402) into synthesized speech (122) having a target prosody specified by the prosody embedding (326) using a text-to-speech (TTS) system (410); further comprising The response (402) is emotionally intelligent 3. The computer-implemented method (600) of claim 1 or 2.

4. performing query interpretation on the transcription (204) of the query (106) further includes identifying a severity of the emotional state (318) of the user; and obtaining the prosodic embedding (326) further comprises: based on the severity of the emotional state (318) of the user. The computer-implemented method (600) of claim 3.

5. 3. The computer-implemented method of claim 1, wherein obtaining the response preamble based on the emotional state of the user further comprises: querying a preamble data store containing a set of different preambles using the identified emotional state of the user, wherein each preamble in the set of different preambles is mapped to a different emotional state.

6. 3. The computer-implemented method of claim 1, wherein obtaining the response preamble based on the emotional state of the user further comprises generating a preamble mapped to the emotional state of the user using a preamble generator configured to receive the emotional state of the user as input.

7. 10. The computer-implemented method of claim 1, wherein determining whether the emotional state of the user includes an emotional need further comprises determining whether the emotional state of the user is associated with an emotional category.

8. 2. The computer-implemented method of claim 1, wherein the operations further include generating the response without obtaining the response preamble when the identified emotional state of the user does not indicate an emotional need.

9. A system (100), comprising: data processing hardware (123); and memory hardware (125) in communication with the data processing hardware (123), the memory hardware (125) storing instructions that, when executed by the data processing hardware (123), cause the data processing hardware (123) to: receiving audio data (202) corresponding to a query (106) spoken by a user and captured by an assistant-enabled device associated with the user; processing the audio data (202) using a speech recognition model (211) to determine a transcription (204) of the query (106); performing query interpretation on the transcription (204) of the query (106); the emotional state (318) of the user who uttered the query (106); and Identifying an action to perform (218); obtaining a response preamble (324) based on the emotional state (318) of the user, wherein obtaining the response preamble (324) based on the emotional state (318) of the user includes determining whether the emotional state (318) of the user indicates an emotional need based on content of the query (106); performing the identified action (218) to obtain information (228) responsive to the query (106); generating a response (402) that includes the obtained response preamble (324) followed by the information (228) responsive to the query (106), the emotionally intelligent response (402); The system (100).

10. 10. The system of claim 9, wherein performing the identified action to obtain the information responsive to the query further comprises querying a search engine using one or more terms in the transcription to obtain the information responsive to the query.

11. The operation is obtaining a prosodic embedding (326) based on the identified emotional state (318) of the user who uttered the query (106); converting the textual representation of the response (402) into synthesized speech (122) having a target prosody specified by the prosody embedding (326) using a text-to-speech (TTS) system (410); further comprising The response (402) is emotionally intelligent A system (100) according to claim 9 or 10.

12. performing query interpretation on the transcription (204) of the query (106) further includes identifying a severity of the emotional state (318) of the user; and obtaining the prosodic embedding (326) further comprises: based on the severity of the emotional state (318) of the user. The system (100) of claim 11.

13. 11. The system of claim 9, wherein obtaining the response preamble based on the emotional state of the user further comprises: querying a preamble data store containing a set of different preambles using the identified emotional state of the user, wherein each preamble in the set of different preambles is mapped to a different emotional state.

14. 11. The system of claim 9 or 10, wherein obtaining the response preamble based on the emotional state of the user further comprises generating a preamble mapped to the emotional state of the user using a preamble generator configured to receive the emotional state of the user as input.

15. 10. The system of claim 9, wherein determining whether the emotional state of the user includes an emotional need further comprises determining whether the emotional state of the user is associated with an emotional category.

16. 10. The system of claim 9, wherein the operations further include generating the response without obtaining the response preamble when the identified emotional state of the user does not indicate an emotional need.

Citation Information

Patent Citations

  • Dialog system and computer program therefor

    JP2018156273A

  • Systems and Methods for Enhancing Responsiveness to Utterances Having Detectable Emotion

    US20190325895A1