An empathetic virtual personal assistant

The VPA system addresses the limitations of conventional VPAs by analyzing emotional states from user inputs to perform accurate actions and maintain user engagement, improving driving safety.

JP7735063B2Active Publication Date: 2025-09-08HARMAN INT IND INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021042045
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-27
Filing Date
2021-03-16
Publication Date
2025-09-08
Estimated Expiration
2041-03-16

AI Technical Summary

Technical Problem

Conventional virtual personal assistants (VPAs) fail to interpret utterances that convey emotional components, leading to incorrect action performance and reduced user engagement, which can divert a user's attention from driving and compromise safety.

Method used

A VPA system that captures various types of user inputs, including vocalizations and non-verbal cues, to identify emotional states and determine appropriate actions, synthesizing outputs that reflect these emotional states to enhance user interaction and engagement.

Benefits of technology

The system enables more accurate action performance and realistic conversations, reducing the need for users to divert their attention from driving and maintaining engagement, thereby enhancing driving safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007735063000001
    Figure 0007735063000001
  • Figure 0007735063000002
    Figure 0007735063000002
  • Figure 0007735063000003
    Figure 0007735063000003
Patent Text Reader

Abstract

To provide a virtual personal assistance responding with emotion.SOLUTION: A virtual private assistant (VPA) is configured to analyze inputs of various types of one or more behaviors related to a user to identify an emotional state of the user on the basis of the inputs. Also, the VPA determines one or more actions to be executed instead of the user on the basis of the inputs and the identified emotional state. Next, the VPA executes one or more actions, and synthesizes an output on the basis of the user's emotional state and the one or more actions. The synthesized output includes one or more semantic elements and one or more emotional elements derived from the user's emotional state. The VPA observes behaviors of the user corresponding to the synthesized output, then, executes various corrections on the basis of the observed behaviors, and improves effectiveness of future dialogues with the user.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Various embodiments relate generally to computer software and virtual personal assistants, and more particularly to emotively responsive virtual personal assistants. [Background technology]

[0002] A "virtual personal assistant" (VPA) is a type of computer program that interacts with a user and performs various actions on their behalf. In doing so, a VPA typically processes vocalizations received from the user and interprets those utterances as one or more commands. The VPA then maps those commands to one or more corresponding actions that it can perform on the user's behalf. After performing an action, the VPA can report the results of the action by conversing with the user through synthesized vocalizations. For example, a VPA may process a vocalization received from a user and interpret the vocalization as a command to "check email." The VPA may map that command to an action to retrieve new emails and perform the action to retrieve the user's new emails. The VPA may then synthesize a vocalization indicating the number of new emails retrieved.

[0003] In some implementations, a VPA can be implemented within a vehicle to allow a user to operate various vehicle functions without significantly diverting their attention from driving. For example, assume a user needs to adjust the vehicle's air conditioning settings to lower the vehicle's interior temperature to a more comfortable level. The user may vocalize the command "turn on air conditioning" to instruct the VPA to turn on the air conditioning. Upon turning on the air conditioning, the VPA may synthesize a vocalization indicating to the user that the associated action has been performed. In this manner, a VPA implemented within a vehicle can help eliminate the need for a user to divert their attention from driving to manually operate various vehicle functions, thereby increasing overall driving safety.

[0004] One drawback of the above techniques is that conventional VPAs can only interpret the semantic components of utterances and therefore cannot correctly interpret utterances that convey information using emotional components. As a result, a VPA implemented in a vehicle may not be able to properly perform a given action on behalf of the user, which may require the user to divert their attention from driving to manually perform the action. For example, suppose a user is listening to the radio and a very loud song suddenly starts playing. The user may immediately instruct the VPA to "Turn the volume down now!" However, the VPA may not be able to interpret the urgency associated with this type of user command and may only lower the volume by one level. To address this misunderstanding, the user must divert their attention from driving and manually lower the radio volume to a more appropriate level, thereby reducing overall driving safety.

[0005] Another drawback of the above techniques is that conventional VPAs often fail to converse realistically with users because they are unable to correctly interpret speech that conveys information using emotional elements. As a result, a VPA implemented in a vehicle can cause users to disengage with the VPA or turn it off entirely, creating situations where the user must divert their attention from driving to manually operate various vehicle functions. For example, assume a user is genuinely excited about a promotion at work and instructs the VPA to identify the fastest route home. If the VPA synthesizes a dull, monotonous speech and reads out related navigation instructions, the user may find the interaction with the VPA discouraging and ultimately turn off the VPA to maintain their excitement level. Such an outcome reduces overall driving safety.

[0006] As noted above, what is needed in the art is a more effective way to interact with a user when a VPA performs actions on the user's behalf. Summary of the Invention [Means for solving the problem]

[0007] Various embodiments include a computer-implemented method for assisting a user while interacting with the user, the computer-implemented method including capturing a first input indicative of one or more behaviors associated with the user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; and outputting the first utterance to the user.

[0008] At least one technical advantage of the disclosed techniques over the prior art is that the disclosed techniques enable a VPA to more accurately determine one or more actions to perform on behalf of a user based on the user's emotional state. Thus, when implemented in a vehicle, the disclosed VPA helps prevent a user from diverting their attention from driving to operate vehicle functions, thereby increasing overall driving safety.

[0009] In order that the above-recited features of the various embodiments may be understood in detail, a more particular description of the inventive concepts briefly summarized above may be rendered by reference to various embodiments, some of which are illustrated in the accompanying drawings. It should be noted, however, that the accompanying drawings illustrate only exemplary embodiments of the inventive concepts and therefore should not be construed as limiting the scope in any manner, as there may be other equally effective embodiments. For example, the present application provides the following: (Item 1) 1. A computer-implemented method for assisting a user while interacting with the user, comprising: capturing a first input indicative of one or more actions associated with the user; determining a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; and The computer-implemented method includes: (Item 2) the user is in a vehicle in which the first input is captured; The computer-implemented method of the preceding paragraph, wherein the first action is performed on behalf of the user by a vehicle subsystem included in the vehicle. (Item 3) determining the first action based on the first emotional state; performing the first action to assist the user; 2. The computer-implemented method of claim 1, further comprising: (Item 4) Identifying the first emotional state of the user includes: identifying a first feature of the first input; Identifying a first emotion type corresponding to the first characteristic; 2. The computer-implemented method of claim 1, comprising: (Item 5) the first input comprises an audio input; 10. The computer-implemented method of claim 1, wherein the first characteristic includes a tone of voice associated with the user. (Item 6) the first input comprises a video input; 10. The computer-implemented method of claim 1, wherein the first feature includes a facial expression made by the user. (Item 7) Identifying the first emotional state of the user includes: determining a first valence value indicative of a location within a spectrum of emotion types based on the first input; determining a first intensity value indicative of a location within a range of intensities corresponding to the location within a spectrum of emotion types based on the first input; 2. The computer-implemented method of claim 1, comprising: (Item 8) 2. The computer-implemented method of claim 1, wherein the first emotional element corresponds to the first valence value and the first intensity value. (Item 9) 2. The computer-implemented method of claim 1, wherein the first emotional element corresponds to at least one of a second valence value or a second intensity value. (Item 10) 2. The computer-implemented method of claim 1, further comprising generating the first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element. (Item 11) A non-transitory computer readable medium storing program instructions that, when executed by a processor, cause the processor to perform steps to assist a user in interacting with the user, the steps comprising: capturing a first input indicative of one or more actions associated with the user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; The non-transitory computer-readable medium. (Item 12) the user is in a vehicle in which the first input is captured; The non-transitory computer-readable medium described in the preceding item, wherein the first operation is performed on behalf of the user by a vehicle subsystem included in the vehicle. (Item 13) determining the first action based on the first emotional state; performing the first action to assist the user; 3. The non-transitory computer-readable medium of any one of the preceding items, further comprising: (Item 14) The step of identifying the first emotional state of the user includes: identifying a first feature of the first input; Identifying a first emotion type corresponding to the first characteristic; Including, the first feature includes a tone of voice associated with the user or a facial expression made by the user; 2. A non-transitory computer-readable medium according to any one of the preceding items. (Item 15) The step of identifying the first emotional state of the user includes: determining a first valence value indicative of a location within a spectrum of emotion types based on the first input; determining a first intensity value indicative of a location within a range of intensities corresponding to the location within a spectrum of emotion types based on the first input; 10. The non-transitory computer-readable medium of any one of the preceding items, comprising: (Item 16) The non-transitory computer-readable medium of any one of the preceding items, further comprising generating the first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element. (Item 17) capturing a second input indicative of at least one action the user performs in response to the output; modifying the response mapping based on the second input and a first objective function that is evaluated to determine how closely the at least one behavior corresponds to a target behavior; 3. The non-transitory computer-readable medium of any one of the preceding items, further comprising: (Item 18) 10. The non-transitory computer-readable medium of claim 1, wherein generating the first utterance includes combining the first emotional element with a first semantic element. (Item 19) generating a transcription of the first input indicating one or more semantic elements contained in the first input; generating the first semantic element based on the one or more semantic elements; 3. The non-transitory computer-readable medium of any one of the preceding items, further comprising: (Item 20) a memory for storing a software application; a processor configured to perform steps when the software application is executed; The steps include: capturing a first input indicative of one or more actions associated with a user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; Including, the system. (Summary) A virtual private assistant (VPA) is configured to analyze various types of input indicating one or more behaviors associated with a user and identify the user's emotional state based on the input. The VPA also determines one or more actions to perform on behalf of the user based on the input and the identified emotional state. The VPA then performs the one or more actions and synthesizes an output based on the user's emotional state and the one or more actions. The synthesized output includes one or more semantic elements and one or more emotional elements derived from the user's emotional state. The VPA observes the user's behavior in response to the synthesized output and then implements various modifications based on the observed behavior to improve the effectiveness of future interactions with the user. [Brief explanation of the drawings]

[0010] [Figure 1A] FIG. 1 illustrates a system configured to implement one or more aspects of various embodiments. [Figure 1B] FIG. 1 illustrates a system configured to implement one or more aspects of various embodiments. [Figure 2] 2 is a more detailed diagram of the VPA of FIG. 1 in accordance with various embodiments. [Figure 3A] 2A-2C illustrate examples of how the VPA of FIG. 1 characterizes and transforms a user's emotional state, according to various embodiments. [Figure 3B] 2A-2C illustrate examples of how the VPA of FIG. 1 characterizes and transforms a user's emotional state, according to various embodiments. [Figure 4A] 2A-2C illustrate examples of how the VPA of FIG. 1 responds to a user's emotional state, according to various embodiments. [Figure 4B] 2A-2C illustrate examples of how the VPA of FIG. 1 responds to a user's emotional state, according to various embodiments. [Figure 4C] 2A-2C illustrate examples of how the VPA of FIG. 1 responds to a user's emotional state, according to various embodiments. [Figure 5A]2A-2C illustrate examples of how the VPA of FIG. 1 compensates for a user's emotional state, according to various embodiments. [Figure 5B] 2A-2C illustrate examples of how the VPA of FIG. 1 compensates for a user's emotional state, according to various embodiments. [Figure 5C] 2A-2C illustrate examples of how the VPA of FIG. 1 compensates for a user's emotional state, according to various embodiments. [Figure 6] 1 is a flow diagram of method steps for synthesizing vocalizations that reflect a user's emotional state, according to various embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0011] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of various embodiments. However, it will be apparent to one skilled in the art that the concepts of the present invention may be practiced without one or more of these specific details.

[0012] As described above, conventional VPAs can only interpret semantic elements of utterances and therefore cannot correctly interpret utterances that convey information using emotional elements. As a result, a VPA implemented in a vehicle may not be able to properly perform certain actions on behalf of a user, which may result in a situation where the user must divert their attention from driving to perform those actions themselves. Furthermore, because conventional VPAs cannot correctly interpret utterances that convey information using emotional elements, the VPA cannot engage in any realistic conversation with the user. As a result, a VPA implemented in a vehicle may cause the user to disengage from the VPA or turn the VPA off entirely, which may result in a situation where the user must divert their attention from driving to operate various vehicle functions.

[0013] To address these issues, various embodiments include a VPA configured to analyze various types of input indicative of one or more behaviors associated with the user. The input may include vocal utterances representing explicit commands for the VPA to execute, as well as nonverbal cues associated with the user, such as facial expressions and / or changes in posture, among others. The VPA identifies the user's emotional state based on the input. The VPA also determines one or more actions to perform on behalf of the user based on the input and the identified emotional state. The VPA then performs the one or more actions and synthesizes an output based on the user's emotional state and the one or more actions. The synthesized output includes one or more semantic elements and one or more emotional elements derived from the user's emotional state. The emotional element(s) of the output can be consistent with or contrast with the user's emotional state, among other possibilities. The VPA observes the user's behavior in response to the synthesized output and then implements various modifications based on the observed behavior to improve the effectiveness of future interactions with the user.

[0014] At least one technical advantage of the disclosed techniques over the prior art is that the disclosed techniques enable a VPA to more accurately determine one or more actions to perform on behalf of a user based on the user's emotional state. Thus, when implemented in a vehicle, the disclosed VPA helps prevent a user from diverting their attention from driving to operate vehicle functions, thereby improving overall driving safety. Another technical advantage of the disclosed techniques is that the disclosed techniques enable a VPA to generate realistic, conversational responses that reflect the user's emotional state. The realistic, conversational responses maintain user engagement with the VPA, reducing situations where the user must turn off the VPA to manually operate vehicle functions, improving overall driving safety. These technical advantages represent one or more technical advances over prior art approaches.

[0015] System Overview 1A and 1B illustrate a system configured to implement one or more aspects of various embodiments. As shown in FIG. 1A, system 100 includes a computing device 110 coupled to one or more input devices 120 and one or more output devices 130.

[0016] Input device 120 is configured to capture input 122 reflecting one or more behaviors associated with user 140. As referred to herein, "behavior" includes any voluntary and / or involuntary act performed by a user. For example, without limitation, "behavior" may include explicit commands issued by the user, facial expressions made by the user, changes in emotion consciously or unconsciously exhibited by the user, as well as changes in the user's posture, heart rate, skin conductivity, pupil dilation, etc. Input device 120 may include a variety of different types of sensors configured to capture different types of data reflecting behaviors associated with the user. For example, without limitation, input device 120 may include an audio capture device that records vocalizations made by the user, an optical capture device that records images and / or video depicting the user, a pupil measurement sensor that measures the dilation of the user's pupils, an infrared sensor that measures blood flow in the user's face and / or body, a heart rate sensor that generates a heart rate reading in beats per minute associated with the user, a galvanic skin response sensor that measures changes in the user's skin conductivity, a body temperature sensor that detects changes in the user's core and / or surface body temperature, an electroencephalogram sensor that detects various brain wave patterns, etc. As described in more detail below, computing device 110 processes input 122 to generate output 132.

[0017] Output device 130 is configured to send output 132 to user 140. While output 132 can include any technically feasible type of data associated with any given sensory type, in practice output device 130 generates and sends audio output to user 140. Thus, output device 130 generally includes one or more audio output devices. For example, without limitation, audio device 130 may include one or more speakers, one or more sound transducers, a set of headphones, a beamforming array, a sound field generator, and / or a sound cone.

[0018] Computing device 110 may be any technically feasible type of computer system, such as a desktop computer, a laptop computer, a mobile device, a virtualized instance of a computing device, or a distributed and / or cloud-based computer system. Computing device 110 includes a processor 112, input / output (I / O) devices 114, and memory 116, all coupled together. Processor 112 includes any technically feasible set of hardware units configured to process data and execute software applications. For example, without limitation, processor 112 may include one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more application-specific integrated circuits (ASICs). I / O devices 114 include any technically feasible set of devices configured to perform input and / or output operations. For example, without limitation, I / O devices 114 may include universal serial bus (USB) ports, serial ports, and / or FireWire ports. In one embodiment, I / O devices 114 may include input devices 120 and / or output devices 130. The memory 116 includes any technically feasible storage medium configured to store data and software applications. For example, without limitation, the memory 116 may include a hard disk, a random access memory (RAM) module, and / or a read-only memory (ROM). The memory 116 includes a virtual private assistant (VPA) 118. The VPA 118 is a software application that, when executed by the processor 112, performs various operations based on the input 122 and generates the output 132.

[0019] In operation, the VPA 118 processes input 122 captured via the input device 120 and identifies an emotional state of the user 140 based on the input. The VPA 118 also determines one or more actions to perform on behalf of the user 140 based on the input 122 and the identified emotional state. The VPA 118 determines the emotional state of the user 140 and one or more actions to perform on behalf of the user 140 using techniques described in more detail below in conjunction with FIGS. 2-3B. The VPA 118 then performs the one or more actions or causes another system to perform the one or more actions. The VPA 118 synthesizes an output 132 based on the emotional state of the user 140 and the one or more actions. The synthesized output includes semantic elements related to the one or more actions and emotional elements derived from and / or influenced by the user's emotional state. The VPA 118 sends the output 132 to the user 140 via the output device 130. The VPA 118 observes the user's 140 behavior in response to the synthesized output and then implements various modifications that may improve the usability of and / or the user's engagement with the VPA 118 .

[0020] As a general matter, system 100 may be implemented as a standalone system or may be integrated with and / or configured to interoperate with any other technically feasible system. For example, without limitation, system 100 may be integrated with and / or configured to interoperate with, among others, a vehicle, a smart home, smart headphones, a smart speaker, a smart television set, one or more Internet of Things (IoT) devices, or a wearable computing system. Figure 1B shows an exemplary implementation in which system 100 is integrated with a vehicle.

[0021] As shown in FIG. 1B , system 100 is integrated with a vehicle 150 in which user 140 rides. System 100 is coupled to one or more subsystems 160(0)-160(N) within the vehicle. Each subsystem 160 provides access to one or more vehicle functions. For example, without limitation, a given subsystem 160 may be a climate control subsystem that provides access to the vehicle's climate control functions, an infotainment subsystem that provides access to the vehicle's infotainment functions, a navigation subsystem that provides access to the vehicle's navigation functions, an autonomous driving subsystem that provides access to the vehicle's autonomous driving functions, etc. A VPA 118 within system 100 is configured to access one or more subsystems 160 to perform actions on behalf of user 140.

[0022] 1A and 1B generally, those skilled in the art will understand that system 100 in general, and VPA 118 in particular, can be implemented in any technically feasible environment to perform actions on behalf of user 140 and synthesize vocalizations based on the emotional state of user 140. Various examples of devices that can implement and / or include a VPA include mobile devices (e.g., mobile phones, tablets, laptops, etc.), wearable devices (e.g., watches, rings, bracelets, headphones, AR / VR head-mounted devices, etc.), consumer products (e.g., gaming, gambling, etc.), smart home devices (e.g., smart lighting systems, security systems, smart speakers, etc.), communication systems (e.g., audio conferencing systems, video conferencing systems, etc.), and the like. The VPA may be deployed in a variety of environments, such as, but not limited to, road vehicle environments (e.g., consumer automobiles, commercial trucks, ride-hailing vehicles, snowmobiles, all-terrain vehicles (ATVs), semi-autonomous and fully autonomous vehicles, etc.), aerospace and / or aviation environments (e.g., airplanes, helicopters, spacecraft, electric vertical take-off and landing aircraft (eVTOLs), etc.), maritime and undersea environments (e.g., boats, ships, jet skis), etc. The various modules that implement the overall functionality of the VPA 118 are described in more detail below in conjunction with FIG. 2.

[0023] Software Overview Figure 2 is a more detailed diagram of the VPA of Figure 1, in accordance with various embodiments. As shown, the VPA 118 includes a semantic analyzer 210, a sentiment analyzer 220, a response generator 230, an output synthesizer 240, and a mapping modifier 250. These various elements are implemented as software and / or hardware modules that interoperate to perform the functionality of the VPA 118.

[0024] In operation, the semantic analyzer 210 receives the input 122 from the user 140 and performs a speech-to-text transcription operation on the utterances contained in the input to generate an input transcript 212. The input transcript 212 includes text data reflecting commands, questions, utterances, and other forms of verbal communication issued by the user 140 to the VPA 118 to elicit a response from the VPA 118. For example, without limitation, the input transcript 212 may include commands indicating actions the user 140 wants the VPA 118 to perform. The input transcript 212 may also indicate questions the user 140 wants the VPA 118 to answer or utterances the user 140 made to the VPA 118. The semantic analyzer 210 sends the input transcript 212 to the response generator 230.

[0025] The sentiment analyzer 220 also receives the input 122 from the user 140 and then performs a sentiment analysis operation on the input 122 to identify an emotional state 222 associated with the user 140. The sentiment analyzer 220 may also identify the emotional state 222 based on the input transcript 212. In generating the emotional state 222, the sentiment analyzer 220 may implement any technically feasible approach for characterizing the emotional state of a living being and, in doing so, may process data contained within the input 122 in any technically feasible format. For example, without limitation, the sentiment analyzer 220 may process vocalizations received from the user to quantify the pitch, tone, timbre, volume, and / or other acoustic features of the vocalizations. The sentiment analyzer 220 may then map those features to a particular emotional state or sentiment metric. In another example, without limitation, the sentiment analyzer 220 may process a video of facial expressions made by the user 140 and then classify the facial expressions as corresponding to a particular emotional state. In one embodiment, emotional state 222 may include a valence value indicating a particular emotional type, an intensity value indicating the intensity at which that emotional type is expressed, and / or an arousal level corresponding to that emotional type, as also described below in conjunction with Figures 3A and 3B. Those skilled in the art will appreciate that emotion analyzer 220 may implement any technically feasible approach to characterizing emotions and may generate emotional state 222 to include any technically feasible data describing emotions. Emotion analyzer 220 sends emotional state 222 to response generator 230.

[0026] Response generator 230 is configured to process input transcript 212 and emotional state 222 to generate actions 232. Each action 232 may correspond to a command received from user 140 and included in input transcript 212. VPA 118 may perform a given action 232 in response to a given command on behalf of user 140, or offload the action to another system to perform on behalf of user 140. For example, without limitation, VPA 118 may offload a given action 232 to one of vehicle subsystems 160 shown in FIG. 1B . In one embodiment, response generator 230 may generate a set of likely actions based on input transcript 212 and then select a subset of those actions to perform based on emotional state 222.

[0027] The response generator 230 is further configured to process the input transcript 212 and the emotional state 222 to generate semantic elements 234. The semantic elements 234 include text data that is synthesized into the output 132 and then sent to the user 140, as described further below. The semantic elements 234 may include words, phrases, and / or sentences that are contextually related to the input transcript 212. For example, without limitation, the semantic elements 234 may include an acknowledgment that a command was received from the user 140. The semantic elements 234 may also describe and / or reference actions 232 and / or the status of performing those actions. For example, without limitation, the semantic elements 234 may include an indication that a particular action 232 was initiated in response to a command received from the user 140.

[0028] The response generator 230 is further configured to process the input transcript 212 and the emotional state 222 to generate emotional elements 236. The emotional elements 236 indicate particular emotional qualities and / or attributes, derived from the emotional state 222, that are incorporated into the output 132 during synthesis. For example, without limitation, a given emotional element 236 may include a particular pitch, tone, timbre, volume, speaking rate, and / or annunciation level at which an utterance should be synthesized to reflect the particular emotional qualities and / or attributes.

[0029] In various embodiments, the response generator 230 is configured to generate audio responses that have the same or similar semantic content but that may vary based on various speech levels, such as (i) the overall tempo, loudness, and pitch of the synthesized voice, (ii) speech emotion parameters described in more detail below, (iii) nonverbal and nonverbal vocalizations, paralinguistic breathing (e.g., laughter, coughs, whistles, etc.), and (iv) non-speech transitions (e.g., beeps, chirps, clicks, etc.). These audio responses vary in perceived emotional effect; for example, the same semantic content can be rendered with speech that feels soft and relaxed to the user, or with speech that feels rushed and abrupt. These variations in perceived emotional effect can be produced using words with soft versus hard sounds and multisyllabic versus abrupt rhythms. For example, sounds such as "l," "m," and "n," and long or diphthongs reinforced by gentle polysyllabic rhythms are interpreted as more "gentle" than hard sounds such as "g" and "k," short vowels, and words with abrupt rhythms. The field of sound symbolism (described, for example, at http: / / grammar.about.com / od / rs / g / soundsymbolismterm.htm) offers a variety of heuristics that attempt to imbue emotion by linking specific sound sequences to specific meanings in speech. The above-mentioned speech emotion parameters generally include: (i) pitch parameters (e.g., accent shape, mean pitch, contour slope, final fall, and pitch range), (ii) timing parameters (e.g., speaking rate and stress frequency), (iii) voice quality parameters (e.g., breathiness, intelligibility, laryngealization, loudness, pause discontinuity, and pitch continuity), and (iv) pronunciation parameters. In addition to audible speech, the audio output may also include non-verbal vocalizations, such as laughter, breathing, hesitations (e.g., "uhm"), and / or non-verbal agreements (e.g., "aha").

[0030] In some cases, the given emotional element 236 may complement or align with the emotional state 222. For example, without limitation, if the emotional state 222 indicates that the user 140 is currently “happy,” the emotional element 236 may include a particular tone of voice generally associated with “happiness.” Conversely, the given emotional element 236 may differ from the emotional state 222. For example, without limitation, if the emotional state 222 indicates that the user 140 is currently “angry,” the emotional element 236 may include a particular tone of voice generally associated with “calm.” Figures 3A and 3B show various examples of how the response generator 230 may generate the emotional element 236 based on the emotional state 222.

[0031] In one embodiment, the response generator 230 may perform a response mapping 238 that maps the input transcription 212 and / or emotional state 222 to one or more actions 232, one or more semantic elements 234, and / or one or more emotional elements 236. The response mapping 238 may be any technically feasible data structure capable of processing one or more inputs and generating one or more outputs. For example, without limitation, the response mapping 238 may include an artificial neural network, a machine learning model, a set of heuristics, a set of conditional statements, and / or one or more lookup tables, among others. In various embodiments, the response mapping 238 may be retrieved from a cloud-based repository of response mappings generated for different users by different instances of the system 100. Additionally, the response mapping 238 may be modified using techniques described in more detail below and then uploaded to the cloud-based repository for use by other instances of the system 100.

[0032] The response generator 230 sends the semantic elements 234 and the emotional elements 236 to the output synthesizer 240. The output synthesizer 240 is configured to combine the semantic elements 234 and the emotional elements 236 to generate the output 132. The output 132 is typically in the form of a synthesized vocalization. The output synthesizer 132 sends the output 132 to the user 140 via the output device 130. Through the above techniques, the VPA 118 uses the emotional state of the user 140 to more effectively interpret input received from the user 140 and more effectively generate vocalizations in response to the user 140. Furthermore, the VPA 118 can adapt based on the user's 140 reaction to a given output 132 to improve usability and engagement with the user 140.

[0033] Specifically, the VPA 118 is configured to capture feedback 242 reflecting one or more actions that the user 140 performed in response to the output 132. The VPA 118 then updates the emotional state 222 to reflect any observed behavioral changes of the user 140. The mapping modifier 250 evaluates one or more objective functions 252 based on the updated emotional state 222 to quantify the effectiveness of the output 132 in eliciting a particular type of behavioral change in the user 140. For example, without limitation, a given objective function 252 may quantify the effectiveness of mapping a given input transcription 212 to a particular action set 232 based on whether the emotional state 222 indicates whether the user 140 is happy or displeased. In another example, without limitation, a given objective function may quantify the effectiveness of selecting a particular semantic element 234 when generating the output 132 based on whether the emotional state 222 indicates whether the user 140 is interested or disinterested. In yet another example, without limitation, a given objective function 252 may quantify the effectiveness of incorporating a "soothing" tone into the output 132 to calm the user 140 when the user 140 is in a "nervous" emotional state.

[0034] As a general matter, a given objective function 252 may represent a target behavior for a user 140, a target emotional state 222 for a user 140, a target state for a user 140, a target level of engagement with a VPA 118, or any other technically feasible goal that can be assessed based on feedback 242. In embodiments in which the response generator 230 includes a response mapping 238, the mapping modifier 250 may update the response mapping 238 to improve subsequent outputs 132. In the described manner, the VPA 118 can adapt to the particular personalities and idiosyncrasies of different users and thus improve over time in its interpretation and engagement with the user 140.

[0035] Exemplary Characterization and Transformation of Emotional States 3A and 3B show examples of how the VPA of FIG. 1 characterizes and translates a user's emotional state, according to various embodiments. As shown in FIG. 3A, the emotional state 222 is defined using a graph 300 that includes a valence axis 302, an intensity axis 304, and a position 306 plotted against the two axes. The valence axis 302 defines a spectrum of positions that may correspond to various emotion types, such as "joy," "happiness," "anger," and "excitement," among others. The intensity axis 304 defines a range of intensities corresponding to each emotion type. The position 306 corresponds to a particular emotion type indicated via the valence axis 302 and a particular intensity at which that emotion is expressed. In the illustrated example, the position 306 corresponds to a high level of joy.

[0036] The sentiment analyzer 220 generates the emotional state 222 based on any of the various types of analysis described above. The response generator 230 then generates emotional components 236 that similarly define a graph 310 including a valence axis 312 and an intensity axis 314. The graph 310 also includes a position 316 that represents an emotional quality to be included in the output 132 during synthesis. In the illustrated example, position 316, like position 306, corresponds to a high level of joy, and therefore the output 132 is generated with an emotional quality intended to compliment the emotional state 222 of the user 140. The response generator 230 can also generate emotional components 236 that are different from the emotional state 222, as shown in FIG. 3B .

[0037] 3B, emotional state 222 includes location 308 corresponding to an elevated level of dissatisfaction. Response generator 230 generates emotional component 236 to include location 318 corresponding to a lower level of cheerfulness. Thus, output 132 is generated with a different emotional quality than emotional state 222, but may alter the emotional state of user 140 by reducing their dissatisfaction through cheerful vocalizations.

[0038] 3A and 3B generally, in various embodiments, response generator 230 may perform response mapping 238 to map particular locations on graph 300 to other locations on graph 310, thereby achieving various types of transformations between emotional states 222 and emotional elements 236, as described above by way of example. Those skilled in the art will appreciate that VPA 118 may implement any technically feasible approach for characterizing emotional states and / or emotional elements other than the exemplary techniques described in connection with FIGS. 3A and 3B.

[0039] As described, system 100 in general, and VPA 118 in particular, can be incorporated into a wide variety of different types of systems, including vehicles. Figures 4A-5C illustrate an exemplary scenario in which system 100 is incorporated into a vehicle and the disclosed techniques enable greater usability, thereby preventing the user 140 from having to divert their attention from driving.

[0040] Example VPA Dialogue 4A-4C illustrate examples of how the VPA of FIG. 1 responds to a user's emotional state, according to various embodiments. As shown in FIG. 4A, system 100 is incorporated into vehicle 150, also shown in FIG. 1B, in which user 140 rides. User 140 excitedly exclaims that he or she forgot about a dentist appointment and asks VPA 118 (within system 100) to provide immediate directions to the dentist. VPA 118 analyzes input 122 and detects specific vocal characteristics commonly associated with feelings of excitement and / or anxiety. As shown in FIG. 4B, VPA 118 then generates output 132 consistent with the detected excitement level. Specifically, VPA 118 urgently states that the fastest route has been found and then reassures user 140 that they will arrive in a timely manner. Subsequently, as shown in FIG. 4C, user 140 expresses relief, which VPA 118 processes as feedback 242 to inform future interactions with user 140. In this example, the VPA 118 promotes user engagement by appearing to respond emotionally to the user 140, thereby encouraging the user 140 to continue using the VPA 118 instead of performing various vehicle-specific actions themselves.

[0041] 5A-5C illustrate an example of how the VPA of FIG. 1 compensates for a user's emotional state, according to various embodiments. In this example, the VPA 118 implements an emotional component distinct from the user's 140 emotional state to effect a change in the user's 140 emotional state. As shown in FIG. 5A, the user 140 expresses frustration at having to drive through traffic. The VPA 118 analyzes the input 122 and detects specific vocal characteristics commonly associated with feelings of frustration and / or dissatisfaction. Then, as shown in FIG. 5B, the VPA 118 generates output 132 with emotional qualities significantly different from these specific feelings. Specifically, the VPA 118 expresses disappointment, apologizes, and confesses that there are no other routes available. Subsequently, as shown in FIG. 4C, the user 140 forgets their previous feelings of frustration and comforts the VPA 118, which processes this as feedback 242 to inform future interactions with the user 140. In this example, the VPA 118 promotes user engagement by appearing emotionally sensitive to the user 140, thereby encouraging the user 140 to continue using the VPA 118 instead of performing various vehicle-specific actions themselves.

[0042] Procedure for performing actions depending on the user's emotional state 6 is a flow diagram of method steps for synthesizing vocalizations that reflect a user's emotional state, according to various embodiments. Although the method steps are described with respect to the systems of FIGS. 1-5C, one skilled in the art will understand that any system configured to perform the method steps in any order is within the scope of the present embodiments.

[0043] As shown, method 600 begins at step 602, where VPA 118 captures input indicative of one or more behaviors associated with a user. VPA 118 interacts with input device 120, shown in FIG. 1, to capture the input. Input device 120 may include a variety of different types of sensors configured to capture different types of data associated with the user, including audio capture devices that record vocalizations made by the user, optical capture devices that record images and / or video depicting the user, pupillometry sensors that measure the user's pupil dilation, infrared sensors that measure blood flow in the user's face and / or body, heart rate sensors that generate heart rate readings associated with the user in beats per minute, galvanic skin response sensors that measure changes in the user's skin conductivity, body temperature sensors that detect changes in the user's core and / or surface body temperature, and electroencephalogram (EEG) sensors that detect various brainwave patterns. As described in more detail below, any and all such data may be processed to identify the user's emotional state.

[0044] In step 604, the VPA 118 determines the user's emotional state based on the input. In doing so, the VPA 118 executes the sentiment analyzer 220 to process any of the above types of data and / or map the data and / or a processed version thereof to the user's emotional state. The sentiment analyzer 220 may use any technically feasible approach to define the user's emotional state. In one embodiment, the sentiment analyzer 220 may describe the user's emotional state via a valence-intensity dataset, such as those shown in FIGS. 3A and 3B. Specifically, the sentiment analyzer 220 analyzes certain properties associated with the input, such as the pitch, timbre, tone, and / or volume associated with the user's vocalizations, among others, and then maps those properties to specific locations within the multidimensional valence-intensity space. Those skilled in the art will appreciate that the VPA 118, when performing step 604, may implement any technically feasible approach to generate data indicative of the user's emotional state.

[0045] In step 606, the VPA 118 determines one or more actions to perform on behalf of the user based on the input captured in step 602 and the emotional state determined in step 604. The VPA 118 executes the response generator 230 to process the transcription of the input to determine one or more relevant actions to perform on behalf of the user. For example, if the input corresponds to a command to play music, the VPA 118 may process the transcription of the input and then activate a stereo system in the vehicle in which the user is riding. The VPA 118 may implement any technically feasible approach for generating a transcription of the input, including speech-to-text, among other approaches. In some embodiments, the response generator 230 may also select a set of relevant actions to perform based on the user's emotional state determined in step 606. For example, if the user is in a “sad” emotional state, the VPA 118 may select a particular radio station that plays sad music.

[0046] In step 608, the VPA 118 performs one or more actions on behalf of the user to assist the user in performing those actions. In some embodiments, the VPA 118 may perform one or more actions by executing one or more corresponding subroutines and / or software functions. In other embodiments, the VPA 118 may be integrated with other systems, and the VPA performs one or more actions by causing those systems to perform those actions. For example, as described above, in embodiments in which the VPA 118 is integrated into a vehicle, the VPA 118 may perform a given action to cause one or more subsystems within the vehicle to perform the given action. In such embodiments, the VPA 118 advantageously performs vehicle-related actions on behalf of the user, thereby preventing the user from having to divert their attention from driving to perform those actions themselves.

[0047] In step 610, the VPA 118 synthesizes an output based on the emotional state identified in step 604 and the one or more actions determined in step 606. The VPA 118 executes the response generator 230 to generate an emotional component of the output in addition to a semantic component of the output. The semantic component of the output includes one or more words, phrases, and / or sentences meaningfully related to one or more actions. For example, if a given action is related to a set of navigation instructions for navigating a vehicle, the semantic component of the output may include words related to the initial navigation instructions. In one embodiment, the response generator 230 may generate the semantic component of the output to have emotional characteristics derived from the user's emotional state. The emotional component of the output may exhibit changes in pitch, tone, timbre, and / or volume intended to evoke a particular emotional response from the user, derived from the user's emotional state. The emotional component of the output may also include other factors that affect how the semantic component of the output is conveyed to the user, such as delivery speed and timing. Based on the semantic and emotional elements, the output synthesizer 240 generates an output and sends the output to the user via the output device 130 .

[0048] In step 612, the VPA 118 observes the user's behavior in response to the output synthesized in step 610. The VPA 118 captures any of the types of data described above that describe various behaviors associated with the user to determine how the user reacts to the output. Specifically, the sentiment analyzer 220 can analyze any captured data to determine how the user's emotional state changes in response to the output. For example, without limitation, the sentiment analyzer 220 may determine that a user who was previously in a "frustrated" emotional state has changed to a "relaxed" emotional state in response to output that includes a "soothing" emotional element.

[0049] In step 614, the VPA 118 modifies the response generator 230 and / or the included response mapping 238 based on the observed behavior. In one embodiment, the VPA 118 may execute a mapping modifier 250 to evaluate one or more objective functions 252 to determine whether to modify the response generator 230 and / or the response mapping 238. Each objective function 252 may reflect the user's target behavioral set, target emotional state, target physical state, etc. For example, without limitation, the objective functions 252 may quantify the user's happiness, connectedness, productivity, and / or enjoyment, and the mapping modifier 250 may adjust the response mapping 238 to maximize one or more of these objectives. The mapping modifier 250 may evaluate each objective function 252 to determine the extent to which the observed behavior corresponds to a given target behavior and then modify the response mapping 238 to increase the extent to which the user exhibits the target behavior. The VPA 118 may implement any technically feasible approach for quantifying the manifestation of a given behavior.

[0050] The VPA 118 implements the method 600 to perform some or all of the various features described herein. In some examples, the VPA 118 is described in connection with an in-vehicle implementation, but those skilled in the art will understand how the disclosed techniques provide certain advantages over the prior art in a wide range of technically feasible implementations. As a general matter, the disclosed techniques enable the VPA 118 to more effectively and accurately interpret utterances received from a user, enabling the VPA 118 to generate more realistic, conversational responses to the user, thereby achieving significantly greater usability compared to prior art approaches.

[0051] In summary, a virtual private assistant (VPA) is configured to analyze various types of inputs that indicate one or more behaviors associated with a user. The inputs may include utterances representing explicit commands for the VPA to execute, emotional states derived from these utterances, and implicit nonverbal cues associated with the user, such as changes in facial expression and / or posture, among others. The VPA identifies the user's emotional state based on the input. The VPA also determines one or more actions to perform on behalf of the user based on the input and the identified emotional state. The VPA then performs the one or more actions and synthesizes an output based on the user's emotional state and the one or more actions. The synthesized output includes one or more semantic elements and one or more emotional elements derived from the user's emotional state. The emotional element(s) of the output can be consistent with or contrast with the user's emotional state, among other possibilities. The VPA observes the user's behavior in response to the synthesized output and then implements various modifications based on the observed behavior to improve the effectiveness of future interactions with the user.

[0052] At least one technical advantage of the disclosed techniques over the prior art is that the disclosed techniques enable a VPA to more accurately determine one or more actions to perform on behalf of a user based on the user's emotional state. Thus, when implemented in a vehicle, the disclosed VPA helps prevent a user from diverting their attention from driving to operate vehicle functions, thereby improving overall driving safety. Another technical advantage of the disclosed techniques is that the disclosed techniques enable a VPA to generate realistic, conversational responses that reflect the user's emotional state. The realistic, conversational responses maintain user engagement with the VPA, reducing situations where the user must turn off the VPA to manually operate vehicle functions, improving overall driving safety. These technical advantages represent one or more technical advances over prior art approaches.

[0053] 1. Some embodiments include a computer-implemented method for assisting a user while interacting with the user, the computer-implemented method including: capturing a first input indicative of one or more behaviors associated with the user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; and outputting the first utterance to the user.

[0054] 2. The computer-implemented method described in clause 1, wherein the user is riding in a vehicle in which the first input is captured and the first action is performed on behalf of the user by a vehicle subsystem included in the vehicle.

[0055] 3. The computer-implemented method described in clause 1 or 2, further comprising: determining the first action based on the first emotional state; and performing the first action to assist the user.

[0056] 4. The computer-implemented method described in any one of clauses 1 to 3, wherein identifying the first emotional state of the user includes identifying a first characteristic of the first input and identifying a first emotional type corresponding to the first characteristic.

[0057] 5. The computer-implemented method of any one of clauses 1-4, wherein the first input includes an audio input and the first characteristic includes a tone of voice associated with the user.

[0058] 6. The computer-implemented method of any one of clauses 1-5, wherein the first input includes a video input and the first feature includes a facial expression made by the user.

[0059] 7. The computer-implemented method described in any one of clauses 1 to 6, wherein identifying the first emotional state of the user includes determining, based on the first input, a first valence value indicative of a position within a spectrum of emotional types, and determining, based on the first input, a first intensity value indicative of a position within a range of intensity corresponding to the position within the spectrum of the emotional types.

[0060] 8. The computer-implemented method of any one of clauses 1 to 7, wherein the first emotional element corresponds to the first valence value and the first intensity value.

[0061] 9. The computer-implemented method of any one of clauses 1 to 8, wherein the first emotional element corresponds to at least one of a second valence value or a second intensity value.

[0062] 10. The computer-implemented method of any one of clauses 1 to 9, further comprising generating the first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element.

[0063] 11. Some embodiments include a non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to assist a user while interacting with the user by performing steps including: capturing a first input indicative of one or more behaviors associated with the user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; and outputting the first utterance to the user.

[0064] 12. The non-transitory computer-readable medium of clause 11, wherein the user is riding in a vehicle in which the first input is captured and the first action is performed on behalf of the user by a vehicle subsystem included in the vehicle.

[0065] 13. The non-transitory computer-readable medium of clause 11 or 12, further comprising the steps of determining the first action based on the first emotional state and performing the first action to assist the user.

[0066] 14. A non-transitory computer-readable medium described in any one of clauses 11 to 13, wherein the step of identifying the first emotional state of the user includes identifying a first characteristic of the first input and identifying a first emotional type corresponding to the first characteristic, the first characteristic including a tone of voice associated with the user or a facial expression made by the user.

[0067] 15. A non-transitory computer-readable medium described in any one of clauses 11 to 14, wherein the step of identifying the first emotional state of the user includes determining, based on the first input, a first valence value indicating a position within a spectrum of emotional types, and determining, based on the first input, a first intensity value indicating a position within a range of intensity corresponding to the position within the spectrum of the emotional types.

[0068] 16. A non-transitory computer-readable medium described in any one of clauses 11 to 15, further comprising a step of generating the first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element.

[0069] 17. A non-transitory computer-readable medium described in any one of clauses 11 to 16, further comprising: capturing a second input indicating at least one action that the user will take in response to the output; and modifying the response mapping based on the second input and a first objective function that is evaluated to determine how closely the at least one action corresponds to a target behavior.

[0070] 18. A non-transitory computer-readable medium described in any one of clauses 11 to 17, wherein the step of generating the first utterance includes combining the first emotional element with a first semantic element.

[0071] 19. A non-transitory computer-readable medium described in any one of clauses 11 to 18, further comprising: generating a transcription of the first input indicating one or more semantic elements contained in the first input; and generating the first semantic element based on the one or more semantic elements.

[0072] 20. Some embodiments include a system comprising: a memory storing a software application; and a processor configured to, when executing the software application, perform steps including: capturing a first input indicative of one or more behaviors associated with a user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; and outputting the first utterance to the user.

[0073] Any and all combinations, in any manner, of any claim element recited in any claim and / or any element described in this application are within the intended scope of the present embodiments and protection.

[0074] The descriptions of various embodiments are presented for purposes of illustration and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments.

[0075] Aspects of the present embodiments may be embodied as a system, a method, or a computer program product. Accordingly, aspects of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, all of which may be referred to herein generically as a "module" or "system" or "computer." Furthermore, aspects of the present disclosure may take the form of a computer program product embodied in one or more computer-readable medium(s) having computer-readable program code embodied therein.

[0076] Any combination of one or more computer-readable medium(s) may be utilized. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of computer-readable storage media would include an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that contains or can store a program for use by or in association with an instruction execution system, apparatus, or device.

[0077] Aspects of the present disclosure are described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine. When the instructions are executed by the processor of the computer or other programmable data processing apparatus, the functions / acts specified in one or more blocks of the flowchart illustrations and / or block diagrams can be implemented. Such a processor may be, but is not limited to, a general-purpose processor, a special-purpose processor, an application-specific processor, or a field-programmable gate array.

[0078] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of code, including one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart diagrams, and combinations of blocks in the block diagrams and / or flowchart diagrams, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0079] While the forgoing is directed to embodiments of the present disclosure, other and further embodiments of the present disclosure may be devised without departing from the basic scope thereof, which scope is defined by the appended claims.

Claims

1. 1. A computer-implemented method for assisting a user while interacting with the user, comprising: capturing a first input indicative of one or more actions associated with the user; determining a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; Including, Identifying the first emotional state of the user includes: determining a first valence value indicative of a first location within a spectrum of emotion types based on the first input; determining a first intensity value indicative of a first location within a range of intensities corresponding to the first location within the emotion-type spectrum based on the first input; 11. A computer-implemented method comprising:

2. A computer-implemented method for assisting a user while interacting with the user, comprising: capturing a first input indicative of one or more actions associated with the user; determining a first emotional state of the user based on the first input; generating a first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element; generating a first utterance incorporating the first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; and outputting the first utterance to the user; capturing a second input indicative of at least one action the user performs in response to the output; modifying the response mapping based on the second input and a first objective function that is evaluated to determine how closely the at least one behavior corresponds to a target behavior; 11. A computer-implemented method comprising:

3. the user is in a vehicle in which the first input is captured; The computer-implemented method of claim 1 or claim 2, wherein the first action is performed on behalf of the user by a vehicle subsystem included in the vehicle.

4. determining the first action based on the first emotional state; performing the first action to assist the user; and The computer-implemented method of claim 1 or claim 2, further comprising:

5. Identifying the first emotional state of the user includes: identifying a first feature of the first input; identifying a first emotion type corresponding to the first characteristic; 3. The computer-implemented method of claim 1 or claim 2, comprising:

6. the first input comprises an audio input; The computer-implemented method of claim 5 , wherein the first characteristic comprises a tone of voice associated with the user.

7. the first input comprises a video input; The computer-implemented method of claim 5 , wherein the first feature comprises a facial expression made by the user.

8. The computer-implemented method of claim 1 , wherein the first emotional element corresponds to the first valence value and the first intensity value.

9. 2. The computer-implemented method of claim 1, wherein the first emotional element corresponds to at least one of a second valence value indicating a second location within the emotion type spectrum that is different from the first location within the emotion type spectrum or a second intensity value indicating a second location within the intensity range that is different from the first location within the intensity range and corresponds to the second location within the emotion type spectrum.

10. The computer-implemented method of claim 1 , further comprising generating the first emotional component based on the first emotional state and a response mapping that converts the emotional state into an emotional component.

11. A non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to: capturing a first input indicative of one or more actions associated with a user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; and executing the above steps to assist the user while interacting with the user; The step of identifying the first emotional state of the user comprises: determining a first valence value indicative of a first location within a spectrum of emotion types based on the first input; determining a first intensity value indicative of a first location within a range of intensities corresponding to the first location within the emotion-type spectrum based on the first input; 1. A non-transitory computer-readable medium comprising:

12. A non-transitory computer-readable medium storing program instructions that, when executed by a processor, cause the processor to: capturing a first input indicative of one or more actions associated with a user; identifying a first emotional state of the user based on the first input; generating a first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element; generating a first utterance incorporating the first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; capturing a second input indicative of at least one action the user performs in response to the output; modifying the response mapping based on the second input and a first objective function that is evaluated to determine how closely the at least one behavior corresponds to a target behavior; and a non-transitory computer-readable medium for causing a user to interact with the user and assist the user by executing the method.

13. the user is in a vehicle in which the first input is captured; The non-transitory computer-readable medium of claim 11 or claim 12, wherein the first action is performed on behalf of the user by a vehicle subsystem included in the vehicle.

14. determining the first action based on the first emotional state; performing the first action to assist the user; 13. The non-transitory computer-readable medium of claim 11 or claim 12, further comprising:

15. The step of identifying the first emotional state of the user comprises: identifying a first feature of the first input; identifying a first emotion type corresponding to the first characteristic; Including, The non-transitory computer-readable medium of claim 11 or claim 12, wherein the first characteristic comprises a tone of voice associated with the user or a facial expression made by the user.

16. 12. The non-transitory computer-readable medium of claim 11, further comprising generating the first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element.

17. 13. The non-transitory computer-readable medium of claim 11 or claim 12, wherein generating the first utterance comprises combining the first emotional element with a first semantic element.

18. generating a transcription of the first input indicating one or more semantic elements contained in the first input; generating the first semantic element based on the one or more semantic elements; 20. The non-transitory computer-readable medium of claim 17, further comprising:

19. a memory for storing a software application; a processor; A system comprising: When the processor executes the software application, capturing a first input indicative of one or more actions associated with a user; identifying a first emotional state of the user based on the first input; generating a first utterance incorporating a first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; configured to run The step of identifying the first emotional state of the user comprises: determining a first valence value indicative of a first location within a spectrum of emotion types based on the first input; determining a first intensity value indicative of a first location within a range of intensities corresponding to the first location within the spectrum of emotion types based on the first input; Including, the system.

20. A method for manufacturing a computer-implemented system comprising: a memory for storing a software application; a processor; A system comprising: When the processor executes the software application, capturing a first input indicative of one or more actions associated with a user; identifying a first emotional state of the user based on the first input; generating a first emotional element based on the first emotional state and a response mapping that converts the emotional state into an emotional element; generating a first utterance incorporating the first emotional element based on the first emotional state, the first utterance being associated with a first action being performed to assist the user; outputting the first utterance to the user; capturing a second input indicative of at least one action the user performs in response to the output; modifying the response mapping based on the second input and a first objective function that is evaluated to determine how closely the at least one behavior corresponds to a target behavior; A system configured to run

Citation Information

Patent Citations

  • Voice interacting method and voice interacting device

    JP2001272991A

  • Voice response system

    JP2018136500A

  • Emotion Type Classification for Interactive Dialog Systems

    JP2018503894A

  • Method and system for emotional conversations between humans and machines

    JP2019012255A

  • Electronic personal interactive device

    US20110283190A1