Information provision system, information provision method, and program

The information provision system adjusts anthropomorphic agent voice output based on driver emotional levels to manage vigilance, addressing excessive or insufficient responses to warnings and improving traffic safety.

JP2026087242APending Publication Date: 2026-05-27HONDA MOTOR CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
HONDA MOTOR CO LTD
Filing Date
2024-11-15
Publication Date
2026-05-27

Smart Images

  • Figure 2026087242000001_ABST
    Figure 2026087242000001_ABST
Patent Text Reader

Abstract

This system provides information that can suppress excessive or insufficient vigilance when using an agent-based HMI to alert users while they are on board a mobile vehicle. [Solution] The information provision system 1 includes a user state recognition unit 11 that recognizes the state of user U riding in the mobile body 100, a user emotion estimation unit 12 that estimates user U's emotion level based on predetermined index values ​​based on the recognition result of user U's state, and an agent control unit 13 that provides user U with an HMI that uses an anthropomorphic agent to output voice in a manner corresponding to the emotion level of the agent based on predetermined index values. When the emotion level of user U estimated by the user emotion estimation unit 12 is within a predetermined judgment range, the agent control unit 13 controls the agent's emotion level to a target level that changes user U's emotion outside the judgment range.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information providing system, an information providing method, and a program.

Background Art

[0002] In recent years, technologies have been proposed for mounting anthropomorphic agents that communicate with users on the HMI (Human Machine Interface) of information providing devices used in moving objects such as vehicles. For example, Patent Document 1 discloses a technology for estimating the emotions of a vehicle driver and outputting response contents that can change the driver's emotions to negative emotions or positive emotions below a predetermined level according to the level of negative emotions when the driver's emotions are negative emotions that do not desire responses from others. Further, Patent Document 2 discloses a technology for starting the output of dialogue voice when a decrease in the driver's arousal level is detected and stopping the output of dialogue voice when a return from the decrease in the driver's arousal level is detected.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] Incidentally, when a vehicle's driver assistance function provides warnings to the driver via voice output from an HMI agent to avoid risks such as collisions between the vehicle and obstacles, it is conceivable to change the manner of voice output by the agent to enhance the effectiveness of the warnings. However, based on the principles of behavioral imitation and emotion transmission, the actions (such as speaking speed) and emotions of the agent with whom the driver is interacting are unconsciously imitated (such as operating speed) and transmitted to the driver. Therefore, depending on the manner of voice output from the agent, changes in the driver's emotions after receiving the warning may lead to excessive or insufficient vigilance towards the driver. Therefore, the objective of this invention is to provide an information provision system that can suppress excessive or insufficient vigilance when alerting passengers on board using an agent-based HMI during transit. This will ultimately contribute to further improving traffic safety and the development of a sustainable transportation system. [Means for solving the problem]

[0005] As a first embodiment for achieving the above objective, an information providing device is provided, comprising: a user state recognition unit that recognizes the state of a user riding in a mobile vehicle; a user emotion estimation unit that estimates the user's emotional level based on a predetermined index value based on the recognition result of the user state recognition unit; and an agent control unit that provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the emotional level of the agent based on the predetermined index value, wherein the agent control unit controls the emotional level of the agent to a target level that changes the user's emotional level outside the predetermined judgment range when the emotional level of the user estimated by the user emotion estimation unit is within a predetermined judgment range.

[0006] The above-described information providing device may be configured to include a risk recognition unit that recognizes a predetermined risk facing the moving object, and the agent control unit may set the determination range according to the predetermined risk when the risk recognition unit recognizes the predetermined risk.

[0007] In the above-described information providing device, the predetermined index value is an emotional valence indicating pleasure and displeasure, and the agent control unit may be configured to control the emotional level of the agent to the target level by changing the fundamental frequency or fluctuation range of the voice output by the agent when the emotional level of the user estimated by the user emotion estimation unit is within the determination range.

[0008] In the above-described information providing device, the predetermined index value is an emotional valence indicating pleasure and displeasure, and the agent control unit may be configured to control the agent's emotional level to the target level by raising or lowering the ending of the utterance output by the agent when the user's emotional level estimated by the user emotional estimation unit is within the determination range.

[0009] In the above-described information providing device, the predetermined index value is an arousal level indicating excitement and calmness, and the agent control unit may be configured to control the agent's emotional level to the target level by changing the speech rate or volume of the voice output by the agent when the user's emotional level estimated by the user emotional estimation unit is within the determination range.

[0010] As a second embodiment for achieving the above objective, there is an information provision method performed by a computer, which includes: a user state recognition step of recognizing the state of a user riding in a mobile vehicle; a user emotion estimation step of estimating the user's emotion level based on a predetermined index value based on the recognition result of the user state recognition step; and an agent control step of providing the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the agent's emotion level based on the predetermined index value, wherein the agent control step controls the agent's emotion level to a target level that changes the user's emotion level outside the predetermined judgment range when the user's emotion level estimated by the user emotion estimation step is within a predetermined judgment range.

[0011] As a third embodiment for achieving the above objective, the computer is configured to function as: a user state recognition unit that recognizes the state of a user riding in a mobile vehicle; a user emotion estimation unit that estimates the user's emotional level based on a predetermined index value, based on the recognition result of the user state recognition unit; and an agent control unit that provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the agent's emotional level based on the predetermined index value. The agent control unit is configured to control the agent's emotional level to a target level that changes the user's emotional level outside the predetermined range when the user's emotional level estimated by the user emotion estimation unit is within a predetermined range. [Effects of the Invention]

[0012] According to the above information provision system, information provision method, and program, it is possible to suppress excessive or insufficient vigilance when issuing warnings to users on board a mobile vehicle using an agent-based HMI. [Brief explanation of the drawing]

[0013] [Figure 1] Figure 1 is an explanatory diagram of how the information provision system is used. [Figure 2] Figure 2 is a diagram showing the configuration of the information provision system. [Figure 3] Figure 3 is an explanatory diagram of the emotion map based on emotional valence and arousal level. [Figure 4] Figure 4 is an explanatory diagram of the characteristics of each region in the emotion map. [Figure 5] Figure 5 is an explanatory diagram of the parameters that change emotional valence and arousal level. [Figure 6] Figure 6 is the first explanatory diagram of the parameters that change the characteristics of the voice. [Figure 7] Figure 7 is a second explanatory diagram of the parameters that change the characteristics of the voice. [Figure 8] Figure 8 is a flowchart of the agent's emotion control process. [Figure 9] Figure 9 is an explanatory diagram illustrating the control mechanism of the agent's emotional level. [Modes for carrying out the invention]

[0014] [1. Usage patterns of the information provision system] Referring to Figure 1, the usage of the information provision system 1 of this embodiment will be described. The information provision system 1 provides a Human Machine Interface (HMI) using an anthropomorphic agent with emotions to the user U, who is the driver of the vehicle 100. The information provision system 1 may be configured as part of the functions of an in-vehicle device such as a navigation device installed in the vehicle 100, or it may be configured as a dedicated in-vehicle device. The vehicle 100 corresponds to the mobile body of this disclosure.

[0015] The information providing system 1 has a communication function and performs wireless communication with a wearable device 70 (such as a smartwatch) worn by the user U and a portable terminal 60 (such as a smartphone, tablet terminal, mobile phone, etc.) brought into the vehicle 100 by the user U. Further, the information providing system 1 communicates with an external communication system such as a content server 210 via a communication network 200.

[0016] The vehicle 100 is equipped with a driver monitoring camera 54 for photographing the user U, a touch panel type display 55, a speaker 56, a microphone 57 for collecting the speech of the user U, a front camera 50 for photographing the front of the vehicle 100, an ADAS (Advanced Driver-Assistance Systems) 2 for performing various driving support controls based on the photographed image of the front camera 50, and a navigation device 3 for guiding the route to the destination.

[0017] The agent installed in the information providing system 1 is an interface agent configured by, for example, an AI (Artificial Intelligence) program. The agent is an interactive interface that analyzes the speech of the user U input to the microphone 57 by speech recognition and outputs information by voice output from the speaker 56 or image display on the display 55 according to the speech content. The agent makes proposals for the destination of route guidance by the navigation device 3, proposals for content to be downloaded and played back from the content server 210, and alerts for a predetermined risk (such as approaching an obstacle) recognized by the ADAS 2 in response to the speech of the user U.

[0018] The information providing system 1 estimates the emotion of the user U based on the photographed image of the user U by the driver monitoring camera 54, the voice of the user U input to the microphone 57, the biometric information (such as heart rate, body temperature, blood pressure, etc.) of the user U detected by the wearable device 70, and the like.

[0019] The information provision system 1 sets the agent's emotion based on the user U's emotions, and the agent controls the characteristics of the voice output from speaker 56 (speech content, intonation, pitch, volume, etc.) according to the set emotion. Furthermore, as will be described in detail later, the information provision system 1 controls the agent's emotion level according to the user U's emotion level in order to make the alerting to the user U more appropriate.

[0020] [2. Configuration of the Information Provision System] Referring to Figures 2 to 7, the configuration of the information provision system 1 will be described. Referring to Figure 2, the information provision system 1 is a control unit equipped with a processor 10, memory 20, communication unit 30, etc. The information provision system 1 is connected to the ADAS 2, speed sensor 51, gyro sensor 52, GNSS (Global Navigation Satellite System) sensor 53, driver monitor camera 54, display 55, speaker 56, and microphone 57 mounted on the vehicle 100.

[0021] The communication unit 30 performs short-range wireless communication with the mobile terminal 60 and the wearable device 70 using specifications such as Bluetooth® and UWB (Ultra Wide Band). The communication unit 30 also communicates with the content server 210 and the like via the communication network 200.

[0022] Memory 20 stores the control program 21 for the information provision system 1, data for the emotion map 22, and other such information. Details of the emotion map 22 will be described later. By reading and executing the program 21, the processor 10 functions as the user state recognition unit 11, the user emotion estimation unit 12, the agent control unit 13, and the risk recognition unit 14.

[0023] The processing performed by the user state recognition unit 11 corresponds to the user state recognition step in the information provision method of this disclosure, and the processing performed by the user emotion estimation unit 12 corresponds to the user emotion estimation step in the information provision method of this disclosure. The processing performed by the agent control unit 13 corresponds to the agent control step in the information provision method of this disclosure.

[0024] The user state recognition unit 11 recognizes the state of user U, who is riding in the vehicle 100, based on at least one of the following: information input via touch operation on the display 55, an image of user U captured by the driver monitor camera 54, the user's voice input into the microphone 57, or biometric information of user U detected by the wearable device 70.

[0025] The user state recognition unit 11 recognizes, for example, the following elements as the state of user U. Element 1: The user U's response to a question such as "How are you feeling today?" displayed on the display 55 when the user U boards the vehicle 100, which is entered via touch operation on the display 55. Second element: User U's facial expressions and behavior as recognized from User U's image. Third element: User U's speech content, intonation, pitch, volume, and intonation. Fourth element: User U's biometric information (heart rate, blood pressure, body temperature, etc.).

[0026] The user emotion estimation unit 12 estimates the emotional level of user U based on the state of user U recognized by the user state recognition unit 11. The user emotion estimation unit 12 estimates the emotional level of user U using the emotion map 22 shown in Figure 3, with emotional valence and arousal level as indicator values ​​for the emotional level. Emotional valence is an indicator value indicating the level of discomfort to pleasure for user U, and arousal level is an indicator value indicating the level of excitement to calmness for user U. Emotional valence and arousal level correspond to predetermined indicator values ​​in this disclosure.

[0027] In regions A, B, C, and D of the emotion map 22, which are demarcated by the emotional valence axis and the arousal axis, the speech sounds differ in the emotional valence parameters of the fundamental frequency and the range of variation of the fundamental frequency, the arousal parameters of speech speed and sound pressure (volume), and the speech pattern parameter of rising and falling intonation at the end of words, as shown in Figure 4.

[0028] The user emotion estimation unit 12 estimates the user U's emotions by, for example, setting the emotional valence and arousal levels based on the first to fourth elements as follows.

[0029] Regarding the first element described above, the user emotion estimation unit 12 sets the emotional valence of user U to a relatively high value (for example, 3 out of ±10) if the response to the question is "as usual," and sets the arousal level to a relatively low value (for example, -3 out of ±10) if the response to the question is "different from usual."

[0030] Regarding the second element described above, the user emotion estimation unit 12 determines, based on the user U's facial expression, that if user U has a smiling expression, it sets the emotional valence to a relatively high value (for example, 5 out of ±10), and if user U has a sulky expression, it sets the emotional valence to a relatively low value (for example, -5 out of ±10). Furthermore, based on the user U's behavior, the user emotion estimation unit 12 determines, based on the user U's behavior, that user U is restless, it sets the arousal level to a relatively high value (for example, 3 out of ±10), and if user U is hardly moving, it sets the arousal level to a relatively low value (for example, -3 out of ±10).

[0031] Regarding the third element described above, the user emotion estimation unit 12, based on the content of user U's utterance, sets the emotional valence to a relatively high value (for example, 5 out of ±10) if it determines that the utterance is positive, such as praising or expressing expectations, and sets the emotional valence to a relatively low value (for example, -5 out of ±10) if it determines that the utterance is negative, such as criticizing something. In addition, if the user emotion estimation unit 12 includes specific keywords in user U's utterance (such as "like" or "super like"), it sets the emotional valence and arousal level values ​​associated with those keywords.

[0032] Furthermore, the user emotion estimation unit 12, based on the pitch of user U's voice, sets the arousal level to a relatively high value (e.g., 3 out of ±10) if user U's voice is above a predetermined pitch, and sets the arousal level to a relatively low value (e.g., -3 out of ±10) if user U's voice is below the predetermined pitch.

[0033] Regarding the fourth element described above, the user emotion estimation unit 12 applies the detected values ​​of user U's biometric information (heart rate, blood pressure, body temperature, etc.) to a pre-prepared correspondence table of detected values, emotional valence, and arousal level to set the emotional valence and arousal level. The correspondence table may be created based on user U's profile entered by user U or user U's biometric information detected in the past.

[0034] Furthermore, the user emotion estimation unit 12 may recognize values ​​for emotional valence and arousal based on changes in the vehicle's speed detected by the speed sensor 51, the vehicle's behavior detected by the gyro sensor 52, and fluctuations in the vehicle's position detected by the GNSS sensor 53. Alternatively, the user emotion estimation unit 12 may estimate the user U's emotions based on either emotional valence or arousal, or on other indicator values.

[0035] The agent control unit 13 provides the user U with an HMI (Human-Machine Interface) provided by an anthropomorphic agent with emotions. The agent control unit 13 analyzes the user U's speech input to the microphone 57 and outputs a response to the speech by outputting the agent's voice from the speaker 56 or displaying the agent's image on the display 55. The agent control unit 13 changes the agent's emotions according to the emotions of the user U estimated by the user emotion estimation unit 12, and changes the output characteristics (intonation, pitch, volume, etc.) of the agent's voice output from the speaker 56 according to the agent's emotions.

[0036] The agent control unit 13 controls the agent's emotional level by using the fundamental frequency (key), the variation range of the fundamental frequency, the speech rate, the volume (sound pressure), and the text content as parameters to change the emotional expression of the voice output by the agent, and by changing these parameters in combination.

[0037] Figures 5 to 7 illustrate the conditions under which the emotional expression of the agent's output voice can be changed by modifying the above parameters. First, Figure 5 shows the control conditions obtained by changing the fundamental frequency, the amplitude of the fundamental frequency vibration, the speech rate, and the volume (sound pressure). As shown in Figure 5, the agent control unit 13 controls the emotional valence in a positive direction by increasing the fundamental frequency or increasing the amplitude of the fundamental frequency vibration. Conversely, the agent control unit 13 controls the emotional valence in an unpleasant direction by decreasing the fundamental frequency or decreasing the amplitude of the fundamental frequency vibration.

[0038] Furthermore, as shown in Figure 5, the agent control unit 13 controls the level of arousal in the direction of excitement by increasing the speech rate or increasing the volume. The agent control unit 13 also controls the level of arousal in the direction of sedation by decreasing the speech rate or decreasing the volume (sound pressure).

[0039] Next, Figure 6 shows the control conditions obtained by changing the fundamental frequency, the variation range of the fundamental frequency, and the speech rate. In Figure 6, the speech text "Drive slowly, there are children" is shown in V11 when output with cheerful voice characteristics, and in V12 when output with frightened voice characteristics. In V11 and V12, the horizontal axis is set to time and the vertical axis is set to frequency.

[0040] In V11, the parameters are set as follows: The fundamental frequency is relatively high. The variation in the fundamental frequency is large. Speech speed...slow. In V12, the parameters are set as follows: Basic frequency... relatively low. The variation in the fundamental frequency is small. Speech speed...fast.

[0041] Next, Figure 7 shows the control conditions by changing the volume and text content. In Figure 7, V21 shows the case where output is generated using cheerful voice characteristics, and V22 shows the case where output is generated using frightened voice characteristics. In V21, the parameters are set as follows: Spoken text: "Drive slowly, there are children." The volume is... too loud. The ending of a sentence... rises in pitch. In V22, the parameters are set as follows: Spoken text: "D-Driving slowly, there are children." The volume is... too low. The ending of a sentence...falls in pitch.

[0042] The risk recognition unit 14 recognizes a predetermined risk facing the vehicle 100 based on risk information input from the ADAS2. An example of a predetermined risk facing the vehicle 100 is approaching a monitored object (other vehicles, pedestrians, obstacles, etc.) in the vicinity of the vehicle 100. As the distance between the vehicle 100 and the monitored object decreases, the degree of risk recognized by the risk recognition unit 14 increases. Alternatively, the risk recognition unit 14 may recognize the degree of concentration of the user U on driving from the user U's facial image captured by the driver monitor camera 54 or the user U's biometric information detected by the wearable device 70, and recognize a state where the degree of concentration has fallen below a predetermined level as a predetermined risk. The risk recognition unit 14 may also use a front camera, inter-vehicle communication device, GNSS unit, V2X (vehicle-to-infrastructure or pedestrian-to-vehicle) communication, or a server to detect objects and their relative positions, and calculate the degree of risk.

[0043] [3. Agent's emotional control processing] The procedure for emotion control processing of the HMI agent, performed by the information provision system 1, will be explained with reference to Figure 9, following the flowchart shown in Figure 8. When the information provision system 1 recognizes that user U has boarded vehicle 100 and turned on the power, it starts the processing according to the flowchart shown in Figure 8.

[0044] In step S1 of Figure 8, the user emotion estimation unit 12 estimates the user's emotions during a predetermined period from the time of boarding, based on the state of user U recognized by the user state recognition unit 11. In the subsequent step S2, the agent control unit 13 sets a base emotion, which is the basis for the emotions to be set for the agent, based on the emotions of user U estimated by the user emotion estimation unit 12 in step S1.

[0045] In the next step S3, the agent control unit 13 sets a determination range for user U's emotion level based on the risk situation recognized by the risk recognition unit 14. Here, Figure 9 illustrates an example of controlling the agent's emotion level so that when the emotion level of user U estimated by the user emotion estimation unit 12 is within the determination range of the emotion level set according to the risk situation, user U's emotion falls outside the determination range.

[0046] Figure 9 shows that E1 represents the case where user U's emotional level is P11 within the judgment range R1 indicating "anxiety" in area C of the emotional map 22, E2 represents the case where user U's emotional level is P21 within the judgment range R2 indicating "anger" in area C of the emotional map 22, and E3 represents the case where user U's emotional level is P31 within the judgment range R3 indicating "drowsiness, fatigue, and boredom" in area D of the emotional map 22.

[0047] The agent control unit 13 sets a determination range R1 when the risk recognition unit 14 recognizes that the vehicle 100 is traveling at high speed, sets a determination range R2 when the risk recognition unit 14 recognizes that the vehicle 100 is traveling on a congested road, and sets a determination range R3 when the risk recognition unit 14 recognizes that the user U's level of alertness is below a predetermined level.

[0048] In step S4, the user emotion estimation unit 12 estimates the emotion level of user U based on the state of user U recognized by the user state recognition unit 11. In the following step S5, the agent control unit 13 determines whether or not the emotion level of user U is within the determination range. If the emotion level of user U is within the determination range, the agent control unit 13 proceeds to step S20; if the emotion level of user U is outside the determination range, the agent control unit 13 proceeds to step S6.

[0049] In step S20, the agent control unit 13 sets the agent's emotion level to a target level outside the judgment range and proceeds to step S7. Also, in step S6, the agent control unit 13 sets the agent's emotion level based on the base emotion set in step S2 and the emotion of user U estimated in step S4, and proceeds to step S7. In step S7, the agent control unit 13 controls the manner of the agent's voice output according to the emotion level set in step S6 or step S20.

[0050] In Figure 9, at E1, the target level of the agent's emotional level is set to P12, which corresponds to "calm" in area B, outside the range of judgment range R1. The agent then outputs the text, "It would be terrible if we had an accident, so let's go slowly!" in a calm voice. As a result, it is expected that, through the principle of behavioral imitation, user U's emotions will change to a calm state, mimicking the agent's emotions. And as user U's emotions change to a "calm state," user U's response to driving operations will slow down, and the urge to go quickly will be suppressed.

[0051] Furthermore, in Figure 9, at E2, the agent's target emotional level is set to P22, which corresponds to "calm" in area B, outside the judgment range R2. The agent then outputs the text, "It's always so congested here, it's really difficult," in a calm voice. This is expected to cause user U's emotions to change to a calm state, mimicking the agent's emotions, through the principle of behavioral imitation. As user U's emotions change to a "calm state," their frustration is reduced, their driving response becomes slower, and their desire to go quickly is curbed. Furthermore, if the agent's emotion is set to correspond to "anger," it may express empathy and sympathy for user U through angry voice output, and then, if the agent's emotion is set to correspond to "calm," it may change to a calm voice output.

[0052] Furthermore, in E3 of Figure 9, the agent's target emotional level is set to P32, which corresponds to "happy" in area A, outside the range of judgment range R3. The agent then outputs the text "Just a little more, let's enjoy it to the end" in a happy voice. This is expected to cause user U's emotions to become happy, mimicking the agent's emotions, through the principle of behavioral imitation. As a result, user U's emotions change to a "happy state," which helps maintain user U's alertness and prevents a decline in user U's concentration. Alternatively, the agent's emotion could be set to correspond to "tiredness," expressing empathy and understanding for user U through tired-sounding voice output, and then the agent's emotion could be set to correspond to "happy," changing the agent's voice output to a happy-sounding voice.

[0053] In the following step S8, the agent control unit 13 determines whether or not the vehicle 100 has been turned off. If the power has not been turned off, the agent control unit 13 proceeds to step S3 and repeats the processes in steps S3-S8 and S20. On the other hand, if the power has been turned off, the agent control unit 13 proceeds to step S9 and terminates the agent's emotion control process.

[0054] [4. Other Embodiments] In the above embodiment, a vehicle 100 was shown as the mobile body of this disclosure, but the mobile body of this disclosure may be any mobile body that provides an HMI using an agent to a user while on board, and may be an aircraft, a ship, or the like.

[0055] In the above embodiment, an information provision system 1 configured with an in-vehicle device mounted on a vehicle 100 was given as an example of the information provision system of the present disclosure. In other embodiments, part or all of the information provision system of the present disclosure may be configured with a mobile terminal 60 brought into the vehicle 100, or an external communication system such as a content server 210.

[0056] When the information provision system is configured using the mobile terminal 60, the HMI using an agent can be provided using the camera, microphone, speaker, and display installed in the mobile terminal 60. In addition, biometric information of user U may be acquired through communication between the wearable device 70 and the mobile terminal 60, and the mobile terminal 60 may estimate the emotions of user U based on the biometric information. Furthermore, the mobile terminal 60 may acquire information on predetermined risks recognized by ADAS2 through communication between the in-vehicle unit of the vehicle 100 and the mobile terminal 60.

[0057] Furthermore, if the information provision system is configured using an external communication system, for example, the vehicle 100 transmits images of user U captured by the driver monitoring camera 54 of the vehicle 100, audio data input to the microphone 57, and information on predetermined risks recognized by the ADAS 2 to the external communication system, and the external communication system sets the agent's emotions.

[0058] In the above embodiment, a risk recognition unit 14 is provided, and the agent control unit 13 sets a determination range for determining the emotional level of user U according to the risk recognized by the risk recognition unit 14. In another embodiment, the risk recognition unit 14 may be omitted, and the determination range may be a range pre-set based on assumptions such as anger, anxiety, sleepiness, etc.

[0059] Figure 2 is a schematic diagram showing the configuration of the information provision system 1 divided by its main processing content, in order to facilitate understanding of the present invention. The configuration of the information provision system 1 may be composed of other divisions. Furthermore, the processing of each component may be performed by one hardware unit or by multiple hardware units. Also, the processing of each component shown in Figure 5 may be performed by one program or by multiple programs.

[0060] [5. Configurations supported by the above embodiment] The above embodiment is a specific example of the following configuration.

[0061] (Configuration 1) An information providing device comprising: a user state recognition unit that recognizes the state of a user riding in a mobile vehicle; a user emotion estimation unit that estimates the user's emotional level based on a predetermined index value, based on the recognition result of the user state recognition unit; and an agent control unit that provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the emotional level of the agent based on the predetermined index value, wherein the agent control unit controls the emotional level of the agent to a target level that changes the user's emotional level outside the predetermined judgment range when the emotional level of the user estimated by the user emotion estimation unit is within a predetermined judgment range.

[0062] According to the information provision device in Configuration 1, when alerting users on board a mobile vehicle using an agent-based HMI, it is possible to suppress excessive or insufficient vigilance.

[0063] (Configuration 2) The information providing device according to Configuration 1, comprising a risk recognition unit that recognizes a predetermined risk facing the moving body, wherein the agent control unit sets the determination range according to the predetermined risk when the risk recognition unit recognizes the predetermined risk.

[0064] According to the information provision device in configuration 2, by setting a judgment range according to a predetermined type of risk, the emotions of the agent when the agent delivers a warning voice message regarding a predetermined risk can be appropriately controlled.

[0065] (Configuration 3) The information providing device according to Configuration 1 or Configuration 2, wherein the predetermined index value is an emotional valence indicating pleasure and displeasure, and the agent control unit controls the emotional level of the agent to the target level by changing the fundamental frequency or fluctuation range of the voice output by the agent when the emotional level of the user estimated by the user emotion estimation unit is within the determination range, wherein the agent control unit controls the emotional level of the agent to the target level by changing the fundamental frequency or fluctuation range of the voice output by the agent.

[0066] According to the information-providing device of configuration 3, when emotional valence is used as a predetermined index value for the emotional level, the agent's emotional level can be controlled to a target level by changing the fundamental frequency or fluctuation range of the output voice.

[0067] (Configuration 4) An information providing device according to any one of Configurations 1 to 3, wherein the predetermined index value is an emotional valence indicating pleasure and displeasure, and the agent control unit controls the emotional level of the agent to the target level by raising or lowering the ending of the utterance output by the agent when the emotional level of the user estimated by the user emotional estimation unit is within the determination range.

[0068] According to the information-providing device of configuration 4, when using emotional valence as a predetermined index value for emotional level, the agent's emotional level can be controlled to a target level by changing the speech rate or volume of the outputted voice.

[0069] (Configuration 5) An information providing device according to any one of Configurations 1 to 4, wherein the predetermined index value is an arousal level indicating excitement and calmness, and the agent control unit controls the agent's emotional level to the target level by changing the speech rate or volume of the voice output by the agent when the user's emotional level estimated by the user emotional estimation unit is within the determination range.

[0070] According to the information-providing device of configuration 5, when the level of arousal is used as a predetermined index value for the emotional level, the emotional level of the agent can be controlled to the target level by changing the speech rate or volume of the output voice.

[0071] (Configuration 6) Information provision method performed by a computer, comprising: a user state recognition step of recognizing the state of a user riding in a mobile vehicle; a user emotion estimation step of estimating the user's emotion level based on a predetermined index value based on the recognition result of the user state by the user state recognition step; and an agent control step of providing the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the agent's emotion level based on the predetermined index value, wherein the agent control step controls the agent's emotion level to a target level to change the user's emotion level outside the predetermined judgment range when the user's emotion level estimated by the user emotion estimation step is within a predetermined judgment range.

[0072] By executing the information provision method of Configuration 6 using a computer, the same effects and benefits as those of the information provision device of Configuration 1 can be obtained.

[0073] (Configuration 7) A computer comprising: a user state recognition unit that recognizes the state of a user riding in a mobile vehicle; a user emotion estimation unit that estimates the user's emotional level based on a predetermined index value, based on the recognition result of the user state recognition unit; and an agent control unit that provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the agent's emotional level based on the predetermined index value, wherein the agent control unit is a program that controls the agent's emotional level to a target level to change the user's emotional level outside the predetermined judgment range when the user's emotional level estimated by the user emotion estimation unit is within a predetermined judgment range.

[0074] By executing the program of Configuration 7 on a computer, the configuration of the information providing device of Configuration 1 can be realized. [Explanation of Symbols]

[0075] 1...Information provision system, 2...ADAS, 10...Processor, 11...User state recognition unit, 12...User emotion estimation unit, 13...Agent control unit, 14...Risk recognition unit, 20...Memory, 21...Program, 22...Emotion map, 30...Communication unit, 50...Front camera, 51...Speed ​​sensor, 52...Gyro sensor, 53...GNSS sensor, 54...Driver monitor camera, 55...Display, 56...Speaker, 57...Microphone, 60...Mobile terminal, 70...Wearable device, 100...Vehicle (mobile object), 200...Communication network, 210...Content server, U...User.

Claims

1. A user status recognition unit that recognizes the status of the user riding in the mobile vehicle, A user emotion estimation unit estimates the user's emotional level based on a predetermined index value, based on the user state recognition result of the user state recognition unit. An agent control unit provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the emotion level of the agent based on the predetermined index value of the agent, Equipped with, The agent control unit controls the agent's emotional level to a target level that causes the user's emotional level to move outside the predetermined range when the user's emotional level estimated by the user's emotional level estimation unit is within a predetermined determination range. Information provision device.

2. The moving body is equipped with a risk recognition unit that recognizes a predetermined risk it is facing, The agent control unit sets the determination range according to the predetermined risk when the risk recognition unit recognizes the predetermined risk. The information providing device according to claim 1.

3. The aforementioned predetermined index value is an emotional valence indicating pleasure and displeasure. The agent control unit controls the emotional level of the agent to the target level by changing the fundamental frequency or fluctuation range of the voice output by the agent when the emotional level of the user estimated by the user emotional estimation unit is within the determination range. The information providing device according to claim 1 or claim 2.

4. The aforementioned predetermined index value is an emotional valence indicating pleasure and displeasure. The agent control unit controls the emotional level of the agent to the target level by raising or lowering the ending of the utterance output by the agent when the emotional level of the user estimated by the user emotional estimation unit is within the determination range. The information providing device according to claim 1 or claim 2.

5. The aforementioned predetermined index value is a level of alertness indicating excitement and sedation, The agent control unit controls the emotional level of the agent to the target level by changing the speech rate or volume of the voice output by the agent when the emotional level of the user estimated by the user emotional estimation unit is within the determination range. The information providing device according to claim 1 or claim 2.

6. A method of providing information performed by a computer, A user status recognition step that recognizes the status of the user riding in the mobile vehicle, A user emotion estimation step, which estimates the user's emotional level based on a predetermined index value, based on the results of the user state recognition step, An agent control step provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the emotion level based on the predetermined index value of the agent, Includes, The agent control step controls the agent's emotional level to a target level that causes the user's emotional level to move outside the predetermined range, when the user's emotional level estimated by the user emotional level estimation step is within a predetermined range. Information provision method.

7. Computers, A user status recognition unit that recognizes the status of the user riding in the mobile vehicle, A user emotion estimation unit estimates the user's emotional level based on a predetermined index value, based on the user state recognition result of the user state recognition unit. An agent control unit provides the user with an HMI (Human Machine Interface) that uses an anthropomorphic agent to output voice in a manner corresponding to the emotion level of the agent based on the predetermined index value of the agent, and make it work The agent control unit controls the agent's emotional level to a target level that causes the user's emotional level to move outside the predetermined range when the user's emotional level estimated by the user's emotional level estimation unit is within a predetermined determination range. program.