Automatic response device, automatic response method, and computer program for automatic response
The automatic response device addresses the lack of emotional reflection in LLMs by estimating occupants' emotions and generating responses that consider their emotional states, enhancing vehicle interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2026-03-31
AI Technical Summary
Existing large language models (LLMs) do not reflect the emotion of the user when generating responses to questions, as only text data representing the content of the question is input, lacking emotional context.
An automatic response device that estimates the emotions of vehicle occupants using in-vehicle sensors and generates responses by inputting these emotions, along with voice information, into a pre-trained generation model to create emotionally reflective outputs.
The system effectively responds to occupants' utterances by incorporating their emotions, enabling personalized and emotionally sensitive interactions within vehicles.
Smart Images

Figure 2026055535000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an automatic response device, an automatic response method, and a computer program for automatic response that automatically respond to utterances from vehicle occupants.
Background Art
[0002] Techniques for appropriately controlling a vehicle with respect to the emotions of vehicle occupants other than the driver have been proposed (see Patent Document 1). The vehicle control device disclosed in Patent Document 1 estimates whether the emotion of an occupant, who is a passenger other than the driver among the passengers in the vehicle, is unpleasant, and when it is estimated that the emotion of the occupant is unpleasant, executes vehicle control based on the preference data of the occupant. Further, when there is no preference data of the occupant, this vehicle control device executes a monitoring process that combines vehicle control and emotion estimation, and acquires the preference data of the occupant based on the result of the monitoring process.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] On the other hand, large language models (LLMs) that automatically generate responses to questions are being studied. However, since only text data representing the content of the question is input to the LLM, the emotion of the user who issued the question is not reflected in the generated response.
[0005] Therefore, an object of the present invention is to provide an automatic response device that can respond to utterances from occupants in a vehicle with a response that reflects the emotion of the occupants.
Means for Solving the Problems
[0006] According to one embodiment, an automatic response device is provided. This automatic response device includes an emotion estimation unit that estimates the emotions of individual occupants of a vehicle based on in-vehicle sensor signals obtained by in-vehicle sensors that detect the conditions inside the vehicle's cabin, and a response generation unit that generates response information by inputting the estimated emotions of each occupant and voice information representing the content of an utterance from any of the occupants into a generation model that has been pre-trained to generate response information for the content of an utterance.
[0007] In one embodiment, the response generation unit generates response information by further inputting the estimated emotions of individual occupants, along with location information representing the occupants' positions, into a generation model.
[0008] In one embodiment, the automatic response device further includes a storage processing unit that, each time response information is generated, stores in a memory unit the combination of the response information, the estimated emotions of each occupant input into the generation model at the time of response information generation, and the voice information. The response generation unit then generates response information by further inputting into the generation model the combination of the estimated emotions of each occupant and the voice information stored in the memory unit for the most recent predetermined period.
[0009] Another embodiment provides an automated response method. This automated response method includes estimating the emotions of individual occupants of a vehicle based on in-vehicle sensor signals obtained by in-vehicle sensors that detect the conditions inside the vehicle's cabin, and generating response information by inputting the estimated emotions of each occupant and speech information representing the content of an utterance from any of the occupants into a generative model that has been pre-trained to generate response information for the content of an utterance.
[0010] In yet another embodiment, an automated response computer program is provided. This automated response computer program includes instructions for a computer to perform the following actions: estimate the emotions of individual occupants of a vehicle based on in-vehicle sensor signals obtained by in-vehicle sensors that detect the conditions inside the vehicle's cabin; and generate response information by inputting the estimated emotions of each occupant and speech information representing the content of an utterance from any of the occupants into a generative model that has been pre-trained to generate response information for the content of an utterance. [Effects of the Invention]
[0011] The automated response system described herein has the effect of being able to respond to speech from occupants in a vehicle in a way that reflects the occupants' emotions. [Brief explanation of the drawing]
[0012] [Figure 1] This is a schematic diagram of a vehicle equipped with an automatic response system. [Figure 2] This is a hardware configuration diagram of an automated response system. [Figure 3] This is a functional block diagram of the processor in an automatic response system. [Figure 4] This diagram illustrates the relationship between the input and response information of the generative model according to this embodiment. [Figure 5] This is an operation flowchart of an automated response system. [Figure 6] This is a modified functional block diagram of the processor in an automatic response device. [Figure 7] This is a modified flowchart of the operation of an automated response system. [Modes for carrying out the invention]
[0013] The following describes the automated response system, the automated response method performed by the automated response system, and the computer program for automated response, with reference to the diagram. This automated response system estimates the emotions of each occupant of the vehicle and generates response information by inputting the estimated emotions of each occupant and the audio information representing the content of an utterance from one of the occupants into a generative model that has been pre-trained to generate response information for the content of the utterance.
[0014] Figure 1 is a schematic diagram of a vehicle equipped with an automatic response device. In this embodiment, the vehicle 1 has a camera 2, a plurality of microphones 3-1 to 3-n (where n is an integer of 2 or more, 4 in the illustrated example), a notification device 4, and an automatic response device 5. The camera 2, microphones 3-1 to 3-n, notification device 4, and automatic response device 5 are connected to each other in a way that allows them to communicate.
[0015] Camera 2 is an example of an in-vehicle sensor and is mounted near the top edge of the windshield, facing into the vehicle interior, so that all occupants of vehicle 1 are included in its target area. Camera 2 generates an image representing the interior of vehicle 1 at predetermined shooting intervals and outputs the generated image to the automatic response device 5. Hereafter, the image generated by camera 2 will be referred to as the in-vehicle image. The in-vehicle image is also an example of an in-vehicle sensor signal.
[0016] Microphones 3-1 to 3-n are another example of in-vehicle sensors, and they pick up the voice uttered by any of the passengers riding in the vehicle 1 and output a voice signal representing that voice. For this purpose, one of the microphones 3-1 to 3-n is attached to each individual seat in the vehicle 1 at a position where the voice uttered by the passenger sitting in that seat can be picked up. In the example shown in FIG. 1, one microphone is attached to the vicinity of each of the driver's seat, the front passenger seat, the right rear seat, and the left rear seat (for example, for the driver's seat and the front passenger seat, the instrument panel, and for the rear seats, the back of the seat in front of them). Note that the microphones 3-1 to 3-n may be attached in an array at a position (for example, near the ceiling in the front of the passenger compartment or the instrument panel) where the voice uttered by any passenger in the vehicle interior can be collected. Each of the microphones 3-1 to 3-n outputs the generated voice signal to the automatic response device 5. The voice signals generated by the individual microphones are another example of in-vehicle sensor signals.
[0017] The notification device 4 is provided in the passenger compartment of the vehicle 1 and notifies the passenger of the response content represented by the response information generated by the automatic response device 5. For this purpose, the notification device 4 has, for example, at least one of a speaker or a display device. Then, when the notification device 4 receives a notification signal representing the response content from the automatic response device 5 to the passenger, it notifies the driver of the response content by voice from the speaker or by displaying a message, an image, or a moving image on the display device. Note that the display device or the speaker included in the notification device 4 may be attached for each seat toward the passenger sitting in that seat.
[0018] The automatic response device 5 generates response information for a speech from any of the passengers in the vehicle 1, notifies the passengers in the vehicle 1 of the generated response information via the notification device 4, or controls the vehicle 1 itself or any device mounted on the vehicle 1 according to the response information.
[0019] FIG. 2 is a hardware configuration diagram of the automatic response device 5. As shown in FIG. 2, the automatic response device 5 includes a communication interface 21, a memory 22, and a processor 23. The communication interface 21, the memory 22, and the processor 23 may each be configured as separate circuits, or may be integrally configured as one integrated circuit.
[0020] The communication interface 21 has an interface circuit for connecting the automatic response device 5 to other devices in the vehicle. The communication interface 21 transmits the in-vehicle image received from the camera 2 and the voice signals received from each of the microphones 3-1 to 3-n to the processor 23. Also, the communication interface 21 outputs the notification signal received from the processor 23 to the notification device 4, or outputs a control command for any of the in-vehicle devices received from the processor 23.
[0021] The memory 22 is an example of a storage unit and includes, for example, a volatile semiconductor memory and a non-volatile semiconductor memory. The memory 22 stores various data used in the automatic response process executed by the processor 23. Specifically, the memory 22 stores parameters that define a generation model for generating response information. Further, the memory 22 may temporarily store the in-vehicle image received from the camera 2 and the voice signals received from each of the microphones 3-1 to 3-n.
[0022] The processor 23 includes one or more CPUs (Central Processing Units) and its peripheral circuits. The processor 23 may further include other arithmetic circuits such as an arithmetic logic unit, a numerical arithmetic unit, or a graphic processing unit. And the processor 23 executes the automatic response process.
[0023] Figure 3 is a functional block diagram of the processor 23 related to automatic response processing. The processor 23 includes an emotion estimation unit 31, a response generation unit 32, a notification processing unit 33, and a control unit 34. Each of these parts of the processor 23 is, for example, a functional module realized by a computer program running on the processor 23. Alternatively, each of these parts of the processor 23 may be a dedicated arithmetic circuit provided on the processor 23.
[0024] The emotion estimation unit 31 estimates the emotions of each passenger in the vehicle 1. To do this, the emotion estimation unit 31 inputs an in-vehicle image into a classifier that has been pre-trained to detect passengers, and detects the sub-region on the in-vehicle image that represents each passenger. The classifier for passenger detection is configured as a DNN with a convolutional neural network (CNN) type architecture such as a Single Shot MultiBox Detector, or as a DNN with an attention mechanism such as a Vision Transformer. Alternatively, the classifier for passenger detection may be configured as a classifier based on a machine learning method other than a DNN, such as an AdaBoost classifier.
[0025] The emotion estimation unit 31 estimates the emotions of each crew member by inputting a subregion representing that crew member into an emotion estimator that has been pre-trained to estimate emotions. The emotion estimator is configured as a CNN, for example, having multiple convolutional layers, one or more fully connected layers, and an output layer in order from the input side. The output layer calculates the probability of each type of emotion (joy, anger, sadness, surprise, anxiety, neutral, etc.) using softmax calculation.
[0026] The emotion estimation unit 31 may also estimate the emotions of individual occupants according to other emotion estimation methods. For example, the emotion estimation unit 31 may estimate the emotions of occupants by detecting multiple feature points of the occupant's face (such as multiple points on the contours of the eyes, nose, and mouth) from a subregion representing the occupant, and inputting the individual feature points into an emotion estimator that estimates emotions based on the detected individual feature points. In this case, the emotion estimation unit 31 detects the individual facial feature points by applying a feature point detector based on an Active Shape Model (ASM) or Active Appearance Models (AAM) to the subregion representing the occupant. The emotion estimator may be configured as a DNN or a Support Vector Machine (SVM).
[0027] Alternatively, the emotion estimation unit 31 may estimate the emotions of individual occupants based on the audio signals generated by microphones 3-1 to 3-n. In this case, the emotion estimation unit 31 estimates the emotions of the occupants at the location of the microphone that generated the audio signal by inputting an audio signal whose average volume over a recent predetermined period is equal to or greater than a predetermined threshold into an emotion estimator that estimates emotions based on the audio signal. In this case, the emotion estimator is configured as a DNN having a recurrent structure such as a recurrent neural network (RNN) or a Long Short-Term Memory (LSTM). Alternatively, the emotion estimator may be configured based on a machine learning method other than a DNN.
[0028] The emotion estimation unit 31 generates an emotion vector for each crew member, representing the estimated emotion of that crew member. The emotion vector is a vector whose elements are the likelihood of each type of emotion. For example, suppose the individual elements of the emotion vector are represented by likelihood values in the order of joy, anger, sadness, surprise, anxiety, and neutrality. And suppose the estimated emotions for a certain crew member are 0.6, 0.02, 0.03, 0.1, 0.05, and 0.2, respectively. In this case, the emotion vector is (0.6, 0.02, 0.03, 0.1, 0.05, 0.2).
[0029] The emotion estimation unit 31 further identifies the position of each occupant. For example, reference points on the interior image corresponding to the positions of individual seats in the vehicle are pre-stored in the memory 22. Then, for each occupant detected from the interior image, the emotion estimation unit 31 identifies the reference point closest to the centroid of the sub-region representing that occupant, and identifies the position of the seat corresponding to the identified reference point as the occupant's position. Furthermore, for occupants whose emotions have been estimated based on audio signals, the emotion estimation unit 31 may identify the position of the seat closest to the microphone that generated the audio signal among the microphones 3-1 to 3-n as the occupant's position. If the microphones 3-1 to 3-n are mounted in an array, the emotion estimation unit 31 may estimate the direction of arrival of the sound based on the phase difference between the audio signals generated by each microphone, and estimate the position of the seat closest to the direction of arrival of the sound as seen from each microphone as the occupant's position. In this case, the direction from the mounting position of each microphone to each seat only needs to be pre-stored in the memory 22. The emotion estimation unit 31 then refers to the directions to each seat that are pre-stored in the memory 22 and identifies the seat that is closest in the direction from which the sound is coming.
[0030] The emotion estimation unit 31 outputs an emotion vector for each occupant, along with location information representing the occupant's position, to the response generation unit 32. The location information includes a string representing the occupant's position (for example, a string representing the seat of the occupant who spoke, such as the driver's seat or the passenger seat), or a vector representing the occupant's position. When the occupant's position is represented by a vector, for example, the vector includes different elements for each seat, and the value of the element corresponding to the seat where the occupant associated with the emotion estimation result is sitting is generated to be different from the value of the elements corresponding to other seats.
[0031] The response generation unit 32 generates response information by inputting the estimated emotions of each crew member and the audio information representing the content of their speech into a generative model that has been pre-trained to generate response information for the content of their speech.
[0032] The response generation unit 32 recognizes the content of the utterance expressed in the speech signal by inputting speech signals from each of the microphones 3-1 to 3-n, where the average volume value over the most recent predetermined period exceeds the speech detection threshold, into the speech recognition model, and generates a string representing the content of the utterance as speech information. Such a speech recognition model may be configured as, for example, a DNN with an attention mechanism, or a DNN with a recursive structure such as an RNN or LSTM. Alternatively, the speech recognition model may be configured as a GMM-HMM based on a mixture of normal distributions and a hidden Markov model, or as a DNN-HMM based on a DNN and a hidden Markov model. The response generation unit 32 may also divide the speech signal into frames with a predetermined time length, extract speech features from each frame, and input the feature quantities for each frame into the speech recognition model in chronological order to recognize the content of the utterance expressed in the speech signal. The feature quantities for each frame may be, for example, predetermined elements of the cepstrum of that frame.
[0033] The response generation unit 32 represents the combination of the content of the utterance and the estimated emotions of each crew member in a single string by adding a string representing the estimated emotions of each crew member before or after the string representing the content of the utterance indicated by the voice information.
[0034] In this embodiment, the generation model is configured as an LLM. The LLM, which serves as the generation model, is configured as a stack of multiple blocks, for example, each containing an attention layer and a feed-forward layer. The response generation unit 32 inputs a string representing the content of the utterance and the estimated emotions of each occupant to the LLM. As a result, the generation model outputs text data representing the response content as response information. In this way, by inputting the estimated emotions of each occupant along with the audio information representing the content of the utterance to the generation model, the generation model can generate response information that takes into account the emotions of each occupant.
[0035] Furthermore, the response information is not limited to the response content notified to each occupant via the notification device 4, but may also include response content that controls the vehicle 1 itself or equipment installed in the vehicle 1.
[0036] In a modified version, the response generation unit 32 may input the estimated emotion of each occupant, along with the occupant's location information, into the generation model. In this case, the response generation unit 32 only needs to insert a string representing the occupant's location (e.g., driver's seat, passenger seat) before and after the string representing the estimated emotion of each occupant in the string input to the generation model. By inputting location information along with the response information in this way, the generation model can reflect the differences in emotion based on the location of each occupant in the response information.
[0037] Figure 4 is an explanatory diagram of the relationship between input and response information in the generation model according to this embodiment. In the example shown in Figure 4, an occupant is seated in the driver's seat and the passenger seat of vehicle 1. Therefore, the emotions of occupant 401, who is the driver seated in the driver's seat, and occupant 402, who is seated in the passenger seat, are estimated. The estimated emotion result 411 for occupant 401 is "joy 0.0 anger 0.4 sadness 0.3 surprise 0.0 anxiety 0.1 neutral 0.2", and the estimated emotion result 412 for occupant 402 is "joy 0.0 anger 0.4 sadness 0.2 surprise 0.0 anxiety 0.1 neutral 0.3".
[0038] Furthermore, either occupant 401 or occupant 402 says "It's hot." In this case, the text data 430 input to the generative model 420 includes the spoken content "It's hot," the occupant's position "Driver's seat (Passenger seat)," and the type of emotion the occupant is feeling, such as "Joy," along with a numerical value indicating the likelihood of that type. If the input to the generative model 420 does not include occupant position information, the string indicating each occupant's position is omitted in the text data 430. The generative model 420 then refers to the text data 430 and outputs text data 440 as response information, representing the response to the spoken content ("I'll turn up the air conditioning").
[0039] In this embodiment, the generation model 420 generates response information to utterances by referring to the results of emotion estimation for each occupant. Therefore, in the above example, the higher the degree of discomfort each occupant feels, that is, the higher the probability of negative emotions such as anger, sadness, and anxiety, the more the generation model 420 can generate response information that includes a response indicating to increase the air conditioning in response to the utterance "It's hot." Conversely, if the degree of discomfort each occupant feels, that is, the lower the probability of negative emotions such as anger, sadness, and anxiety, the generation model 420 can generate response information that confirms whether or not to increase the air conditioning in response to the utterance "It's hot."
[0040] Furthermore, by inputting the estimated emotions of each occupant along with their location information into the generative model 420, the generative model 420 can generate response information that takes into account the differences in emotions among the occupants. For example, if the level of discomfort of occupant 402 sitting in the passenger seat is higher than that of occupant 401 sitting in the driver's seat, the generative model 420 can generate response information that includes a response to increase the air conditioning on the passenger side.
[0041] In a modified version, the generative model may have an input layer for emotion estimation results, separate from the stacked blocks, into which vectors representing the estimated emotions of each individual crew member are input. In this case, only the string corresponding to the audio information is input to the innermost block of the stacked blocks, and the vectors representing the estimated emotions of each crew member are input to the input layer for emotion estimation results. The vectors representing the estimated emotions of each crew member, input to that input layer, are then taken into a block of the generative model that has a cross-attention mechanism that calculates cross-attention between the vectors representing the estimated emotions of each crew member and the outputs from previous blocks. In this case, tuning techniques such as LoRA may be applied to the learning of the generative model regarding the acquisition of the vectors representing the estimated emotions of each crew member.
[0042] In another modification, the generative model may be configured to output separately response information for notifying the occupants of vehicle 1 via notification device 4, and response information for controlling vehicle 1 itself or equipment installed in vehicle 1. In this case, the generative model is configured such that a stack of blocks branches off from the middle, with one or more blocks generating notification response information and one or more blocks generating control response information in parallel. Each block after the branch is also configured to include an attention layer and a feed-forward layer. The notification response information and the control response information are determined separately according to the output probability calculated by a softmax operation on the output from the corresponding final stage block. In this case as well, a tuning method such as LoRA may be applied to the learning of the generative model for the part of the notification response information and the control response information that is added to the base model.
[0043] The response generation unit 32 outputs the generated response information to the notification processing unit 33 and the control unit 34.
[0044] The notification processing unit 33 outputs response information via the notification device 4. To do this, the notification processing unit 33 generates a notification signal representing the response content included in the response information and outputs the generated notification signal to the notification device 4 via the communication interface 21. For example, the notification processing unit 33 generates an audio signal representing the response content as a notification signal based on the text data representing the response content included in the response information, according to a predetermined speech synthesis method. The notification processing unit 33 then outputs the notification signal to the speaker of the notification device 4, causing the speaker to output audio representing the response content. Alternatively, the notification processing unit 33 includes text data representing the response content in the notification signal. The notification processing unit 33 then causes the display device of the notification device 4 to display the text data representing the response content.
[0045] Furthermore, if the response information generated by the generation model represents the control content of the equipment of vehicle 1, the processing by the notification processing unit 33 may be omitted.
[0046] The control unit 34 controls the device specified in the response information, according to the response. The control unit 34 determines the device to be controlled and the control command by referring to a control reference table that shows the correspondence between the text data representing the response information, the device to be controlled (including the vehicle 1 itself), and the control command for executing that control. The control unit 34 then outputs the specified control command to the electronic control unit (ECU) of the device to be controlled via the communication interface 21.
[0047] The controlled equipment may include, in addition to the air conditioning system, windows, door locks, interior lights, and seats. The control unit 34 controls the equipment according to the response, by opening and closing the windows, locking and unlocking the doors, turning the interior lights on or off, or adjusting the seat position of any of the occupants' seats.
[0048] Furthermore, if the text data representing the response does not contain any of the multiple words registered in the control reference table that identify the controlled device, the response information does not control the device. In this case, the control unit 34 does not output a control signal.
[0049] Figure 5 is an operation flowchart of the automatic response process according to this embodiment. The processor 23 executes the automatic response process according to this operation flowchart.
[0050] The emotion estimation unit 31 estimates the emotions of each crew member (step S101). The response generation unit 32 generates speech information representing the content of speech by any of the crew members based on the speech signals generated by any of the microphones 3-1 to 3-n (step S102). The response generation unit 32 then generates response information by inputting the speech information and the estimated emotions of each crew member into the generation model (step S103). As described above, the response generation unit 32 may also input the estimated emotions of each crew member into the generation model along with the crew member's position information.
[0051] The notification processing unit 33 notifies each occupant of the response content included in the response information via the notification device 4 (step S104). The control unit 34 also controls the vehicle 1 itself or the equipment installed in the vehicle 1 according to the response content (step S105).
[0052] As explained above, this automated response system estimates the emotions of each occupant in the vehicle and generates response information by inputting the estimated emotions of each occupant and audio information representing the content of a utterance from one of the occupants into a generative model. Therefore, this automated response system can respond to utterances from occupants in the vehicle in a way that reflects the emotions of the occupants.
[0053] In a modified version, the automatic response device 5 may use a combination of response information generated in the most recent predetermined period and the estimated results of the emotions of individual occupants input into the generation model when generating the response information, along with voice information, to generate the latest response information.
[0054] Figure 6 is a functional block diagram of the processor 23 of the automatic response device 5 according to this modified example. The processor 23 includes an emotion estimation unit 31, a response generation unit 32, a notification processing unit 33, a control unit 34, and a storage processing unit 35. Each of these parts of the processor 23 is, for example, a functional module realized by a computer program running on the processor 23. Alternatively, each of these parts of the processor 23 may be a dedicated arithmetic circuit provided on the processor 23. This modified example differs from the above embodiment in that the processor 23 has a storage processing unit 35 and in some of the processing of the response generation unit 32. Therefore, these differences will be explained below.
[0055] The storage processing unit 35 stores in memory 22, each time response information is generated, the combination of the response information and the information input to the generation model to generate that response information (the emotion estimation results and voice information of each crew member, and location information if location information is input), associated with the date and time the response information was generated. The information input to the generation model to generate the response information may be simply referred to as input information below.
[0056] When generating response information, the response generation unit 32 inputs the combination of response information and input information described above for the most recent predetermined period, along with the estimated emotion and voice information of each crew member, into the generation model. In this modified version, the location information of each crew member may also be input into the generation model along with the estimated emotion when generating response information. The generation model then generates response information by referring to the response information and input information for the most recent predetermined period, along with the estimated emotion (and location information) and voice information of each crew member. In this modified version, the generation model also refers to the past response information generation history, so it can generate more appropriate response information.
[0057] The response generation unit 32 may generate text data to input to the generation model by representing the response information and input information for the most recent predetermined period as strings, and combining these strings with strings representing the estimated emotions (and location information) and voice information of each crew member. Alternatively, when the generation model generates response information, the output of any of the individual blocks of the generation model, including the attention layer and the feed-forward layer, may be stored in the memory 22. The output of that block may then be recursively input to that block when generating the next response information.
[0058] Figure 7 is an operation flowchart of the automated response process according to this modified example. The processor 23 executes the automated response process according to this operation flowchart.
[0059] The emotion estimation unit 31 estimates the emotion of each crew member (step S201). The response generation unit 32 generates speech information representing the content of speech by any of the crew members based on the speech signals generated by any of the microphones 3-1 to 3-n (step S202). The response generation unit 32 then generates response information by inputting the speech information, the estimated emotion of each crew member, and the combination of response information and input information for the most recent predetermined period into the generation model (step S203). In this modified example, the response generation unit 32 may also input the estimated emotion of each crew member into the generation model along with the crew member's location information.
[0060] The notification processing unit 33 notifies each occupant of the response content included in the response information via the notification device 4 (step S204). The control unit 34 controls the vehicle 1 itself or the equipment installed in the vehicle 1 according to the response content (step S205). The storage processing unit 35 stores the combination of the response information and the input information input to the generation model to generate the response information in the memory 22 (step S206).
[0061] As explained above, this modified automated response system refers to the history of past response and input information combinations when generating response information. Therefore, this automated response system can generate more appropriate response information.
[0062] In the above embodiments and modifications, the response generation unit 32 may further input an in-vehicle image or a sub-region representing an individual occupant on the in-vehicle image to the generation model when generating response information. In this case, the generation model is configured as a Vision Language Model (VLM). By inputting an in-vehicle image or a sub-region representing an occupant in this way, the generation model can generate response information by referring to the occupant's state. When a sub-region representing an occupant is input to the generation model, the response generation unit 32 can either crop the sub-region representing that occupant from the in-vehicle image for each occupant, or mask the area by converting the values of individual pixels in the area other than the sub-region representing each occupant to predetermined pixel values.
[0063] Furthermore, in the above embodiments or each of its modifications, the response generation unit 32 may further input to the generation model, when generating response information, at least one of the following signals: a sensor signal obtained from a sensor installed in the vehicle 1 for detecting the behavior of the vehicle 1; a sensor signal obtained from a sensor for detecting the conditions inside the vehicle 1 or the conditions around the vehicle 1; an amount of operation performed by the occupant on the vehicle 1 (accelerator opening, brake amount, steering angle); and a signal representing a setting amount for an in-vehicle device. The sensor for detecting the behavior of the vehicle 1 is, for example, a speed sensor or an acceleration sensor. The sensor for detecting the conditions inside the vehicle 1 or the conditions around the vehicle 1 is, for example, a thermometer, an illuminometer, or a rain sensor. The amount of operation performed by the occupant on the vehicle 1 is, for example, an accelerator opening, a brake amount, or a steering angle. The setting amount for an in-vehicle device is, for example, an air conditioning set temperature, a set airflow rate, the open / closed state of the windows, or the set volume of the audio system. Hereafter, these signals will be referred to as vehicle state information. By inputting vehicle status information into the generation model, the generation model can refer to the status of vehicle 1 or the in-vehicle equipment, thereby enabling the generation of more appropriate response information.
[0064] The response generation unit 32 can generate text data to input to the generative model by converting the type and signal value of each sensor signal included in the vehicle state information into a string, and then combining the converted string with a string representing the estimated emotion (and location information) and voice information of each occupant. Alternatively, similar to the case where the estimated emotion of each occupant is input separately from the voice information, the generative model may be provided with an input layer for inputting vehicle state information, separate from the block into which the estimated emotion (and location information) and voice information of each occupant are input. In this case, only the strings corresponding to the estimated emotion (and location information) and voice information of each occupant are input to the input-side block of the stacked blocks, and the vehicle state information is input to the input layer for vehicle state information input. The vehicle state information input to that input layer is then taken in by a block of the stacked blocks of the generative model that has a cross-attention mechanism that calculates cross-attention between the vehicle state information and the output from the previous block. In this case as well, tuning methods such as LoRA may be applied to the learning of the generative model regarding the acquisition of vehicle state information.
[0065] In the above embodiment or each of its modifications, a server (not shown) that can communicate via an in-vehicle wireless communication terminal (not shown) may have a generation model, and the server may generate response information by executing the processing of the response generation unit 32 and the storage processing unit 35. In this case, the automatic response device 5 generates a question signal that includes the estimated emotion and location information of each occupant, and text data representing the content of their speech, and transmits the generated question signal to the server via the wireless communication terminal. The server then generates response information by inputting the text data contained in the received question signal into the generation model, and transmits the generated response information to the wireless communication terminal of the vehicle 1 that sent the question signal. As in the above modification, the server may also input combinations of response information and input information for the most recent predetermined period into the generation model. The automatic response device 5 only needs to execute the processing of the notification processing unit 33 and the control unit 34 according to the response information received via the wireless communication terminal. According to this modification, the automatic response device 5 can use a larger generation model than the generation model realized with in-vehicle hardware resources to generate response information, and can generate more appropriate response information.
[0066] In other modifications, the processing of either the notification processing unit 33 or the control unit 34 may be omitted.
[0067] The computer program that implements the automated response processing according to the above embodiment or modification may be provided as a computer program product, for example, in the form of being recorded on a computer-readable portable recording medium.
[0068] As described above, those skilled in the art can make various modifications within the scope of the present invention to suit the implemented form. [Explanation of Symbols]
[0069] 1 Vehicle, 2 Camera, 3-1~3-n Microphone, 4 Notification device, 5 Automatic response device, 21 Communication interface, 22 Memory, 23 Processor, 31 Emotion estimation unit, 32 Response generation unit, 33 Notification processing unit, 34 Control unit, 35 Storage processing unit
Claims
1. An emotion estimation unit that estimates the emotions of individual occupants of the vehicle based on in-vehicle sensor signals obtained from in-vehicle sensors that detect the conditions inside the vehicle's interior, A response generation unit that generates response information by inputting the estimated emotions of each crew member and audio information representing the content of an utterance from any of the crew members into a generative model that has been pre-trained to generate response information for the content of the utterance, An automatic response device having the following features.
2. The automatic response device according to claim 1, wherein the response generation unit generates the response information by further inputting the estimated emotion of each occupant, along with location information representing the occupant's position, into the generation model.
3. Each time the aforementioned response information is generated, the storage processing unit further stores the combination of the response information, the estimated results of the emotions of each occupant input into the generation model at the time of generating the response information, and the voice information in the storage unit. The automatic response device according to claim 1 or 2, wherein the response generation unit generates the response information by further inputting the combinations stored in the storage unit for the most recent predetermined period into the generation model.
4. Based on the in-vehicle sensor signals obtained by in-vehicle sensors that detect the conditions inside the vehicle's interior, the emotions of each individual occupant of the vehicle are estimated. The response information is generated by inputting the estimated emotions of each crew member and the audio information representing the content of an utterance from any of the crew members into a generative model that has been pre-trained to generate response information for the content of the utterance. An automated response method that includes the following.
5. Based on the in-vehicle sensor signals obtained by in-vehicle sensors that detect the conditions inside the vehicle's interior, the emotions of each individual occupant of the vehicle are estimated. The response information is generated by inputting the estimated emotions of each crew member and the audio information representing the content of an utterance from any of the crew members into a generative model that has been pre-trained to generate response information for the content of the utterance. An automated computer program for instructing a computer to perform a specific action.
Citation Information
Patent Citations
Vehicle control device
JP2024084568A