Vehicle dialogue system
By integrating interactive AI and multiple sensor data processing technologies in the on-board dialogue system, the problem that the existing on-board voice assistant system cannot respond naturally is solved, and the natural response and personalized processing of driver voice input in the on-board environment is realized.
Patent Information
- Application Number
- JP2023181851
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2025-05-08
AI Technical Summary
Existing on-board voice assistant systems cannot respond naturally to drivers’ voice inputs, and interactive AI cannot obtain appropriate responses in on-board environments unless the driver mentions driving status and vehicle status in sequence.
A vehicle dialogue system is designed. Using interactive AI, the driver's voice input is converted into text data through the sound input unit, the voice recognition unit, the input control unit, the voice synthesis unit and the audio output unit, and the driver's status information is input into the interactive AI, generating response information and converting it into voice output.
The on-board dialogue system can naturally respond to the driver's voice input, understand the driver and the vehicle's status, provide personalized and appropriate response, and improve the naturalness and effectiveness of on-board interaction.
Smart Images

Figure 2025071571000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a dialogue system for a vehicle. [Background technology]
[0002] A voice agent service has been proposed that responds by voice when a user speaks (Patent Document 1). In the voice agent service of Patent Document 1, it is proposed to return predetermined noise data if the user's speech content is not a search. However, in the voice agent service of Patent Document 1, since the noise data is extracted from among predetermined data, a natural response like that of a human conversation is not returned.
[0003] In recent years, conversational AI such as ChatGPT has been proposed that can learn from the vast amount of information available on the Internet and have more natural conversations. However, such conversational AI cannot return an appropriate answer unless an appropriate prompt (text data) is input. For this reason, when using conversational AI in a vehicle, there was a problem in that an appropriate response would not be returned unless the driver spoke out the driver's condition and the vehicle's condition one by one. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2014-98844 A Summary of the Invention [Problem to be solved by the invention]
[0005] The present invention has been made in view of the above-mentioned circumstances, and an object of the present invention is to provide a vehicle dialogue system that responds appropriately to the driver's voice. [Means for solving the problem]
[0006] In order to achieve the above object, the vehicle dialogue system according to the present invention has the following features. A dialogue system for a vehicle using a dialogue type AI that outputs response information made of text data when input information made of text data is input, A voice input unit for inputting a voice uttered by a driver; a voice recognition unit that converts the voice input by the voice input unit into text data; an input control unit that inputs the text data converted by the voice recognition unit and text data indicating the state of the driver or the vehicle driven by the driver when the driver speaks, or an instruction corresponding to the state, as the input information to the dialogue AI; A voice synthesis unit that converts the response information from the dialogue AI into voice; A voice output unit that outputs the voice converted by the voice synthesis unit. It is a dialogue system for vehicles. Effect of the Invention
[0007] According to the present invention, it is possible to provide a vehicle dialogue system that responds appropriately to the driver's voice.
[0008] The present invention has been briefly described above. Furthermore, the details of the present invention will be further clarified by reading the following description of the embodiment of the present invention (hereinafter, referred to as "embodiment") with reference to the accompanying drawings. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing an embodiment of a dialogue system for a vehicle according to the present invention. [Diagram 2] FIG. 2 is a flowchart showing a processing procedure of the microcomputer shown in FIG. [Diagram 3] FIG. 3 is an explanatory diagram showing an example of text data transmitted from the vehicle dialogue system to the dialogue AI when the ignition is turned on. [Figure 4] FIG. 4 is an explanatory diagram showing an example of a conversation between the vehicular dialogue system and the driver. [Diagram 5] FIG. 5 is an explanatory diagram showing an example of a conversation between the vehicular dialogue system and the driver. [Figure 6] FIG. 6 is an explanatory diagram showing an example of a conversation between the vehicular dialogue system and the driver. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Specific embodiments of the present invention will be described below with reference to the accompanying drawings.
[0011] The vehicle dialogue system 1 of this embodiment is a system that is mounted on a vehicle and dialogues with a driver using a dialogue type AI (Artificial Intelligence) 10. The dialogue type AI 10 is configured, for example, by ChatGPT, and when input information S1 consisting of text data is input, response information S2 consisting of text data is output.
[0012] The vehicle dialogue system 1 includes a microphone 2 as a voice input unit, a communication module 3, a microcomputer 4 (hereinafter abbreviated as "microcomputer 4"), and a speaker 5 as a voice output unit. The microphone 2 inputs the voice spoken by the driver to the microcomputer 4. The communication module 3 is for communicating with the dialogue type AI 10 via an Internet communication network (not shown), and is composed of a circuit, an antenna, etc. for connecting to the Internet communication network.
[0013] The microcomputer 4 has memories such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and a CPU (Central Processing Unit) that operates according to programs stored in the memories.
[0014] The microcomputer 4 has a voice recognition unit 41, a voice dialogue unit 42 as an input control unit, a first determination unit, an equipment control unit, and a second determination unit, and a voice synthesis unit 43. The voice recognition unit 41 converts the voice input by the microphone 2 into text data and inputs it to the voice dialogue unit 42. The voice dialogue unit 42 adds text data indicating the state of the driver or the vehicle driven by the driver when the driver spoke, or an instruction according to the state, to the text data converted by the voice recognition unit 41, and inputs it to the interactive AI 10 as input information S1.
[0015] Further, sensor information S3 from a plurality of sensors mounted on the vehicle is input to the voice dialogue unit 42. Possible sensors include an illuminance sensor that measures the illuminance outside the vehicle, a temperature sensor that measures the outside air temperature, a GPS (Global Positioning System) that detects the vehicle position, a seating sensor that detects whether or not a person is seated in a seat, a speed sensor that detects the vehicle speed, etc. The sensor information S3 input from the sensor may be not only a measurement value measured by the sensor, but also a moving average or weighted average of the measurement value, or a value calculated using the measurement value, the moving average or weighted average of the measurement value.
[0016] Furthermore, human information S4, which is a detection result from a driver monitor that detects the driver's state (whether the driver is dozing, driving absent-mindedly, or not looking at the road) based on an image of the driver's face, is input to the voice dialogue unit 42. Furthermore, warning information S5, which is a detection result from an abnormality detection unit that detects vehicle abnormalities (engine abnormality, oil pressure abnormality, water temperature abnormality, charging abnormality, etc.) and turns on a warning lamp, is input to the voice dialogue unit 42.
[0017] The voice dialogue unit 42 is connected to a vehicle device 11 mounted on the vehicle and can control the vehicle device 11. The vehicle device 11 may be, for example, an air conditioner, a display such as a head-up display, or an audio device mounted on the vehicle.
[0018] The voice dialogue unit 42 receives response information S2 from the dialogue-type AI 10 and outputs the received response information S2 to the voice synthesis unit 43. The voice synthesis unit 43 converts the response information S2 into voice and outputs it to the speaker 5. The speaker 5 outputs the voice converted by the voice synthesis unit 43.
[0019] Next, the operation of the vehicle dialogue system 1 having the above-mentioned configuration will be described with reference to the flowchart shown in Fig. 2. When the microcomputer 4 detects that the ignition is on (IG-ON) (Sp1), it transmits text data indicating the state of the driver and the vehicle to the dialogue AI 10 (Sp2). An example of the text data indicating the state of the driver and the vehicle is, as shown in Fig. 3, "The person you are speaking to is about to get into the car. Please talk to him / her with the driver in mind." This allows the dialogue AI 10 to understand that the person you are speaking to (the driver) is in the car, and the dialogue content thereafter is based on the assumption that the person is in the car.
[0020] Also, in Sp2, the microcomputer 4 may transmit text data of information on the vehicle equipment 11 that can be controlled. An example of this text data is "We can suggest adjusting the temperature and airflow of the air conditioner, adjusting the brightness of the head-up display (HUD), and controlling the audio, depending on the conversation with the person you are talking to." This allows the conversational AI 10 to understand that the driver can adjust the temperature and airflow of the air conditioner, adjust the brightness of the head-up display, and control the audio in the vehicle he is riding in.
[0021] Also, in Sp2, the microcomputer 4 may determine the number of people riding in the vehicle based on the detection result from the seating sensor, and transmit the determined number of people to the conversational AI 10. An example of this text data is "Only the driver is riding in the vehicle" or "There are two people riding in the vehicle, including the driver." This allows the conversational AI 10 to understand whether there is only one driver or other people riding in the vehicle.
[0022] Next, the microcomputer 4 acquires the above-mentioned sensor information S3, human information S4, and warning information S5 (Sp3). After that, when the driver speaks (Y in Sp4), the microcomputer 4 performs a voice recognition process to convert the voice spoken by the driver into text data (Sp5). Next, the microcomputer 4 adds the driver and vehicle state when the driver spoke, or a command (text data) according to the state, to the text data converted by the voice recognition process (Sp6), and transmits the text data to the interactive AI 10 as input information S1 (Sp7).
[0023] In Sp6, the microcomputer 4 may add text data indicating the sensor information S3, the person information S4, and the warning information S5 acquired in Sp3 as text data indicating the state of the driver and the vehicle when the driver spoke. In addition, the microcomputer 4 may add text data indicating whether the vehicle is driving, stopped, or waiting at a traffic light as text data indicating the state of the vehicle.
[0024] Furthermore, the microcomputer 4 determines whether the driver's driving load is high when the driver speaks. Whether the driving load is high can be determined based on, for example, the speed from a speed sensor (sensor information S3) or the person information S4. Alternatively, it may be determined based on pre-registered attributes of the driver (whether the driver is accustomed to driving or not, whether the driver is familiar with vehicle equipment or not). If the microcomputer 4 determines that the driving load is high when the driver speaks, it may give text data to shorten the response information S2 as a command according to the state. An example of the text data at this time may be "Please summarize your answer."
[0025] Thereafter, the microcomputer 4 receives response information S2 corresponding to the input information S1 transmitted in Sp7 (Sp8), converts the received response information S2 into voice, and outputs it from the speaker 5 (Sp9). Next, when the microcomputer 4 detects that the ignition is off (IG-OFF) (Y in Sp10), it ends the process. On the other hand, when the microcomputer 4 does not detect that the ignition is off (N in Sp10), it returns to Sp3 again.
[0026] Furthermore, after executing Sp3, if the driver has not spoken (N in Sp4), or if a predetermined proposal condition based on the information S3 to S5 acquired by Sp3 is satisfied (Sp11), the microcomputer 4 converts the proposal content into voice data and outputs it from the speaker 5 (Sp12), and then returns to Sp3. As a proposal condition for Sp11, for example, if the speed measured by the speed sensor (sensor information S3) is fast, the proposal condition to reduce the speed is satisfied, and a proposal to reduce the speed is output from the speaker 5.
[0027] The above-mentioned vehicle dialogue system 1 transmits the state of the driver and the vehicle when the driver is speaking to the dialogue AI 10. This allows the dialogue AI 10 to understand the state of the driver and the vehicle when the driver is speaking and respond appropriately to the driver's voice.
[0028] The above-mentioned vehicular dialogue system 1 transmits a command to shorten the response information to the dialogue AI 10 when the driving load is high when the driver speaks. This enables the dialogue AI 10 to shorten the response when the driving load is high when the driver speaks.
[0029] For example, as shown in the conversation examples in Figures 4 and 5, a case will be described where the warning lamp turns on while driving and the driver says, "A red triangular exclamation mark is on. What is that?" When the microcomputer 4 of the vehicular dialogue system 1 determines that the driving load is high, it adds text data saying "Please summarize the answer" to the text data of the voice spoken by the driver, as shown in Figure 4, and transmits it to the dialogue AI 10.
[0030] On the other hand, if microcomputer 4 judges that the driving load is low, it does not give the text data "Please summarize your answer" as shown in Figure 5. As a result, when the driving load is low, as shown in Figure 5, a detailed and long phrase like that written in the vehicle manual is returned, "The red warning light in the shape of an exclamation mark is a master warning. The master warning lights up when other warning lights or indicator lights come on, or when a warning message is displayed in the multi-information display, and also sounds a buzzer depending on the content of the warning."
[0031] On the other hand, if the driving load is high, you will receive a simple but short reply that summarizes the vehicle manual: "The red warning light in the shape of an exclamation mark is the master warning. It comes on when a highly urgent abnormality occurs, so if you are driving, stop the car immediately and contact your dealer."
[0032] The microcomputer 4 of the vehicle dialogue system 1 described above transmits the sensor information S3 from the sensor mounted on the vehicle when the driver speaks as the vehicle status to the dialogue AI 10. This allows the dialogue AI 10 to understand the vehicle status and respond more appropriately to the driver's voice.
[0033] The microcomputer 4 of the vehicle dialogue system 1 described above inputs the detection result of the abnormality detection unit when the driver speaks as the vehicle state. This allows the dialogue AI 10 to understand that a warning lamp is on and respond appropriately to the driver's voice. This allows the dialogue AI 10 to respond appropriately to, for example, the driver's utterance, "Some kind of lamp is on. What is that?"
[0034] The microcomputer 4 of the vehicle dialogue system 1 described above transmits text data of information on the vehicle device 11 that can be controlled in Sp2. This allows the dialogue AI 10 to suggest control of the vehicle device 11 in response to the driver's utterance. Although not shown in the flowchart of FIG. 2, the microcomputer 4 controls the vehicle device 11 based on the text data converted by the voice recognition unit 41 and the response information S2 from the dialogue AI 10.
[0035] For example, as shown in the conversation example in FIG. 6, a case will be described where the driver says, "The sun is shining brighter than usual today, isn't it?" In addition to the text data "The sun is shining brighter than usual today, isn't it?", the vehicle dialogue system 1 transmits the vehicle's position information and illuminance information (text data) to the dialogue AI 10 in Sp7. Furthermore, the vehicle dialogue system 1 transmits that it can control the HUD in Sp2. As a result, the dialogue AI 10 can obtain weather and temperature information from the position information, predict that the HUD is becoming difficult to see, and transmit response information S2 saying, "The weather forecast says it will be sunny all day. The temperature is expected to rise to 25 degrees during the day. Isn't the HUD display difficult to see at this brightness?"
[0036] Furthermore, the microcomputer 4 of the vehicle dialogue system 1 analyzes the subsequent conversation between the vehicle dialogue system 1 and the driver, and controls the HUD, which is the vehicle device 11, to increase the brightness by five levels.
[0037] In addition, the vehicle dialogue system 1 assigns the driver and vehicle status to the first utterance from the driver in a series of conversations, but does not assign the driver and vehicle status to subsequent utterances. That is, in the conversation example shown in Fig. 6, the vehicle dialogue system 1 assigns the driver and vehicle status only to the text data of the first utterance by the driver, "The sun is stronger and more dazzling than usual today," but does not assign driver and vehicle status information to the subsequent text data, "Now that you mention it, I might be having a harder time seeing than usual" and "It's okay. Thank you."
[0038] The microcomputer 4 of the vehicle dialogue system 1 transmits the number of passengers to the dialogue AI 10. This allows the dialogue AI 10 to understand whether the driver is alone or there are other passengers in the car, and respond appropriately to the driver's voice.
[0039] The present invention is not limited to the above-described embodiment, and can be appropriately modified, improved, etc. In addition, the material, shape, size, number, arrangement location, etc. of each component in the above-described embodiment are arbitrary as long as the present invention can be achieved, and are not limited.
[0040] According to the embodiment described above, the information on the controllable vehicle devices 11 and the number of passengers, which do not change from moment to moment, are transmitted immediately after the ignition is turned on, but this is not limited to this. The information on the controllable vehicle devices 11 and the number of passengers may be added to the text data of the voice uttered by the driver and transmitted.
[0041] According to the above-described embodiment, the driver's condition is determined from an image of the driver's face, but this is not limited to the above. The driver's condition may be determined based on, for example, a measurement value from a measuring device worn by the driver to measure the heart rate.
[0042] According to the above-described embodiment, the microcomputer 4 reads out the response information S2 to the end in Sp9, and then returns to Sp4 to accept the driver's speech, but this is not limited to the above. The microcomputer 4 may have a so-called barge-in function. In the barge-in function, when the microcomputer 4 detects speech while reading out the response information S2 in Sp9, the microcomputer 4 may stop reading out the response information S2, return to Sp5, and perform voice recognition of the driver's speech. The barge-in function may be set to be on or off by a user such as the driver.
[0043] Furthermore, the microcomputer 4 may read out the response information S2 in response to an utterance by the driver that interrupts the response information S2 while the response information S2 is being read out, and then read out the response information S2 that was being read out again.
[0044] Also, if the result of voice recognition of the utterance interrupted in Sp5 is command information to stop speaking, such as "stop talking" or "too long," the microcomputer 4 may stop proceeding to Sp7 for transmitting the voice recognition result to the conversational AI 10, and may return to Sp4 to accept the next utterance from the driver. In other words, if the command information is to stop speaking, the microcomputer 4 does not re-read the response information S2 that was in the middle of being read out.
[0045] If the information is not a command to stop speaking, the microcontroller 4 proceeds to Sp7 to transmit the voice result, etc. to the conversational AI 10.
[0046] According to the above-mentioned embodiment, the vehicle dialogue system 1 communicates with the dialogue AI 10 via the Internet communication network, but this is not limited to this. The dialogue AI 10 may be built into a control board in the vehicle. In this case, dialogue is possible even in a poor communication environment.
[0047] Here, the features of the above-described embodiment of the vehicular dialogue system according to the present invention will be briefly summarized and listed in the following [1] to [6].
[0048] [1] A dialogue system (1) for a vehicle using a dialogue type AI (10) that outputs response information (S2) made of text data when input information (S1) made of text data is input, A voice input unit (2) for inputting a voice uttered by a driver; a voice recognition unit (41) that converts the voice input by the voice input unit (2) into text data; an input control unit (42) that inputs the text data converted by the voice recognition unit (41) and text data indicating the state of the driver or the vehicle driven by the driver when the driver spoke, or an instruction corresponding to the state, to the dialogue AI (10) as the input information (S1); a voice synthesis unit (43) that converts the response information (S2) from the dialogue AI (10) into voice; and a voice output unit (5) that outputs the voice converted by the voice synthesis unit (43). Vehicle dialogue system (1).
[0049] According to the configuration [1] above, the conversational AI (10) can understand the state of the driver and the vehicle when the driver is speaking and can respond appropriately to the driver's voice.
[0050] [2] In the vehicular dialogue system (1) according to [1], a first determination unit (42) that determines whether or not a driving load of the driver is high, the input control unit (42) inputs, when the first determination unit (42) determines that the driving load is high when the driver speaks, the text data indicating that the response information (S2) is to be shortened to the dialogue-type AI (10) as the text data indicating the command. Vehicle dialogue system (1).
[0051] According to the configuration [2] above, if the driving load when the driver speaks is high, the response of the dialogue type AI (10) can be made shorter.
[0052] [3] In the vehicular dialogue system (1) according to [1], The input control unit (42) receives sensor information from a sensor mounted on the vehicle, The input control unit (42) inputs the sensor information when the driver speaks to the dialogue type AI (10) as the state of the vehicle. Vehicle dialogue system (1).
[0053] According to the configuration [3] above, the conversational AI (10) can understand the vehicle's condition and respond more appropriately to the driver's voice.
[0054] [4] In the vehicular dialogue system (1) according to [1], The input control unit (42) receives a detection result from an abnormality detection unit that detects an abnormality in the vehicle and turns on a warning lamp, The input control unit (42) inputs the detection result of the abnormality detection unit when the driver speaks to the interactive AI (10) as the state of the vehicle. Vehicle dialogue system (1).
[0055] According to the configuration [4] above, the conversational AI (10) can understand that the warning lamp is on and respond appropriately to the driver's voice.
[0056] [5] In the vehicular dialogue system (1) according to [1], A device control unit that controls a vehicle device (11), The input control unit (42) inputs text data indicating information of the vehicle equipment (11) that can be controlled by the equipment control unit (42) to the interactive AI (10); The device control unit (42) controls the vehicle device (11) based on the text data converted by the voice recognition unit (41) and the response information (S2) from the dialogue-type AI (10). Vehicle dialogue system (1).
[0057] According to the configuration of [5] above, the conversational AI (10) can suggest control of the vehicle equipment (11) in response to the driver's utterances.
[0058] [6] In the vehicular dialogue system (1) according to [1], A second determination unit (42) for determining the number of people riding in the vehicle, The input control unit (42) inputs the number of people determined by the second determination unit (42) to the interactive AI (10). Vehicle dialogue system (1).
[0059] According to the configuration of [6], the conversational AI (10) can understand whether the driver is alone or whether there are other passengers in the car, and respond appropriately to the driver’s voice. [Explanation of symbols]
[0060] 1 Vehicle dialogue system 2 Microphone (audio input section) 5 Audio output section 10 Conversational AI 41 Voice Recognition Unit 42 Voice dialogue unit (input control unit, first determination unit, device control unit, second determination unit) 43 Voice synthesis section S1 Input information S2 Response Information
Claims
1. A dialogue system for a vehicle using a dialogue type AI that outputs response information made of text data when input information made of text data is input, A voice input unit for inputting a voice uttered by a driver; a voice recognition unit that converts the voice input by the voice input unit into text data; an input control unit that inputs the text data converted by the voice recognition unit and text data indicating the state of the driver or the vehicle driven by the driver when the driver speaks, or an instruction corresponding to the state, as the input information to the interactive AI; A voice synthesis unit that converts the response information from the conversational AI into voice; A voice output unit that outputs the voice converted by the voice synthesis unit. Dialogue system for vehicles.
2. 2. The vehicle dialogue system according to claim 1, A first determination unit that determines whether or not the driving load of the driver is high, The input control unit inputs the text data indicating that the response information is to be shortened to the dialogue type AI as the text data indicating the command when the first determination unit determines that the driving load is high when the driver speaks. Dialogue system for vehicles.
3. 2. The vehicle dialogue system according to claim 1, The input control unit receives sensor information from a sensor mounted on the vehicle, The input control unit inputs the sensor information when the driver speaks to the interactive AI as the state of the vehicle. Dialogue system for vehicles.
4. 2. The vehicle dialogue system according to claim 1, The input control unit receives a detection result from an abnormality detection unit that detects an abnormality in the vehicle and turns on a warning lamp, The input control unit inputs the detection result of the abnormality detection unit when the driver speaks to the interactive AI as the state of the vehicle. Dialogue system for vehicles.
5. 2. The vehicle dialogue system according to claim 1, A device control unit for controlling vehicle devices, The input control unit inputs text data indicating information of the vehicle equipment that can be controlled by the equipment control unit to the interactive AI, The equipment control unit controls the vehicle equipment based on the text data converted by the voice recognition unit and the response information from the conversational AI. Dialogue system for vehicles.
6. 2. The vehicle dialogue system according to claim 1, A second determination unit that determines the number of people in the vehicle, The input control unit inputs the number of people determined by the second determination unit to the interactive AI. Dialogue system for vehicles.
Citation Information
Patent Citations
Voice interacting device
JP2003108191A
Interactive device and interactive method
JP2017067849A
Information processing device and information processing method
JP2020112732A
Information processing device, information processing method, and program
JP2022103675A
Information processing apparatus, information processing system, and information processing method
WO2019064928A1