Conversation system for vehicle
By combining vehicle sensor information and driver status detection with a conversational AI system, the problem of natural response of voice agent services in vehicles has been solved, enabling appropriate responses to driver voice and control of vehicle equipment, thus improving the naturalness and applicability of the vehicle dialogue system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YAZAKI CORP
- Filing Date
- 2024-10-18
- Publication Date
- 2026-05-05
AI Technical Summary
Existing vehicle voice agent services fail to provide natural conversational responses and cannot obtain appropriate responses unless appropriate prompts are entered, especially when using conversational AI in vehicles, where driver and vehicle status information is not effectively utilized.
The system employs a conversational AI system, which combines vehicle sensor information and driver status detection with a voice input unit, a voice recognition unit, an input control unit, a voice synthesis unit, and a voice output unit to achieve appropriate responses to the driver's voice.
It provides appropriate responses to driver voice commands, understands driver and vehicle status, adapts to different driving loads and alarm conditions, controls vehicle equipment, and enables natural dialogue.
Smart Images

Figure CN121986376A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a dialogue system for vehicles. Background Technology
[0002] A voice proxy service that responds via voice when a user speaks has been proposed (Patent Document 1). The voice proxy service in Patent Document 1 proposes returning pre-defined chat data when the user's spoken content is not a search. However, in the voice proxy service of Patent Document 1, since the chat data is extracted from pre-defined data, a natural response, such as that of conversing with a person, is not returned.
[0003] In recent years, conversational AI (artificial intelligence) has been proposed, capable of learning from the vast amounts of information available on the internet and achieving more natural conversations, such as ChatGPT. However, unless appropriate prompts (text data) are input, this conversational AI will not return appropriate responses. Therefore, when using conversational AI in vehicles, appropriate responses may not be obtained unless the driver states the driver's status, the vehicle's status, etc., one by one.
[0004] Reference List
[0005] Patent documents
[0006] Patent Document 1: JP2014-098844A Summary of the Invention
[0007] Technical issues
[0008] The present invention was made in view of the above circumstances, and the object of the present invention is to provide a dialogue system for a vehicle that responds appropriately to the driver's voice.
[0009] Solution to the problem
[0010] To achieve the above objectives, the dialogue system for vehicles according to the present invention has the following features.
[0011] A dialogue system for a vehicle using conversational AI is provided, which outputs response information including text data when it receives input information including text data. The dialogue system for a vehicle includes: a voice input unit configured to input voice spoken by a driver; a voice recognition unit configured to convert the voice input by the voice input unit into text data; an input control unit configured to input the text data converted by the voice recognition unit, along with text data representing the state of the driver or the vehicle driven by the driver at the time of the driver's speech, or text data of a command corresponding to the state, as input information to the conversational AI; a voice synthesis unit configured to convert the response information from the conversational AI into speech; and a voice output unit configured to output the speech converted by the voice synthesis unit.
[0012] Beneficial effects of the present invention
[0013] According to the present invention, a dialogue system for a vehicle that responds appropriately to the driver's voice can be provided.
[0014] The invention has been briefly described above. Details of the invention can be understood by referring to the following description of embodiments for carrying out the invention (hereinafter referred to as "Examples") in conjunction with the accompanying drawings. Attached Figure Description
[0015] Figure 1 This is a block diagram illustrating a dialogue system for a vehicle according to an embodiment of the present invention;
[0016] Figure 2 It is shown Figure 1 The flowchart shown illustrates the processing procedure of the microcomputer.
[0017] Figure 3 This is a schematic diagram illustrating an example of text data transmitted from a dialogue system for a vehicle to a conversational AI when the ignition is turned on.
[0018] Figure 4 This is a schematic diagram illustrating an example of a dialogue between a vehicle's dialogue system and the driver;
[0019] Figure 5 This is a schematic diagram illustrating an example of a dialogue system used in a vehicle and its interaction with the driver; and
[0020] Figure 6 This is a schematic diagram illustrating an example of a dialogue system used in a vehicle and its interaction with the driver.
[0021] Reference tag list
[0022] 1. Dialogue system for vehicles
[0023] 2 microphones (voice input unit)
[0024] 5. Voice output unit
[0025] 10 Conversational AI
[0026] 41 Speech Recognition Unit
[0027] 42. Voice dialogue unit (input control unit, first decision unit, device control unit, second decision unit)
[0028] 43 Speech Synthesis Unit
[0029] S1 Input Information
[0030] S2 Response Information Detailed Implementation
[0031] Specific embodiments of the present invention will now be described with reference to the accompanying drawings.
[0032] The dialogue system 1 for a vehicle according to this embodiment is a system installed in a vehicle and communicating with the driver using conversational artificial intelligence (AI) 10. The conversational AI 10 is implemented, for example, ChatGPT, and outputs response information S2 including text data when it receives input information S1 including text data.
[0033] The dialogue system 1 for vehicles includes a microphone 2 as a voice input unit, a communication module 3, a microcomputer 4, and a speaker 5 as a voice output unit. The microphone 2 inputs the driver's spoken voice into the microcomputer 4. The communication module 3 is used to communicate with the conversational AI 10 via an Internet communication network (not shown), and includes circuitry, an antenna, etc., for connecting to the Internet communication network.
[0034] The microcomputer 4 includes: memory, such as random access memory (RAM) and read-only memory (ROM); and a central processing unit (CPU) that operates according to a program stored in the memory and controls the entire dialogue system 1 for the vehicle.
[0035] The microcomputer 4 includes: a speech recognition unit 41; a voice dialogue unit 42 serving as an input control unit, a first determination unit, a device control unit, and a second determination unit; and a speech synthesis unit 43. The speech recognition unit 41 converts speech input from the microphone 2 into text data and inputs this text data to the voice dialogue unit 42. The voice dialogue unit 42 adds text data representing the state of the driver or the vehicle being driven by the driver when the driver speaks, or a command corresponding to that state, to the text data converted by the speech recognition unit 41, and inputs this text data as input information S1 to the conversational AI 10.
[0036] In addition, sensor information S3 from multiple sensors installed on the vehicle is input to the voice dialogue unit 42. Examples of sensors include an illuminance sensor that measures the illuminance outside the vehicle, a temperature sensor that measures the outside air temperature, a Global Positioning System (GPS) that detects the vehicle's position, a seat sensor that detects whether a person is sitting in a seat, and a speed sensor that detects the vehicle's speed. The sensor information S3 received from the sensors can be not only the measured values obtained by the sensors, but also the moving average and weighted average of the measured values, as well as values calculated using the measured values and the moving average and weighted average of the measured values.
[0037] In addition, personnel information S4 is input to the voice dialogue unit 42. This personnel information S4 is the detection result from the driver monitor, which detects the driver's state (whether the driver is dozing off, driving carelessly, or not paying attention) based on an image obtained by imaging the driver's face. Furthermore, alarm information S5 is input to the voice dialogue unit 42. This alarm information S5 is the detection result from the anomaly detection unit that detects vehicle anomalies (engine anomaly, hydraulic anomaly, water temperature anomaly, charging anomaly, etc.) and illuminates the warning lights.
[0038] The voice interaction unit 42 is connected to and can control the vehicle equipment 11 installed in the vehicle. Examples of vehicle equipment 11 include air conditioning installed in the vehicle, displays such as head-up displays, and audio equipment.
[0039] The voice dialogue unit 42 receives response information S2 from the conversational AI 10 and outputs the received response information S2 to the speech synthesis unit 43. The speech synthesis unit 43 converts the response information S2 into speech and outputs the speech to the speaker 5. The speaker 5 outputs the speech converted by the speech synthesis unit 43.
[0040] Next, we will refer to Figure 2The flowchart shown describes the operation of the dialogue system 1 for a vehicle with the above configuration. When ignition on (IG-ON) is detected (Sp1), the microcomputer 4 sends text data representing the status of the driver and vehicle to the conversational AI 10 (Sp2). As an example of the text data representing the status of the driver and vehicle, such as... Figure 3 As shown, this includes the ability to use a message that means "The conversation partner is about to enter the vehicle. Please consider that the person is inside the vehicle when conducting the conversation." Therefore, the conversational AI 10 understands the state of the conversation partner (driver) inside the vehicle, and thereafter, the content of the conversation is based on the premise that the conversation partner is inside the vehicle.
[0041] In Sp2, the microcomputer 4 can also send text data containing information about the vehicle equipment 11 that it can control. As an example of text data, a message could be used stating, "Based on the conversation with the dialogue partner, I can suggest adjusting the air conditioning temperature and fan speed, adjusting the head-up display (HUD) brightness, and controlling the audio." Therefore, the conversational AI 10 can understand that in the vehicle where the driver is located, it can adjust the air conditioning temperature and fan speed, adjust the head-up display brightness, and control the audio settings.
[0042] In Sp2, the microcomputer 4 can also determine the number of people in the vehicle based on detection results from the seat sensors and send the determined number to the conversational AI 10. As an example of text data, a message indicating "only the driver is in the vehicle" or "two people, including the driver, are in the vehicle" can be used. Therefore, the conversational AI 10 can understand whether the driver is alone or whether another person is also in the vehicle.
[0043] Subsequently, the microcomputer 4 acquires the aforementioned sensor information S3, personnel information S4, and alarm information S5 (Sp3). Then, when the driver speaks (Y in Sp4), the microcomputer 4 performs speech recognition processing (Sp5) to convert the driver's speech into text data. Next, the microcomputer 4 adds the driver's and vehicle's state at the time of the driver's speech, or the command corresponding to that state (text data), to the text data converted by the speech recognition processing (Sp6), and sends this text data as input information S1 to the conversational AI 10 (Sp7).
[0044] In Sp6, the microcomputer 4 can add text data acquired in Sp3, representing sensor information S3, personnel information S4, and alarm information S5, as text data indicating the driver's and vehicle's status when the driver speaks. The microcomputer 4 can also add text data indicating whether the vehicle is being driven, stopped, or waiting at a traffic light as text data indicating the vehicle's status.
[0045] Furthermore, the microcomputer 4 determines whether the driver's driving load is high when the driver speaks. This determination can be based on, for example, speed from a speed sensor (sensor information S3) and personnel information S4. It can also be based on pre-registered driver attributes (whether the driver is skilled at driving and familiar with the vehicle's equipment). When it is determined that the driving load is high when the driver speaks, the microcomputer 4 can add text data to shorten the response information S2 as a command corresponding to the state. An example of such text data could be a message indicating "Please provide a brief response."
[0046] Subsequently, microcomputer 4 receives response information S2 corresponding to the input information S1 sent in Sp7 (Sp8), converts the received response information S2 into speech, and outputs the speech from speaker 5 (Sp9). Next, when an ignition device disconnection (IG-OFF) (Y in Sp10) is detected, microcomputer 4 terminates the process. On the other hand, when no ignition device disconnection is detected (N in Sp10), microcomputer 4 returns to Sp3.
[0047] After executing Sp3, if the driver does not speak (N in Sp4), then when the predetermined proposal condition (Sp11) is met based on the information S3 to S5 obtained from Sp3, the microcomputer 4 converts the proposal content into voice data and outputs the voice data from the speaker 5 (Sp12) and then returns to Sp3. As a proposal condition for Sp11, for example, if the speed measured by the speed sensor (sensor information S3) is high, then the proposal condition for reducing the speed is met, and a proposal to reduce the speed is output from the speaker 5.
[0048] The aforementioned dialogue system 1 for vehicles sends the driver's and vehicle's states at the time the driver speaks to the conversational AI 10. Therefore, the conversational AI 10 is able to respond appropriately to the driver's voice after understanding the driver's and vehicle's states at the time the driver speaks.
[0049] When the driver's workload is high due to speaking, the aforementioned dialogue system 1 for vehicles sends a command to the conversational AI 10 to shorten the response information. Therefore, the response of the conversational AI 10 can be shortened when the driver's workload is high due to speaking.
[0050] For example, such as Figure 4 and Figure 5 The example dialogue shown describes a situation where the warning light illuminates while driving and the driver says, "A red triangle with an exclamation mark appeared, what is this?" This occurs when a high driving load is determined, such as... Figure 4As shown, the microcomputer 4 of the dialogue system 1 for the vehicle adds the text data of "Please summarize your answer" to the text data of the driver's spoken voice, and sends the added text data to the conversational AI 10.
[0051] On the one hand, when the driving load is determined to be low, such as Figure 5 As shown, microcomputer 4 does not add the text data for "Please summarize your answer". Therefore, when the driving load is low, such as Figure 5 As shown, you can return to the long, detailed description written in the vehicle manual: "The red warning light in the shape of an exclamation mark is the main alarm. When other warning lights or indicator lights are on, or when an alarm message is displayed on the multi-information display, the main alarm lights up simultaneously and also sounds a buzzer depending on the content of the alarm."
[0052] On the other hand, when the driving load is high, one can refer back to the simple and concise statement in the vehicle manual: "The red warning light in the shape of an exclamation mark is a major warning. This light illuminates when an emergency of high urgency occurs, so please stop immediately and contact the dealer while driving."
[0053] The microcomputer 4 of the aforementioned vehicle dialogue system 1 sends sensor information S3 from sensors installed on the vehicle as the vehicle's state when the driver speaks. Therefore, the dialogue AI 10 is able to respond more appropriately to the driver's voice after understanding the vehicle's state.
[0054] The microcomputer 4 of the vehicle's dialogue system 1 receives the detection results from the anomaly detection unit when the driver speaks, taking this as the vehicle's state. Therefore, the conversational AI 10 is able to appropriately respond to the driver's voice after understanding that a warning light has illuminated. For example, it can appropriately answer the driver's question, "The light is on, what is this?"
[0055] The microcomputer 4 of the aforementioned dialogue system 1 for vehicles sends text data containing information about vehicle equipment 11 that can be controlled in Sp2. Thus, the conversational AI 10 can control vehicle equipment 11 based on the driver's verbal suggestions. Although in Figure 2 The flowchart is not shown, but the microcomputer 4 controls the vehicle equipment 11 based on text data converted by the speech recognition unit 41 and response information S2 from the conversational AI 10.
[0056] For example, as in Figure 6As shown in the dialogue example, the system describes the situation where the driver says, "The sun is brighter than usual today." In addition to the text data "The sun is brighter than usual today," the vehicle's dialogue system 1 sends the vehicle's location information and illuminance information (text data) to the conversational AI 10 in Sp7. Furthermore, the vehicle's dialogue system 1 sends information indicating that the HUD can be controlled in Sp2. Therefore, the conversational AI 10 can obtain information about the weather and temperature from the location information and predict that the HUD will be difficult to see, and can send the response information S2: "The weather forecast says it will be sunny all day. The expected temperature is 25 degrees Celsius during the day. Is it difficult to see the HUD display under that brightness?"
[0057] The microcomputer 4 of the vehicle dialogue system 1 analyzes the subsequent dialogue between the vehicle dialogue system 1 and the driver, and controls the HUD, which is a vehicle device 11, to increase the brightness by five levels.
[0058] The following configuration can be adopted: In a single dialogue round, the dialogue system 1 for the vehicle adds the driver's and vehicle's status to the driver's initial utterance, but does not add the driver's and vehicle's status to subsequent utterances. That is, in... Figure 6 In the dialogue example shown, the dialogue system 1 for the vehicle only adds the status of the driver and the vehicle to the text data of the driver's first utterance, "The sun is brighter than usual today," and does not add the driver and vehicle status information to the subsequent text data of "Speaking of which, it may be harder to see than usual" and "It's okay, thank you."
[0059] The microcomputer 4 of the aforementioned dialogue system 1 for vehicles sends the number of people in the vehicle to the conversational AI 10. Therefore, the conversational AI 10 responds appropriately to the driver's voice after understanding whether the driver is alone or if others are also in the vehicle.
[0060] This invention is not limited to the above embodiments, and can be appropriately modified and improved. The material, shape, size, quantity, and arrangement of the components in the above embodiments are freely selectable and not limited, as long as the invention can be realized.
[0061] According to the above embodiment, information about the controllable vehicle equipment 11, which does not change constantly, and the number of people in the vehicle, is sent immediately after the ignition is turned on. Alternatively, the invention is not limited thereto. Information about the controllable vehicle equipment 11 and the number of people in the vehicle may also be added to the text data of the voice spoken by the driver and sent.
[0062] According to the above embodiment, the driver's state is determined based on an image obtained by imaging the driver's face. Alternatively, the invention is not limited thereto. The driver's state can be determined based, for example, measurements from a heart rate monitor worn by the driver.
[0063] According to the above embodiment, after the microcomputer 4 fully broadcasts the response information S2 in Sp9, it returns to Sp4 and receives the driver's voice. Alternatively, the invention is not limited thereto. The microcomputer 4 may have a so-called interruption function. In the interruption function, when a voice is detected while the response information S2 is being broadcast in Sp9, the microcomputer 4 may stop broadcasting, return to Sp5, and perform speech recognition on the driver's voice. The interruption function can be turned on and off by a user such as the driver.
[0064] During the process of broadcasting response information S2, the microcomputer 4 can broadcast response information S2 in response to the driver's interruption, and then rebroadcast the response information S2 that was being broadcast previously.
[0065] When the speech recognition result in Sp5 is a stop speech command such as "stop talking" or "too much talk," the microcomputer 4 can stop proceeding to Sp7, which sends the speech recognition result to the conversational AI 10, and instead return to Sp4 to receive the driver's next speech. That is, in the case of a stop speech command, the microcomputer 4 will not replay the previously played response information S2.
[0066] In cases other than command messages to stop speaking, the microcomputer 4 sends the voice results, etc., to the conversational AI 10's Sp7.
[0067] According to the above embodiment, the dialogue system 1 for a vehicle communicates with the conversational AI 10 via an internet communication network, but the invention is not limited thereto. The conversational AI 10 can be mounted on a control panel in the vehicle. In this case, dialogue can be conducted even in poor communication environments.
[0068] Here, the features of the embodiments of the dialogue system for vehicles according to the present invention described above are briefly summarized and listed in the following [1] to [6]. [1]
[0070] A dialogue system (1) for a vehicle using a conversational AI (10), the conversational AI (10) outputting response information (S2) including text data when receiving input information (S1) including text data, the dialogue system (1) for a vehicle comprising:
[0071] A voice input unit (2) is configured to input a voice issued by the driver;
[0072] A speech recognition unit (41) is configured to convert speech input by the speech input unit (2) into text data;
[0073] The input control unit (42) is configured to input text data converted by the speech recognition unit (41) and text data representing the state of the driver or the vehicle driven by the driver when the driver speaks, or text data of the command corresponding to the state, as input information (S1) into the conversational AI (10).
[0074] A speech synthesis unit (43) configured to convert response information (S2) from a conversational AI (10) into speech; and
[0075] The speech output unit (5) is configured to output speech converted by the speech synthesis unit (43).
[0076] According to the configuration in [1], the conversational AI (10) is able to respond appropriately to the driver's voice after understanding the state of the driver and the vehicle when the driver speaks. [2]
[0078] In the dialogue system (1) for vehicles described in [1],
[0079] The dialogue system (1) for the vehicle also includes a first determination unit (42) configured to determine whether the driver's driving load is high, wherein,
[0080] When the first determination unit (42) determines that the driver’s driving load is high when the driver speaks, the input control unit (42) inputs the text data used to shorten the response information (S2) as text data representing the command into the conversational AI (10).
[0081] According to the configuration in [2], when the driver’s driving load is high when the driver speaks, the response of the conversational AI (10) can be shortened. [3]
[0083] In the dialogue system (1) for vehicles described in [1],
[0084] Sensor information from sensors installed on the vehicle is input to the input control unit (42), and
[0085] The input control unit (42) inputs the sensor information when the driver speaks as the vehicle status into the conversational AI (10).
[0086] According to the configuration in [3], the conversational AI (10) is able to respond more appropriately to the driver's voice after understanding the state of the vehicle. [4]
[0088] In the dialogue system (1) for vehicles described in [1],
[0089] The detection results from the anomaly detection unit, which detects anomalies in the vehicle and illuminates the warning lights, are input to the input control unit (42).
[0090] The input control unit (42) inputs the detection result of the abnormality detection unit when the driver speaks as the vehicle status into the conversational AI (10).
[0091] According to the configuration in [4], the conversational AI (10) is able to respond appropriately to the driver's voice after understanding that the warning light is on. [5]
[0093] In the dialogue system (1) for vehicles described in [1],
[0094] The dialogue system (1) for the vehicle also includes a device control unit configured to control vehicle equipment (11), wherein,
[0095] The input control unit (42) inputs text data representing information about vehicle equipment (11) that can be controlled by the device control unit (42) into the conversational AI (10), and
[0096] The device control unit (42) controls the vehicle equipment (11) based on text data converted by the voice recognition unit (41) and response information (S2) from the conversational AI (10).
[0097] According to the configuration in [5], the conversational AI (10) is able to control the vehicle equipment (11) based on the driver's verbal suggestions. [6]
[0099] In the dialogue system (1) for vehicles described in [1],
[0100] The dialogue system (1) for vehicles also includes a second determination unit (42) configured to determine the number of people in the vehicle, and
[0101] The input control unit (42) inputs the number of people determined by the second determination unit (42) into the conversational AI (10).
[0102] According to the configuration in [6], the conversational AI (10) responds appropriately to the driver’s voice after understanding whether the driver is alone or whether other people are also in the vehicle.
[0103] Although the present invention has been described in detail with reference to specific embodiments, it will be apparent to those skilled in the art that various changes and modifications can be made without departing from the spirit and scope of the invention.
[0104] This application is based on Japanese patent application (JP2023-181851A) filed on October 23, 2023, the contents of which are incorporated herein by reference.
[0105] Industrial applicability
[0106] According to the present invention, a dialogue system for a vehicle that appropriately responds to the driver's voice can be provided. The present invention, having this effect, can be used in a vehicle dialogue system.
Claims
1. A dialogue system for vehicles using conversational AI, wherein the conversational AI outputs response information including text data when it receives input information including text data, the dialogue system for vehicles comprising: A voice input unit configured to input voice produced by the driver; A speech recognition unit configured to convert the speech input by the speech input unit into text data; An input control unit is configured to input the text data converted by the speech recognition unit, along with text data representing the state of the driver or the vehicle driven by the driver at the time the driver speaks, or text data of a command corresponding to that state, as input information into the conversational AI. A speech synthesis unit configured to convert the response information from the conversational AI into speech; and A voice output unit configured to output the voice converted by the voice synthesis unit.
2. The dialogue system for vehicles according to claim 1, further comprising: A first determination unit is configured to determine whether the driver's driving load is high, wherein... If the first determination unit determines that the driving load is high when the driver speaks, the input control unit will input the text data used to shorten the response information as the text data representing the command to the conversational AI.
3. The dialogue system for a vehicle according to claim 1, wherein, Sensor information from sensors installed on the vehicle is input to the input control unit, and The input control unit inputs the sensor information when the driver speaks as the vehicle's status into the conversational AI.
4. The dialogue system for a vehicle according to claim 1, wherein, The detection result from the anomaly detection unit, which detects anomalies in the vehicle and illuminates the warning lights, is input to the input control unit, and The input control unit inputs the detection result of the anomaly detection unit when the driver speaks as the vehicle's status into the conversational AI.
5. The dialogue system for a vehicle according to claim 1, further comprising: The equipment control unit is configured to control vehicle equipment, wherein... The input control unit inputs text data representing information about the vehicle equipment that can be controlled by the device control unit into the conversational AI, and The device control unit controls the vehicle equipment based on text data converted by the speech recognition unit and response information from the conversational AI.
6. The dialogue system for a vehicle according to claim 1, further comprising: The second determination unit is configured to determine the number of people in the vehicle, wherein... The input control unit inputs the number of people determined by the second determination unit into the conversational AI.
Citation Information
Patent Citations
Interaction support device, interaction system, interaction support method, and program
JP2014098844A
Rotary electric machine
JP2023181851A