Vehicle dialogue system
The vehicle dialogue system addresses the lack of intuitive feedback in conventional voice agent services by using interactive AI and visual system state indicators, significantly improving user experience and convenience.
Patent Information
- Application Number
- JP2023181853
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-23
- Publication Date
- 2025-05-08
AI Technical Summary
Conventional voice agent services in vehicles often lack intuitive display feedback, making it difficult for users to understand whether the system is actively recognizing voice input or not, which can lead to user frustration and difficulty in usage.
A vehicle dialogue system that employs interactive AI, incorporating a voice input unit, speech recognition, input control, speech synthesis, audio output, and a drawing processor to convert system states into visible information for intuitive display.
The system allows users to intuitively grasp the system's state, enhancing user convenience by providing clear visual feedback during voice recognition, response generation, and audio output states.
Smart Images

Figure 2025071573000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to a dialogue system for a vehicle. [Background technology]
[0002] Conventionally, a voice agent service has been proposed that responds by voice when a user speaks (see Patent Document 1). In the voice agent service of Patent Document 1, when a user speaks, it is determined whether the content of the utterance is a search or not so that the method of response can be changed according to the user's intention, and if the content of the utterance is not a search, predetermined chat data is returned as a response. In addition, in recent years, a voice agent service using a dialogue AI such as ChatGPT has been proposed that can learn a huge amount of information on the Internet and can have a more natural conversation. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2014-98844 A Summary of the Invention [Problem to be solved by the invention]
[0004] However, in the past, when a voice agent service recognized a user's speech, some displayed an abstract message such as a wavy line on the display mounted on the vehicle, while others had no display at all. This made it difficult for users to intuitively tell whether the agent service (system) was accepting voice recognition, which made it difficult for users to use.
[0005] The present invention has been made in consideration of the above-mentioned circumstances, and has an object to provide a vehicle dialogue system that is highly convenient and enables a user to intuitively grasp the system status. [Means for solving the problem]
[0006] In order to achieve the above object, the vehicle dialogue system according to the present invention has the following features. A dialogue system for a vehicle using a dialogue type AI that outputs response information made of text data when input information made of text data is input, a voice input unit for inputting voice spoken by a user; a voice recognition unit that converts the voice input by the voice input unit into text data; an input control unit that inputs the text data converted by the voice recognition unit and text data indicating the state of the user or the vehicle driven by the user when the user speaks, or a command corresponding to the state, as the input information to the dialogue AI; A voice synthesis unit that converts the response information from the dialogue AI into voice; a voice output unit that outputs the voice converted by the voice synthesis unit; a rendering processing unit that converts a system state of the vehicular dialogue system into visual information; a display unit that outputs the visual information converted by the drawing processing unit, The system state includes at least a standby state in which the user is waiting for speech before the user starts speaking; a voice recognition state in which the voice uttered by the user is recognized from the start of the utterance to the end of the utterance by the user; a response generation state in which the response information is generated during a period from the end of the utterance to the end of the conversion of the response information into the voice; a voice output state in which the voice converted by the voice synthesis unit is outputted from the voice output unit while the voice converted by the voice synthesis unit is outputted from the voice output unit; It is a dialogue system for vehicles. Effect of the Invention
[0007] According to the present invention, the drawing processing unit converts the system states, which transition between a standby state, a voice recognition state, an answer generation state, and a voice output state, into visual information and outputs the information to the display unit. This allows the user to intuitively understand the state of the system. In other words, the present invention is more convenient than the conventional technology.
[0008] The present invention has been briefly described above. Furthermore, the details of the present invention will be further clarified by reading the following description of the embodiment of the present invention (hereinafter, referred to as "embodiment") with reference to the accompanying drawings. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing an embodiment of a dialogue system for a vehicle according to the present invention. [Diagram 2] FIG. 2 is a diagram showing an example of state transition of the vehicular dialogue system shown in FIG. [Diagram 3] FIG. 3 is a flowchart illustrating an example of a processing procedure of the microcomputer illustrated in FIG. [Figure 4] FIG. 4 is a flowchart showing an example of a processing procedure of the microcomputer shown in FIG. 1, which is a continuation of the flowchart shown in FIG. [Diagram 5] FIG. 5 is an explanatory diagram showing an example of how a character is displayed. [Figure 6] FIG. 6 is an explanatory diagram showing an example of how a character is displayed. [Figure 7] FIG. 7 is an explanatory diagram showing an example of how a character is displayed. [Figure 8] FIG. 8 is an explanatory diagram showing an example of how a character is displayed. [Figure 9] FIG. 9 is an explanatory diagram showing an example of how a character is displayed. [Figure 10] FIG. 10 is an explanatory diagram showing an example of how a character is displayed. [Figure 11] FIG. 11 is an explanatory diagram showing an example of how characters are displayed. [Figure 12]FIG. 12 is an explanatory diagram showing an example of how a character is displayed. [Figure 13] FIG. 13 is a block diagram showing another embodiment of the vehicular dialogue system of the present invention. [Figure 14] FIG. 14 is an explanatory diagram showing an example of displaying a character based on the state of the vehicle. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Specific embodiments of the present invention will be described below with reference to the accompanying drawings.
[0011] The vehicle dialogue system 1 of this embodiment is a system that is mounted on a vehicle and dialogues with a user (a driver or a passenger in the vehicle) using a dialogue type AI (Artificial Intelligence) 10. The dialogue type AI 10 is configured, for example, by ChatGPT, and when input information S1 consisting of text data is input, response information S2 consisting of text data is output.
[0012] The vehicle dialogue system 1 includes a microphone 2 as a voice input unit, a communication module 3, a microcomputer 4 (hereinafter abbreviated as "microcomputer 4"), a speaker 5 as a voice output unit, and a display 6 as a display unit. The microphone 2 inputs voice spoken by a user to the microcomputer 4. The communication module 3 is for communicating with the dialogue type AI 10 via an Internet communication network (not shown), and is composed of a circuit, an antenna, etc. for connecting to the Internet communication network.
[0013] The microcomputer 4 has memories such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and a CPU (Central Processing Unit) that operates according to programs stored in the memories.
[0014] The microcomputer 4 has a voice recognition unit 41, a voice dialogue unit 42 as an input control unit, a voice synthesis unit 43, and a drawing processing unit 44. The voice recognition unit 41 converts the voice input by the microphone 2 into text data and inputs it to the voice dialogue unit 42. The voice dialogue unit 42 adds text data indicating the state of the user or the vehicle driven by the user when the user spoke, or a command corresponding to the state, to the text data converted by the voice recognition unit 41, and inputs it to the interactive AI 10 as input information S1.
[0015] Furthermore, vehicle information S3 from various devices mounted on the vehicle that detect the vehicle condition is input to the voice dialogue unit 42. The vehicle information S3 is, for example, sensor information from a plurality of sensors and warning information that is a detection result from an abnormality detection unit that detects abnormalities in the vehicle (engine abnormality, oil pressure abnormality, water temperature abnormality, charging abnormality, etc.) and turns on a warning lamp.
[0016] Possible sensors include an illuminance sensor that measures the illuminance outside the vehicle, a temperature sensor that measures the outside air temperature, a GPS (Global Positioning System) that detects the vehicle position, a seating sensor that detects whether or not a person is sitting in a seat, a speed sensor that detects the vehicle speed, etc. The sensor information (vehicle information S3) input from the sensor may be not only a measurement value measured by the sensor, but also a moving average or weighted average of the measurement values, or a value calculated using the measurement value, the moving average or weighted average of the measurement values.
[0017] In addition, the voice dialogue unit 42 receives human information S4, which is the detection result from a driver monitor that detects the driver's state (whether the driver is dozing, driving aimlessly, or not, etc.) based on an image of the driver's face.
[0018] The voice dialogue unit 42 is connected to a vehicle device 11 mounted in the vehicle and can control the vehicle device 11. The vehicle device 11 may be, for example, an air conditioner, an audio device, or the like mounted in the vehicle.
[0019] The voice dialogue unit 42 receives response information S2 from the dialogue-type AI 10 and outputs the received response information S2 to the voice synthesis unit 43. The voice synthesis unit 43 converts the response information S2 into voice and outputs it to the speaker 5. The speaker 5 outputs the voice converted by the voice synthesis unit 43.
[0020] The voice dialogue unit 42 inputs the system state (standby state, voice recognition state, answer generation state, and voice output state) of the vehicular dialogue system 1, and outputs the input system state to the drawing processing unit 44. The drawing processing unit 44 converts the system state into visible information such as character CR (see FIG. 5) and letters ST1, ST2 (see FIG. 11 and FIG. 12), and outputs it to the display 6. The user can set a desired character for the character CR in advance. The display 6 outputs the visible information converted by the drawing processing unit 44. Note that the display 6 may be, for example, a display device such as a meter, HUD, or center display mounted on the vehicle.
[0021] The vehicle dialogue system 1 may be configured such that the processing of the dialogue AI 10, which was previously performed on the cloud, is incorporated into the microcomputer 7 (see FIG. 13). That is, in this configuration, the vehicle dialogue system 1 includes a microphone 2, a microcomputer 4, a speaker 5, a display 6, and a microcomputer 7 incorporating the dialogue AI 10a. With this configuration, dialogue processing is possible even in a poor reception environment.
[0022] Next, a description will be given of the system states and state transitions of the vehicle dialogue system 1. As shown in Fig. 2, the vehicle dialogue system 1 transitions between at least four system states, which are a standby state for waiting for a user's utterance, a voice recognition state for performing voice recognition on the contents of the user's utterance while the user is speaking, an answer generation state for generating an answer (response) to the contents of the utterance after the user's utterance ends, and a voice output state for outputting a voice converted by the voice synthesis unit 43 after the answer generation ends.
[0023] More specifically, as shown in FIG. 2, when the user starts speaking in the standby state, the standby state transitions to the voice recognition state. Similarly, when the user stops speaking in the voice recognition state, the voice recognition state transitions to the answer generation state. Similarly, when the generation of the answer ends in the answer generation state, the answer generation state transitions to the voice output state. Similarly, when the reading out of the answer (voice output) ends in the voice output state, the voice output state transitions to the standby state.
[0024] Next, the operation of the vehicle dialogue system 1 having the above-mentioned configuration will be described with reference to the flowcharts shown in Figs. 3 and 4. When the microcomputer 4 detects the ignition on (IG-ON) or the user's operation, it transmits text data indicating the state of the user and the vehicle to the dialogue AI 10. An example of the text data indicating the state of the user and the vehicle is "The person you are talking to is about to get in the car. Please talk to them while taking into consideration the fact that they are in the car." This allows the dialogue AI 10 to understand that the user is in the car, and the dialogue content thereafter is based on the assumption that the user is in the car. At this time, the microcomputer 4 also acquires the vehicle information S3 and the person information S4.
[0025] When the system is started as described above, the system goes into a standby state waiting for a user's utterance, and the microcomputer 4 outputs the character CR to the display 6 in conjunction with the system state (Sp1).
[0026] At this time, examples of the display of the character CR are as shown in FIG. 5 when the system is started up or during normal operation, whereas in the standby state, for example, the character CR may be asleep (see FIG. 6), or the size of the character CR may be reduced and displayed in the corner of the display 6 (see FIG. 7), or a combination of these may be used.
[0027] Next, if the microcontroller 4 detects the user's speech in the standby state (Yes in Sp2), the system state transitions from the standby state to the voice recognition state, and the microcontroller 4 outputs the character CR to the display 6 in conjunction with the system state (Sp3).
[0028] As a display example of the character CR at this time, the size of the character CR may be the same as that in normal times (see FIG. 5), and the character CR may perform a nodding motion in response to the user's speech (see FIG. 8).
[0029] On the other hand, when the microcomputer 4 does not detect the user's speech in the standby state (No in Sp2), the microcomputer 4 returns to the process of Sp1. That is, in this case, the system state is maintained in the standby state.
[0030] Next, if the user has finished speaking (Yes in Sp4), the microcomputer 4 performs a voice recognition process to convert the user's voice (content) into text data (Sp5). On the other hand, if the user has not finished speaking (No in Sp4), the microcomputer 4 returns to the process of Sp3. That is, in this case, the system state is maintained in the voice recognition state.
[0031] Next, if the display of the user's speech content is set to ON (Yes in Sp6), the microcomputer 4 causes the display 6 to display the text data converted in Sp5 (Sp7, see ST1 in FIG. 11). In other words, in Sp7, the user's speech content is transcribed on the display 6. On the other hand, if the display of the user's speech content is set to OFF (No in Sp6), the microcomputer 4 does not display the text data converted in Sp5 on the display 6 (Sp8).
[0032] In addition, in Sp7, whether or not to display the text data on the display 6 may be switched depending on the state of the vehicle. More specifically, for example, in consideration of safe driving, the text data may be displayed on the display 6 only when the vehicle is stopped. In other words, the text data may not be displayed on the display 6 when the vehicle is traveling. This reduces the risk of the driver staring at the display 6, thereby achieving safe driving.
[0033] Also, in Sp7, the display mode of the text data on the display 6 may be changed depending on the state of the vehicle. More specifically, for example, in consideration of safe driving, the entire text data may be displayed on the display 6 when the vehicle is stopped, but summary text data that summarizes the text data may be displayed on the display 6 when the vehicle is traveling. Note that when the amount of text in the text data is small, it can be displayed without summarizing. This reduces the risk of the driver staring at the display 6, and thus enables safe driving.
[0034] Thereafter, the system state transitions from the voice recognition state to the answer generation state, and the microcomputer 4 outputs the character CR to the display 6 in conjunction with the system state (Sp9).
[0035] As an example of how the character CR is displayed at this time, the character CR may have a facial expression that looks as if it is thinking, and a cloud-shaped speech bubble indicating that the character is thinking may be displayed near the character CR (see FIG. 9).
[0036] Next, the microcontroller 4 adds the state of the user and the vehicle when the user spoke, or commands (text data) corresponding to the state, to the text data converted by the voice recognition processing, and transmits it to the interactive AI 10 as input information S1 (Sp10).
[0037] At this time, the microcomputer 4 may add text data indicating the acquired vehicle information S3 and person information S4 as text data indicating the state of the user and the vehicle when the user spoke. Also, the microcomputer 4 may add text data indicating whether the vehicle is driving, stopped, or waiting at a traffic light as text data indicating the state of the vehicle.
[0038] Next, when the microcontroller 4 receives response information S2 corresponding to the input information S1 sent in Sp10 (Yes in Sp11), the system state transitions from the answer generation state to the voice output state, and the microcontroller 4 outputs the character CR to the display 6 in conjunction with the system state (Sp12), and converts the received response information S2 into voice and outputs it from the speaker 5 (Sp13).
[0039] Possible display examples of the character CR at this time include the character CR moving its mouth in accordance with the sound output from the speaker 5, i.e., the character CR acting as if it is speaking, or a so-called cartoon symbol representing the character CR's speaking appearance being displayed near its mouth, or a combination of these (see FIG. 10).
[0040] When reading out the answer in Sp13, if the display of the read-out contents is set to ON (Yes in Sp14), the microcomputer 4 displays the response information S2 received in Sp11 (more specifically, the answer to be read out in Sp13) on the display 6 (Sp15, see ST2 in Figs. 11 and 12). In other words, in Sp13, the answer to be read out is transcribed on the display 6.
[0041] As an example of the display of the read-out contents at this time, when the spoken contents are displayed (Yes in Sp6, Sp7), the spoken contents text ST1 and the read-out contents text ST2 may be displayed in a so-called chat format (see FIG. 11).
[0042] On the other hand, when the spoken content is not displayed in Sp8 (No in Sp6, Sp8), a possible display example of the read-out content is that the text ST2 of the read-out content is displayed in a form similar to so-called subtitles.
[0043] In addition, in Sp15, whether or not to display the response information S2 on the display 6 may be switched depending on the state of the vehicle. More specifically, for example, in consideration of safe driving, the response information S2 may be displayed on the display 6 only when the vehicle is stopped. In other words, the response information S2 may not be displayed on the display 6 when the vehicle is traveling. This reduces the risk that the driver will focus on the display 6, and thus allows for safe driving.
[0044] Also, in Sp15, the display mode of the response information S2 on the display 6 may be changed depending on the state of the vehicle. More specifically, for example, in consideration of safe driving, the entire text of the response information S2 may be displayed on the display 6 when the vehicle is stopped, but summary information that summarizes the response information S2 may be displayed on the display 6 when the vehicle is moving. Note that when the amount of text in the response information S2 is small, it can be displayed without being summarized. This reduces the risk of the driver staring at the display 6, and thus enables safe driving.
[0045] On the other hand, if the display of the read-out contents is set to OFF (No in Sp14), the microcomputer 4 does not display the response information S2 received in Sp11 on the display 6 (Sp16, see FIG. 10).
[0046] Thereafter, when the reading of the answer read out in Sp13 is completed (Yes in Sp17), the microcomputer 4 detects whether the system is on or off (Sp18). On the other hand, when the reading of the answer read out in Sp13 is not completed (No in Sp17), the microcomputer 4 returns to the process in Sp12. That is, in this case, the system state is maintained in the voice output state.
[0047] If the system is OFF (Yes in Sp18), the microcomputer 4 ends the process. Accordingly, the display of the character CR may be ended, or the character may be displayed as shown in FIG.
[0048] In contrast, if the system is ON (No at Sp18), the microcontroller 4 returns to processing at Sp1, the system state transitions from the voice output state to the standby state, and the microcontroller 4 outputs the character CR to the display 6 in conjunction with the system state (Sp1, see Figures 6 and 7).
[0049] As described above, in this embodiment, the display of the character CR is output according to the state transition so that the user can intuitively grasp the system state. On the other hand, the character CR can also be displayed according to the driving situation.
[0050] For example, after the microcomputer 4 detects the ignition on (IG-ON) or the user's operation and starts up the system, the character CR can be displayed on the windshield WS in the HUD and can point to navigation or accident-prone locations based on the acquired vehicle information S3 (see FIG. 14). This allows the user to intuitively grasp not only the system status but also the vehicle status (e.g., driving status, etc.).
[0051] The character CR may be configured to be movable between the center display, the HUD, and an electronic mirror such as a rearview mirror.
[0052] The present invention is not limited to the above-described embodiment, and can be appropriately modified, improved, etc. In addition, the material, shape, size, number, arrangement location, etc. of each component in the above-described embodiment are arbitrary as long as the present invention can be achieved, and are not limited.
[0053] As described above, according to this embodiment, the drawing processing unit 44 converts the system states, which transition between a standby state, a voice recognition state, an answer generation state, and a voice output state, into visual information and outputs the information to the display 6. This allows the user to intuitively understand the state of the system. That is, this embodiment is more convenient than the conventional embodiment.
[0054] Furthermore, according to this embodiment, the character CR as the visible information is displayed in a different manner corresponding to each system state (see Figs. 6 to 10). This allows the user to intuitively understand the state of the system by looking at the character CR.
[0055] Furthermore, according to this embodiment, the drawing processing unit 44 displays the text data converted by the voice recognition unit 41 as visible information on the display 6, allowing the user to check whether the system is correctly recognizing his / her speech and to obtain all information provided by the system via voice without any omissions.
[0056] Furthermore, according to this embodiment, the drawing processing unit 44 displays the response information S2 as visible information on the display 6, allowing the user to check whether the system is correctly recognizing his or her speech, and to obtain all information provided by the system via voice without any omissions.
[0057] Furthermore, according to this embodiment, by making the display of the character CR smaller in the standby state, it is possible to avoid obstructing other information when the user wishes to see it, thereby improving the visibility of the display.
[0058] Furthermore, according to this embodiment, by visualizing the contents of the user's speech and the voice output by the system, the user can feel reassured that the system is operating normally, and can obtain information from the system without missing anything. In other words, according to this embodiment, information can be conveyed to the user accurately.
[0059] Here, the features of the above-described embodiment of the vehicular dialogue system according to the present invention will be briefly summarized and listed in the following [1] to [6].
[0060] [1] A dialogue system (1) for a vehicle using a dialogue type AI (10) that outputs response information (S2) made of text data when input information made of text data is input, A voice input unit (microphone 2) for inputting voice spoken by a user; a voice recognition unit (41) that converts the voice inputted through the voice input unit (microphone 2) into text data; an input control unit (voice dialogue unit 42) that inputs the text data converted by the voice recognition unit (41) and text data indicating the state of the user or the vehicle driven by the user when the user speaks, or an instruction corresponding to the state, as the input information (S1) to the dialogue AI (10); a voice synthesis unit (43) that converts the response information (S2) from the dialogue AI (10) into voice; a voice output unit (speaker 5) that outputs the voice converted by the voice synthesis unit (43); a rendering processing unit (44) for converting the system state of the vehicle dialogue system (1) into visual information; a display unit (display 6) that outputs the visual information converted by the drawing processing unit (44), The system state includes at least a standby state in which the user is waiting for speech before the user starts speaking; a voice recognition state in which the voice uttered by the user is recognized from the start of the utterance to the end of the utterance by the user; a response generation state in which the response information (S2) is generated during a period from the end of the utterance to the end of the conversion of the response information (S2) into the voice; a voice output state in which the voice converted by the voice synthesis unit (43) is outputted from the voice output unit (speaker 5), while the voice converted by the voice synthesis unit (43) is outputted from the voice output unit (speaker 5), Vehicle dialogue system (1).
[0061] According to the configuration of [1] above, the drawing processing unit converts the system state, which transitions between the standby state, the voice recognition state, the answer generation state, and the voice output state, into visual information and outputs it to the display unit. This allows the user to intuitively understand the state of the system. In other words, this configuration is more convenient than the conventional one.
[0062] [2] The vehicle dialogue system (1) according to the above [1], The visual information includes a character (CR) preset by the user, The character (CR) is Each of the system states is displayed in a different manner. Vehicle dialogue system (1).
[0063] According to the configuration of [2] above, the characters as visual information are displayed in different ways corresponding to each system state, allowing the user to intuitively understand the state of the system by looking at the characters.
[0064] [3] The vehicle dialogue system (1) according to the above [1] or [2], When the voice uttered by the user is set to be displayed on the display unit (display 6), the drawing processing unit (44) causes the display unit (display 6) to display the text data converted by the voice recognition unit (41) as the visible information only when the vehicle is stopped. Vehicle dialogue system (1).
[0065] According to the configuration of [3] above, the drawing processing unit displays the text data converted by the voice recognition unit as visible information on the display unit, so that the user can check whether the system is correctly recognizing the user's speech, and can obtain all information provided by voice from the system without missing anything. Also, according to this configuration, the drawing processing unit displays the text data converted by the voice recognition unit as visible information on the display unit only when the vehicle is stopped. In other words, the text data converted by the voice recognition unit is not displayed as visible information on the display unit when the vehicle is traveling. In this way, this configuration takes safety driving into consideration, and the driver can achieve safe driving.
[0066] [4] The vehicle dialogue system (1) according to the above [1] or [2], When the voice uttered by the user is set to be displayed on the display unit (display 6), the drawing processing unit (44) displays the text data converted by the voice recognition unit (41) as the visible information on the display unit (display 6) when the vehicle is stopped, and displays summary information summarizing the text data converted by the voice recognition unit as the visible information on the display unit (display 6) when the vehicle is traveling. Vehicle dialogue system (1).
[0067] According to the configuration of [4] above, the drawing processing unit displays the text data converted by the voice recognition unit as visible information on the display unit, so that the user can check whether the system is correctly recognizing the user's speech, and can obtain all information provided by voice from the system side. Also, according to this configuration, the drawing processing unit displays the text data converted by the voice recognition unit as visible information on the display unit when the vehicle is stopped, and displays summary information that summarizes the text data converted by the voice recognition unit as visible information on the display unit when the vehicle is running, so that safe driving is taken into consideration and the driver can achieve safe driving.
[0068] [5] A vehicle dialogue system (1) according to any one of the items [1] to [4], When the voice converted by the voice synthesis unit (43) is set to be displayed on the display unit (display 6), the drawing processing unit (44) displays the response information (S2) as the visible information on the display unit (display 6) only when the vehicle is stopped. Vehicle dialogue system (1).
[0069] According to the configuration of [5] above, the drawing processing unit displays the response information as visible information on the display unit, so that the user can check whether the system is correctly recognizing the user's speech, and can obtain all information provided by the system through voice. Also, according to this configuration, the drawing processing unit displays the response information as visible information on the display unit only when the vehicle is stopped. In other words, the response information is not displayed as visible information on the display unit when the vehicle is moving. In this way, this configuration takes safe driving into consideration, and the driver can achieve safe driving.
[0070] [6] A vehicle dialogue system (1) according to any one of the items [1] to [4], When the voice converted by the voice synthesis unit (43) is set to be displayed on the display unit (display 6), the drawing processing unit (44) displays the response information (S2) as the visible information on the display unit (display 6) when the vehicle is stopped, and displays summary information summarizing the response information (S2) as the visible information on the display unit (display 6) when the vehicle is traveling. Vehicle dialogue system (1).
[0071] According to the configuration of [6] above, the drawing processing unit displays the response information as visible information on the display unit, so that the user can check whether the system is correctly recognizing the user's speech, and can obtain all information provided by the system through voice. Also, according to this configuration, the drawing processing unit displays the response information as visible information on the display unit when the vehicle is stopped, and displays summary information summarizing the response information as visible information on the display unit when the vehicle is moving, so that safe driving is taken into consideration and the driver can achieve safe driving. [Explanation of symbols]
[0072] 1 Vehicle dialogue system 2 Microphone (audio input section) 3. Communication Module 4. Microcomputer 5 Speaker (audio output section) 6 Display (Display unit) 10 Conversational AI 41 Voice Recognition Unit 42 Voice dialogue section 43 Voice synthesis section 44 Drawing processing section CR character S1 Input information S2 Response Information S3 vehicle information S4 People information ST1,ST2 characters
Claims
1. A dialogue system for a vehicle using a dialogue type AI that outputs response information made of text data when input information made of text data is input, a voice input unit for inputting voice spoken by a user; a voice recognition unit that converts the voice input by the voice input unit into text data; an input control unit that inputs the text data converted by the voice recognition unit and text data indicating the state of the user or the vehicle driven by the user when the user speaks, or an instruction corresponding to the state, as the input information to the interactive AI; A voice synthesis unit that converts the response information from the conversational AI into voice; a voice output unit that outputs the voice converted by the voice synthesis unit; a rendering processing unit that converts a system state of the vehicular dialogue system into visual information; a display unit that outputs the visual information converted by the drawing processing unit, The system state includes at least a standby state in which the user is waiting for speech before the user starts speaking; a voice recognition state in which the voice uttered by the user is recognized from the start of the utterance to the end of the utterance by the user; a response generation state in which the response information is generated during a period from the end of the utterance to the end of the conversion of the response information into the voice; a voice output state in which the voice converted by the voice synthesis unit is outputted from the voice output unit while the voice converted by the voice synthesis unit is outputted from the voice output unit; Dialogue system for vehicles.
2. 2. A vehicle dialogue system according to claim 1, the visual information includes a character preset by the user; The character is: Each of the system states is displayed in a different manner. Dialogue system for vehicles.
3. 2. A vehicle dialogue system according to claim 1, When the voice spoken by the user is set to be displayed on the display unit, the drawing processing unit causes the display unit to display the text data converted by the voice recognition unit as the visible information only when the vehicle is stopped. Dialogue system for vehicles.
4. 2. A vehicle dialogue system according to claim 1, When the voice uttered by the user is set to be displayed on the display unit, the drawing processing unit causes the display unit to display the text data converted by the voice recognition unit as the visible information when the vehicle is stopped, and causes the display unit to display summary information summarizing the text data converted by the voice recognition unit as the visible information when the vehicle is traveling. Dialogue system for vehicles.
5. 2. A vehicle dialogue system according to claim 1, When the voice converted by the voice synthesis unit is set to be displayed on the display unit, the drawing processing unit causes the display unit to display the response information as the visible information only when the vehicle is stopped. Dialogue system for vehicles.
6. 2. A vehicle dialogue system according to claim 1, When the voice converted by the voice synthesis unit is set to be displayed on the display unit, the drawing processing unit causes the display unit to display the response information as the visible information when the vehicle is stopped, and causes the display unit to display summary information summarizing the response information as the visible information when the vehicle is traveling. Dialogue system for vehicles.
Citation Information
Patent Citations
Interaction support device, interaction system, interaction support method, and program
JP2014098844A