Vehicle dialogue system

The vehicle dialogue system addresses the limitations of conventional voice agent services by using interactive AI to process passenger voice inputs and generate natural responses, enhancing passenger convenience.

JP2025071572APending Publication Date: 2025-05-08YAZAKI CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023181852
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-23
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Conventional voice agent services in vehicles lack the ability to provide natural and appropriate responses to passenger queries, as they rely on pre-stored response data rather than dynamic interaction.

Method used

A vehicle dialogue system utilizing interactive AI, which includes a voice input unit, speech recognition unit, control unit, speech synthesis unit, and audio output unit, to process passenger voice inputs, understand passenger attributes and vehicle conditions, and generate natural responses.

Benefits of technology

The system enables interactive AI to provide appropriate and natural responses to passenger queries, improving passenger convenience by understanding passenger attributes and vehicle conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071572000001_ABST
    Figure 2025071572000001_ABST
Patent Text Reader

Abstract

To provide a vehicle dialogue system capable of appropriately responding to voices of passengers and improving passenger convenience.SOLUTION: A vehicle dialogue system 1 is mounted on a passenger vehicle and uses interactive AI 30. A voice recognition unit of an agent 20 included in a control unit 11 converts the voice input by a voice input unit 14 into text data, and the control unit inputs first text data converted by the voice recognition unit and second text data indicating an instruction corresponding to a state of the vehicle, as input information, to the interactive AI. A voice synthesis unit of the agent converts response information from the interactive AI into voice, and a voice output unit 15 outputs the voice converted by the voice synthesis unit. The control unit also controls the input information according to the attributes of passengers.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a vehicle dialogue system to be installed in a passenger vehicle. [Background technology]

[0002] Conventionally, a voice agent service that responds by voice when a user speaks has been proposed (see Patent Document 1). In recent years, such a voice agent service (conversation assistant device) has also been installed in passenger vehicles such as taxis (see Patent Document 2). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2014-98844 A [Patent Document 2] JP 2021-105770 A Summary of the Invention [Problem to be solved by the invention]

[0004] The voice agent service of Patent Document 1 judges whether the spoken content of a user is a search or not when the user speaks so that the method of response can be changed according to the user's intention, and if the spoken content is not a search, returns predetermined chat data as a response.The voice agent service (conversation assistant device) of Patent Document 2 judges whether the spoken content of a passenger can be automatically responded to when the passenger speaks so as to reduce the burden on the driver in relation to conversation with the passenger, and if so, returns pre-stored response data as a response. However, in these voice agent services (conversation assistant devices), response data is extracted from responses stored in advance, so natural responses like those returned in a conversation with a person are not returned.

[0005] In recent years, conversational AI such as ChatGPT has been proposed that can learn from the vast amount of information available on the Internet and have more natural conversations. This type of conversational AI returns an appropriate response when an appropriate prompt (text data) is input.

[0006] In other words, there is room for improvement in conventional voice agent services by utilizing conversational AI as described above.

[0007] The present invention has been made in consideration of the above-mentioned circumstances, and an object of the present invention is to provide a vehicle dialogue system that can appropriately respond to passenger voices and improve passenger convenience. [Means for solving the problem]

[0008] In order to achieve the above object, the vehicle dialogue system according to the present invention has the following features. A dialogue system for vehicles using a dialogue type AI, which is installed in a passenger vehicle and outputs response information consisting of text data when input information consisting of text data is input, a voice input unit for inputting voice spoken by a passenger; a voice recognition unit that converts the voice input by the voice input unit into text data; a control unit that inputs, as the input information, first text data, which is the text data converted by the voice recognition unit, and second text data, which indicates a command according to a vehicle state, to the dialogue AI; A voice synthesis unit that converts the response information from the dialogue AI into voice; a voice output unit that outputs the voice converted by the voice synthesis unit; The control unit controls the input information according to attributes of the passengers. It is a dialogue system for vehicles. Effect of the Invention

[0009] According to the present invention, the conversational AI can understand the attributes of passengers and the condition of the vehicle and respond appropriately to the passengers' voices, thereby improving passenger convenience.

[0010] The present invention has been briefly described above. Furthermore, the details of the present invention will be further clarified by reading the following description of the embodiment of the present invention (hereinafter, referred to as "embodiment") with reference to the accompanying drawings. [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 is a block diagram showing an example of the configuration of a vehicle dialogue system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram showing an example of the configuration of the agent shown in FIG. [Diagram 3] FIG. 3 is a diagram showing a first example of a processing procedure of the vehicular dialogue system shown in FIG. [Figure 4] FIG. 4 is a flowchart illustrating a first example of the procedure for generating a prompt shown in FIG. [Diagram 5] FIG. 5 corresponds to FIG. 4 and is a flow chart showing a first example of the processing procedure for generating the output information shown in FIG. [Figure 6] FIG. 6 is a flowchart illustrating a second example of the prompt generation processing procedure illustrated in FIG. [Figure 7] FIG. 7 corresponds to FIG. 6 and is a flow chart showing a second example of the processing procedure for generating output information shown in FIG. [Figure 8] FIG. 8 is a diagram showing a second example of the processing procedure of the vehicular dialogue system shown in FIG. [Figure 9] FIG. 9 is a flowchart illustrating an example of the prompt generation processing procedure illustrated in FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] Hereinafter, a vehicle dialogue system according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0013] The vehicle dialogue system 1 is a system that is mounted on a passenger vehicle and dialogues with passengers aboard the passenger vehicle by operating a passenger terminal 10 using a dialogue type AI (Artificial Intelligence) 30. The dialogue type AI 30 is composed of, for example, ChatGPT, and outputs response information also made up of text data when input information made up of text data is input.

[0014] The passenger terminal 10 includes a control unit 11 including an agent 20, a communication unit 12, a memory unit 13, a microphone 14 as an audio input unit, a speaker 15 as an audio output unit, an operation unit 16, and a display unit 17 (see FIG. 1).

[0015] The control unit 11 has a memory such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and a CPU (Central Processing Unit) that operates according to a program stored in the memory, and controls the vehicular dialogue system 1 as a whole.

[0016] The communication unit 12 is for communicating with the interactive AI 30 via an Internet communication network (not shown), and is composed of a circuit, an antenna, etc. for connecting to the Internet communication network. The communication unit 12 also communicates with a taximeter and a navigation system, which are necessary for the driver's operation, via wired or wireless communication.

[0017] The storage unit 13 is a means for storing data, and includes a recording medium such as a flash memory, and stores data used by the control unit for control purposes and part of the translation data.

[0018] The microphone 14 inputs the voices uttered by the passengers to the control unit 11. The speaker 15 outputs a voice converted by a voice synthesis unit 23 of the agent 20, which will be described later.

[0019] The operation unit 16 is a display screen with a touch panel function, buttons, etc., and receives various operations from passengers. The display unit 17 is a display or the like that visualizes various types of information, and is a medium that transmits various types of information to passengers.

[0020] As described above, the control unit 11 is capable of responding appropriately to passenger voices and includes an agent function (referred to as "agent 20" in this specification) for improving passenger convenience. The agent 20 has a voice recognition unit 21, a voice dialogue unit 22 as the "control unit" of the present invention, and a voice synthesis unit 23 (see FIG. 2). The agent 20 may also have a UI (User Interface) drawing unit. The UI drawing unit performs drawing processes such as voice input / output, destination display, fare calculation, and character display.

[0021] The voice recognition unit 21 converts the voice input by the microphone 14 into text data and inputs it to the voice dialogue unit 22. The voice dialogue unit 22 adds text data indicating the vehicle state when the passenger spoke or an instruction according to the vehicle state to the text data converted by the voice recognition unit 21, and inputs it to the dialogue AI 30 as input information. In addition, the voice dialogue unit 22 acquires response information from the dialogue AI 30, performs output processing on the text data of the acquired response information as necessary, and outputs it to the voice synthesis unit 23. The voice synthesis unit 23 converts the response information into voice and outputs it to the speaker 15.

[0022] Next, an example of the operation of the vehicle dialogue system 1 having the above-mentioned configuration will be described with reference to Fig. 3. When the agent 20 detects that the ignition of the passenger vehicle is on (IG-ON), the driver's operation, or a passenger boarding the vehicle, it acquires vehicle information indicating the vehicle state (Sp1). After that, when the passenger speaks (inputs) (Sp2), the terminal input information is transmitted to the agent 20 via the passenger terminal 10 (Sp3), and the agent 20, having acquired the terminal input information, performs a prompt generation process (Sp4).

[0023] Then, the agent 20 sends the prompt (text data) generated in Sp4 as input information to the interactive AI 30 (Sp5). After that, the agent 20 receives response information corresponding to the input information from the interactive AI 30 (Sp6) and performs output information creation processing (Sp7).

[0024] Then, the agent 20 transmits the output information created in Sp7 to the passenger terminal as terminal output information (Sp8), and voice is output by the passenger terminal 10 (Sp9). This allows a response to the passenger's voice.

[0025] Here, examples of the prompt generation process (Sp4) and the output information creation process (Sp7) will be described.

[0026] First, an example of the prompt generation process (Sp4) when a passenger is a foreign language speaker (i.e., a speaker of a language other than Japanese) will be described with reference to FIG. 4, and an example of the output information creation process (Sp7) in this case will be described with reference to FIG. 5.

[0027] When the prompt generation process is started, the agent 20 performs a voice recognition process to recognize the voice spoken by the passenger (Sp11). At this time, the agent 20 inquires about the language used by communicating with data stored in the storage unit 13 and an external database server via the communication unit 12. The agent 20 then performs a text conversion process to convert the voice recognized in Sp11 into text data (Sp12).

[0028] The agent 20 then determines whether the language of the text data converted into text in Sp12 is a predetermined language (in this example, English) (Sp13). If the text data is in English ("Yes" in Sp13), a command (text data) corresponding to the vehicle state is added to the text data generated in Sp12 (Sp14), and the prompt generation process ends. On the other hand, if the text data is not in English ("No" in Sp13), the text data is translated into English (Sp15), and a command (text data) corresponding to the vehicle state is added to the English-translated text data in Sp15 (Sp14), and the prompt generation process ends.

[0029] Thereafter, the agent 20 performs the output information creation process (Sp7) through Sp5 and Sp6. When the output information creation process starts, the agent 20 determines whether the text data at the time of conversion is in English or not (Sp21).

[0030] If the text data at the time of conversion is in English ("Yes" in Sp21), a voice synthesis process is performed to convert the response information into voice (Sp22), and the output information creation process ends. On the other hand, if the text data at the time of conversion is not in English ("No" in Sp21), the response information is translated into the original language (Sp23), and a voice synthesis process is performed to convert this response information into voice (Sp22), and the output information creation process ends. In this way, the input information is controlled by the agent 20 in accordance with the attributes of the passenger, so that an appropriate response is made to the passenger's voice.

[0031] The predetermined language is assumed to be the language of the country in which the conversational AI was developed, and is a language that can provide highly accurate response results. This also applies to the second predetermined language described below.

[0032] The text data given in Sp14 is in a predetermined language (English in this example). The vehicle state in Sp14 includes location information and the like. For example, if a passenger in a passenger vehicle in Kyoto City utters, "I want to go to the red tunnel," it is possible to generate text data such as, "Show me the red tunnel in Kyoto" by giving a command according to the location information (vehicle state) in Sp14, and it is considered that appropriate response information can be obtained from the conversational AI 30. The same applies to each of the examples described below.

[0033] Next, an example of the prompt generation process (Sp4) when the passenger speaks a first predetermined language (Japanese in this example) will be described with reference to Fig. 6, and an example of the output information creation process (Sp7) in this case will be described with reference to Fig. 7. The first predetermined language is the language of the country in which the passenger vehicle is traveling.

[0034] When the prompt generation process is started, the agent 20 performs a voice recognition process to recognize the voice spoken by the passenger (Sp31). The agent 20 then performs a text conversion process to convert the voice recognized in Sp31 into text data (Sp32), and adds a command (text data) corresponding to the vehicle state to this text data (Sp33), at which point the prompt generation process ends. A process is performed in parallel with this, and the text data generated in Sp33 is translated into a second predetermined language (English in this example) (Sp34), at which point the prompt generation process ends. As a result, prompts (text data) in Japanese and English are generated.

[0035] Thereafter, the agent 20 transmits Japanese and English text data as input information to the dialogue AI in Sp5, and after receiving Japanese and English response information in Sp6, output information creation processing is performed (Sp7).

[0036] When the output information creation process starts, the agent 20 compares Japanese and English response information (text data), selects the response information with high accuracy (Sp41), and determines whether the selected response information is in Japanese (Sp42).

[0037] If the response information is in Japanese ("Yes" in Sp42), a voice synthesis process is performed to convert the response information into voice (Sp43), and the output information creation process ends. On the other hand, if the response information is not in Japanese ("No" in Sp42), the response information is translated into Japanese (Sp44), and a voice synthesis process is performed to convert this response information into voice (Sp43), and the output information creation process ends. In this way, the input information is controlled by the agent 20 in accordance with the attributes of the passenger, so that an appropriate response is made to the passenger's voice.

[0038] Next, a further example of the operation of the vehicular dialogue system 1 will be described with reference to Fig. 8. When the agent 20 detects that the ignition of a passenger vehicle is turned on (IG-ON), the driver's operation, or a passenger boarding the vehicle, it acquires vehicle information indicating the vehicle state (Sp1a). After that, the agent 20 determines the passenger's attributes (Sp2a) and performs a prompt generation process (Sp3a).

[0039] The attributes of the passenger may be determined, for example, from the contents of the passenger's speech, or from at least one of the location information and the destination. Also, the driver may operate the driver terminal, and is not particularly limited.

[0040] The prompt (text data) generated in Sp3a is sent as input information to the dialogue AI 30 (Sp4a). After that, the agent 20 receives response information corresponding to the input information from the dialogue AI 30 (Sp5a), and performs output information creation processing such as voice synthesis processing (Sp6a).

[0041] Then, the agent 20 transmits the output information created in Sp6a to the passenger terminal as terminal output information (Sp7a), and voice is output by the passenger terminal 10 (Sp8a). This allows a response to the passenger's voice.

[0042] Here, an example of the prompt generation process (Sp3a) when the passenger's attribute is determined to be a tourist in Sp2a will be described with reference to FIG. 9. When the prompt generation process is started, the agent 20 generates text data corresponding to at least one of the location information and the destination (Sp51). At this time, the generated text data is, for example, tourist information corresponding to the location information and the destination. The agent 20 then gives an instruction (text data) corresponding to the vehicle state to the text data generated in Sp51 (Sp52), and the prompt generation process ends. That is, when the passenger's attribute is a tourist, the vehicular dialogue system 1 functions like a so-called tour guide. In this way, the agent 20 controls the input information according to the passenger's attribute, and an appropriate response is made to the passenger's voice.

[0043] As described above, according to this embodiment, the interactive AI 30 can understand the passenger's attributes and vehicle status and respond appropriately to the passenger's voice, thereby improving passenger convenience.

[0044] Furthermore, according to this embodiment, if the specified language (in this example, English) is the language of the country in which the interactive AI 30 was developed, or the like, and a language that can produce highly accurate response results, then if the passenger's voice is not in English, the text data can be translated into English and input into the interactive AI 30, making it possible to respond appropriately to the passenger's voice depending on the passenger's attributes.

[0045] Furthermore, according to this embodiment, the first specified language (in this example, Japanese) is the language of the country in which the passenger vehicle is traveling, and the second specified language (in this example, English) is the language of the country in which the conversational AI was developed, or the like, which can provide highly accurate response results. By inputting Japanese text data and English text data into the conversational AI 30 and selecting response information suitable for the passenger from among this response information, it is possible to respond appropriately to the passenger's voice depending on the passenger's attributes.

[0046] Furthermore, according to this embodiment, text data corresponding to at least one of the vehicle's position information and the destination is generated and input to the interactive AI 30, so that the vehicle dialogue system functions like a so-called tour guide. That is, according to this embodiment, it is possible to respond appropriately to the passenger's voice according to the passenger's attributes.

[0047] <Other forms> The present invention is not limited to the above-described embodiment, and can be appropriately modified, improved, etc. In addition, the material, shape, size, number, arrangement location, etc. of each component in the above-described embodiment are arbitrary as long as the present invention can be achieved, and are not limited.

[0048] In the above embodiment, voice input is assumed, but input may also be made by operating the operation unit 16, for example. In this case, the agent 20 only needs to convert the terminal input information into text. Note that if the terminal input information is text data, text conversion is not necessary. This process may be performed by the voice dialogue unit 22, and is not particularly limited.

[0049] In the above embodiment, it is assumed that voice is output, but for example, an image or the like may be displayed on the display unit 17. In this case, the agent 20 may display the response information directly on the display unit 17, or convert it into image data. This process may be performed by the voice dialogue unit 22, and is not particularly limited.

[0050] In the above embodiment, the passenger's speech can be translated into English by performing the steps Sp11, Sp12, Sp13, and Sp15 in the prompt generation process (Sp4) and the steps Sp21 and Sp22 in the output information creation process (Sp7). That is, the passenger's speech can be translated into a desired language by changing the part corresponding to "English" in the above to the desired language.

[0051] The above-described embodiments and other embodiments can be combined as appropriate.

[0052] Here, the features of the above-described embodiment of the vehicular dialogue system according to the present invention will be briefly summarized and listed in the following [1] to [4].

[0053] [1] A vehicle dialogue system (1) using a dialogue type AI (30) that is installed in a passenger vehicle and that outputs response information made of text data when input information made of text data is input, A voice input unit (microphone 14) for inputting voices uttered by passengers; a voice recognition unit (21) that converts the voice inputted through the voice input unit (microphone 14) into text data; a control unit (voice dialogue unit 22) that inputs, as the input information, first text data, which is the text data converted by the voice recognition unit (21), and second text data, which indicates a command according to a vehicle state, to the dialogue AI; A voice synthesis unit (23) that converts the response information from the dialogue AI (30) into voice; a voice output unit (speaker 15) that outputs the voice converted by the voice synthesis unit (23); The control unit (voice dialogue unit 22) controls the input information according to the attributes of the passenger. Dialogue system for vehicles.

[0054] According to the configuration [1] above, the conversational AI can understand the passengers' attributes and the vehicle status and respond appropriately to the passengers' voices, thereby improving passenger convenience.

[0055] [2] The vehicle dialogue system (1) according to the above [1], When the first text data is not in a predetermined language, the control unit (voice dialogue unit 22) inputs third text data obtained by translating the first text data into a predetermined language as the input information to the dialogue AI (30). Dialogue system for vehicles.

[0056] According to the configuration [2] above, if the specified language is a language of the country in which the conversational AI was developed, or the like, that can provide highly accurate response results, then if the passenger's voice (first text data) is not in the specified language, the text data can be translated into the specified language and input to the conversational AI, making it possible to respond appropriately to the passenger's voice depending on the passenger's attributes.

[0057] [3] The vehicle dialogue system (1) according to the above [1] or [2], When the first text data is in a first predetermined language, the control unit (voice dialogue unit 22) inputs the first text data and third text data obtained by translating the first text data into a second predetermined language as the input information to the dialogue AI (30), and selects response information suitable for the passenger from the response information in the first predetermined language and the response information in the second predetermined language. Dialogue system for vehicles.

[0058] According to the configuration [3] above, the first specified language is the language of the country in which the passenger vehicle travels, and the second specified language is a language that can obtain highly accurate response results, such as the language of the country in which the conversational AI was developed. By inputting text data in the first specified language (first text data) and text data in the second specified language (third text data) to the conversational AI and selecting response information suitable for the passenger from among these pieces of response information, it is possible to respond appropriately to the passenger's voice according to the passenger's attributes.

[0059] [4] A vehicle dialogue system (1) according to any one of the items [1] to [3], The control unit (voice dialogue unit 22) generates fourth text data according to at least one of the vehicle position information and the destination, and inputs the fourth text data as the input information to the dialogue AI. Dialogue system for vehicles.

[0060] According to the configuration of [4] above, text data (fourth text data) corresponding to at least one of the vehicle's position information and the destination is generated, and this is input to the dialogue AI, so that the dialogue system for the vehicle functions like a so-called tour guide. That is, according to the above configuration, it is possible to respond appropriately to the voice of the passenger according to the passenger's attributes. [Explanation of symbols]

[0061] 1 Vehicle dialogue system 10 Passenger terminals 11 Control section 12 Communications Department 13 Storage section 14. Mike 15 Speakers 20. Agent 21 Voice Recognition Unit 22 Voice dialogue section 23 Voice synthesis section 30 Conversational AI

Claims

1. A vehicle dialogue system using a dialogue type AI, which is mounted on a passenger vehicle and outputs response information consisting of text data when input information consisting of text data is input, a voice input unit for inputting voice spoken by a passenger; a voice recognition unit that converts the voice input by the voice input unit into text data; a control unit that inputs, as the input information, first text data, which is the text data converted by the voice recognition unit, and second text data, which indicates a command according to a vehicle state, to the interactive AI; A voice synthesis unit that converts the response information from the conversational AI into voice; a voice output unit that outputs the voice converted by the voice synthesis unit; The control unit controls the input information according to attributes of the passengers. Dialogue system for vehicles.

2. 2. A vehicle dialogue system according to claim 1, When the first text data is not in a predetermined language, the control unit inputs third text data obtained by translating the first text data into a predetermined language as the input information to the interactive AI. Dialogue system for vehicles.

3. 2. A vehicle dialogue system according to claim 1, When the first text data is in a first predetermined language, the control unit inputs the first text data and third text data obtained by translating the first text data into a second predetermined language as the input information to the interactive AI, and selects response information suitable for the passenger from the response information in the first predetermined language and the response information in the second predetermined language. Dialogue system for vehicles.

4. 2. A vehicle dialogue system according to claim 1, The control unit generates fourth text data corresponding to at least one of the vehicle position information and the destination, and inputs the fourth text data as the input information to the interactive AI. Dialogue system for vehicles.

Citation Information

Patent Citations

  • Interaction support device, interaction system, interaction support method, and program

    JP2014098844A

  • Conversation assistant apparatus and method

    JP2021105770A