Dialogue system for vehicle

Through the voice input, recognition, control and synthesis unit of the dialogue AI system, combined with the control of the vehicle equipment, the problem that voice agent services in the prior art cannot operate the vehicle equipment and make natural responses is solved, and the voice operation and natural interaction of the vehicle dialogue system is realized.

CN120476444APending Publication Date: 2025-08-12YAZAKI CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202480006341.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2024-10-18
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing voice proxy service cannot operate the vehicle equipment through voice, and cannot provide a natural response to the words that do not intend to operate the vehicle equipment.

Method used

The dialogue AI system is adopted, including a voice input unit, a voice recognition unit, a device control unit, a voice synthesis unit and a voice output unit. The driver's voice is converted into text data through the voice recognition unit. The device control unit controls the vehicle equipment, and the voice synthesis unit generates a voice response, and uses dialogue AI to generate a natural response when determining non-operating instructions.

Benefits of technology

The function of operating the on-board equipment through voice is realized, and a natural response to non-operating instructions such as chatting is made, improving the voice interaction capabilities of the on-board equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476444A_ABST
    Figure CN120476444A_ABST
Patent Text Reader

Abstract

A voice interaction unit (42) determines whether or not the text data converted by the voice recognition unit (41) indicates an operation command of the in-vehicle device (11). When the voice interaction unit (42) determines that the text data represents an operation instruction of the vehicle-mounted device (11), the voice interaction unit (42) controls the vehicle-mounted device (11) according to the operation instruction. When the voice interaction unit (42) determines that the text data does not represent the operation instruction of the vehicle-mounted device (11), the voice interaction unit (42) inputs the text data of the voice as input information (S1) to the interactive AI (10).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a dialogue system for a vehicle. Background Art

[0002] A voice agent service has been proposed that responds to user utterances using voice (Patent Document 1). The voice agent service in Patent Document 1 is a system that responds and returns replies only to predefined utterance data. Therefore, when utterances other than those defined are uttered, the voice agent service in Patent Document 1 either responds with "I don't understand" or associates the utterance with the closest utterance among the defined utterances. As a result, the service does not return natural responses, as would be expected from a conversation with a human.

[0003] In recent years, conversational AI (generative AI), such as ChatGPT, has been proposed, which enables more natural conversations by learning from the vast amount of information available on the internet. However, while this type of conversational AI is suitable for casual conversations such as small talk, it lacks the ability to operate in-vehicle devices. Consequently, there is a problem: when the driver utters a word to operate a device, this type of conversational AI is unable to handle the situation.

[0004] Reference List

[0005] Patent Literature

[0006] Patent Document 1: JP2014-98844A Summary of the Invention

[0007] Technical issues

[0008] The present disclosure has been made in view of the above circumstances, and an object thereof is to provide a dialogue system for a vehicle that can operate an in-vehicle device by voice and also can make a natural response to words not intended to operate the in-vehicle device, such as small talk.

[0009] Solutions to the Problem

[0010] In order to achieve the above-mentioned object, a dialogue system for a vehicle according to the present disclosure has the following features.

[0011] A vehicle dialogue system utilizing conversational AI, wherein the conversational AI outputs response information consisting of text data in response to receiving input information consisting of text data, the vehicle dialogue system comprising:

[0012] a voice input unit configured to input a voice uttered by a driver;

[0013] a speech recognition unit configured to convert the speech input through the speech input unit into text data;

[0014] a device control unit configured to, in response to determining that the text data converted by the speech recognition unit represents an operation instruction for an in-vehicle device, control the in-vehicle device according to the operation instruction;

[0015] a speech synthesis unit configured to, in response to determining that the text data does not represent the operation instruction for the in-vehicle device, convert the response information output from the conversational AI into speech based on input information consisting of the text data converted by the speech recognition unit; and

[0016] A speech output unit is configured to output the speech converted by the speech synthesis unit.

[0017] Effects of the Invention

[0018] According to the vehicle dialogue system of the present disclosure, it is possible to achieve the following effects: an in-vehicle device can be operated by voice, and a natural response can also be made to utterances not intended to operate the in-vehicle device, such as small talk.

[0019] The present disclosure has been briefly described above. Further, details of the present disclosure will be clarified by reading the following modes for carrying out the present disclosure (hereinafter referred to as “embodiments”) with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a block diagram illustrating a conversation system for a vehicle according to an embodiment of the present disclosure.

[0021] Figure 2 Shows the storage Figure 1 An example of a data table in the ROM is shown.

[0022] Figure 3 Is shown equipped with Figure 1 A view of the periphery of the dashboard of a vehicle is shown, showing a vehicular dialogue system.

[0023] Figure 4 The first embodiment shows the structure of Figure 1 Flowchart of the processing procedure of the microcomputer of the vehicle communication system shown.

[0024] Figure 5 It shows Figure 4 An illustration of the operation in Sp5 in FIG.

[0025] Figure 6 The second embodiment shows the structure Figure 1Flowchart of the processing procedure of the microcomputer of the vehicle communication system shown.

[0026] Figure 7 is a block diagram of a vehicle dialogue system according to another embodiment.

[0027] Figure 8 is a block diagram of a vehicle dialogue system according to another embodiment.

[0028] Figure 9 is a block diagram of a vehicle dialogue system according to another embodiment.

[0029] Figure 10 Shows the storage Figure 1 Another example of a data table in a ROM is shown.

[0030] Reference Mark List

[0031] 1, 1B, 1C, 1D vehicle dialogue systems

[0032] 2 microphones (voice input unit)

[0033] 5 speakers (voice output unit)

[0034] 7ROM (storage unit)

[0035] 10. Conversational AI

[0036] 11In-vehicle equipment

[0037] 12 servers (speech recognition unit, judgment unit)

[0038] 13DB (storage unit)

[0039] 41 speech recognition units

[0040] 42 Voice dialogue unit (determination unit, device control unit, first input control unit, second input control unit)

[0041] 44 speech synthesis units DETAILED DESCRIPTION

[0042] (First embodiment)

[0043] Hereinafter, a first embodiment of the present disclosure will be described with reference to the accompanying drawings.

[0044] A vehicular dialogue system 1 according to the first embodiment is a system installed in a vehicle and interacts with a driver using a conversational artificial intelligence (AI) 10. Conversational AI 10 is implemented by, for example, ChatGPT, and outputs response information S2 consisting of text data when input information S1 consisting of text data is input.

[0045] The vehicle dialogue system 1 includes a microphone 2 as a voice input unit, a communication module 3, a microcomputer 4, a speaker 5 as a voice output unit, a display 6, and a read-only memory (ROM) 7 (storage unit). Microphone 2 inputs the driver's speech into microcomputer 4. Communication module 3 is used to communicate with conversational AI 10 via an internet communication network (not shown) and includes circuits, an antenna, and other components for connecting to the internet communication network. In this embodiment, communication module 3, microcomputer 4, and ROM 7 (described later) are mounted on the same control board 100.

[0046] The microcomputer 4 includes, for example, a memory such as a random access memory (RAM) or a ROM, and a central processing unit (CPU), operates according to a program stored in the memory, and controls the entire vehicular dialogue system 1 .

[0047] Microcomputer 4 includes a speech recognition unit 41, a speech dialogue unit 42, a speech synthesis unit 44, and a graphics processing unit 45. Speech recognition unit 41 converts speech input from microphone 2 into text data and inputs the text data to speech dialogue unit 42. Speech dialogue unit 42 inputs the text data converted by speech recognition unit 41 into conversational AI 10 as input information S1.

[0048] The voice dialogue unit 42 receives vehicle information S3 and personal information S4 as input. The microcomputer 4 is connected to sensors or devices installed in the vehicle via a communication network provided in the vehicle, such as a controller area network (CAN). The vehicle information S3 is information indicating the vehicle status obtained from the sensors or devices installed in the vehicle.

[0049] like Figure 2 As shown, the ROM 7 stores a data table including an operation instruction for the vehicle-mounted device 11, instruction text data corresponding to the operation instruction for the vehicle-mounted device 11, and response text data corresponding to the instruction text data. Figure 2 In the example shown, one instruction text data item is stored for one operation instruction. However, the present disclosure is not limited thereto. Multiple instruction text data items may be stored for one operation instruction. For example, in addition to "turn on the air conditioner," instruction text data items such as "heat" and "cold" may also be associated with the "turn on the air conditioner" operation instruction.

[0050] Personal information S4 is input to the voice dialogue unit 42, which is a detection result from a driver monitor that detects the driver's state (whether the driver is dozing off, driving carelessly, or driving inattentively) based on an image obtained by capturing the driver's face.

[0051] The voice dialogue unit 42 is connected to the in-vehicle device 11 mounted on the vehicle and can control the in-vehicle device 11. Examples of the in-vehicle device 11 include an air conditioner mounted on the vehicle, a motor for opening and closing windows, a headlight, and an electronic control unit (ECU) for controlling an adaptive cruise control (ACC) function.

[0052] The speech dialogue unit 42 serves as a determination unit that compares the text data converted by the speech recognition unit 41 with the text data. Figure 2 The voice dialogue unit 42 compares the text data converted by the voice recognition unit 41 with the instruction text data shown, and when a match occurs, determines that the text data converted by the voice recognition unit 41 represents an operation instruction for the in-vehicle device 11. If a match does not occur, the voice dialogue unit 42 determines that the text data converted by the voice recognition unit 41 does not represent an operation instruction for the in-vehicle device 11.

[0053] If the voice dialogue unit 42 determines that the text data represents an operation instruction for the in-vehicle device 11, the voice dialogue unit 42 functions as a device control unit and controls the in-vehicle device 11 according to the operation instruction corresponding to the compared instruction text data. If the voice dialogue unit 42 determines that the text data represents an operation instruction for the in-vehicle device 11, the voice dialogue unit 42 inputs the response text data corresponding to the matched instruction text data to the voice synthesis unit 44.

[0054] When the voice dialogue unit 42 determines that the text data does not represent an operation instruction for the in-vehicle device 11, the voice dialogue unit 42 transmits a prompt including the text data converted by the voice recognition unit 41 as input information S1 to the conversational AI 10. The voice dialogue unit 42 inputs the response information S2 from the conversational AI 10 and outputs the input response information S2 to the voice synthesis unit 44. The voice synthesis unit 44 converts the response text data or the response information S2 into voice and outputs the voice to the speaker 5. The speaker 5 outputs the voice converted by the voice synthesis unit 44.

[0055] While outputting the voice from the speaker 5, the voice dialogue unit 42 outputs a display request for displaying the character on the display 6 to the drawing processing unit 45. Figure 3 As shown, the display 6 is provided on the instrument panel between the driver's seat and the passenger seat. The graphics processing unit 45 outputs an image to the display 6 in which the character seems to be speaking in synchronization with the voice output from the speaker 5.

[0056] Next, refer to Figure 4 The flowchart shown describes the operation of the vehicle dialogue system 1 having the above configuration. If the microcomputer 4 detects that the vehicle dialogue system 1 is turned on, such as when the ignition is turned on, the microcomputer 4 starts Figure 4First, the microcomputer 4 enters a standby state until the driver starts to speak (Sp1). If the driver speaks (Yes in Sp2), the microcomputer 4 performs a speech recognition process to convert the driver's speech into text data (Sp3).

[0057] If the utterance has not yet ended (No in Sp4), the process returns to Sp3 and the microcomputer 4 continues the voice recognition process. On the other hand, if the utterance has ended (Yes in Sp4), the microcomputer 4 performs a determination process to determine whether the text data of the voice converted by the voice recognition process represents an operation instruction for the in-vehicle device 11 (Sp5).

[0058] In the judgment process, the microcomputer 4 compares the text data of the speech with the Figure 2 For example, when the text data of the voice is "turn on ACC", Figure 5 As shown, the microcomputer 4 sequentially compares the text data with the instruction text data "turn on the air conditioner", "open the window", "turn on the headlights", and "turn on the ACC".

[0059] If there is instruction text data that matches the text data of the speech, the microcomputer 4 determines that the text data of the speech represents an operation instruction for the in-vehicle device 11. The microcomputer 4 is not limited to determining a match based on a complete match of the text data, but may determine a match when the matching rate of the words is equal to or greater than a certain value.

[0060] Next, if the microcomputer 4 determines through the determination process that the voice text data represents an operation instruction for the in-vehicle device 11 (Yes in Sp6), the microcomputer 4 controls the in-vehicle device 11 according to the operation instruction corresponding to the matched instruction text data (Sp7). For example, if the voice text data matches the instruction text data "turn on the ACC function" in the determination process, the microcomputer 4 sends a request to turn on the ACC function to the ECU that controls the ACC function according to the operation instruction.

[0061] Next, microcomputer 4 retrieves the response text data corresponding to the matched command text data from the data table (Sp8). Microcomputer 4 then performs speech synthesis processing to convert the retrieved response text data into speech and output the speech from speaker 5. After reading the response text data (Sp9), processing proceeds to Sp10. For example, in the determination process, if the speech text data matches the command text data "Activate ACC function," microcomputer 4 retrieves the response text data "ACC is activated. Speed is set to xx km / h. Following distance is set to close" (Sp8) and reads the response text data from speaker 5.

[0062] On the other hand, if the microcomputer 4 determines through the determination process that the text data of the voice does not represent an operation instruction for the vehicle-mounted device 11 (No in Sp6), the microcomputer 4 creates a prompt including the text data of the voice (Sp11). In Sp11, the microcomputer 4 may create only the text data of the voice as a prompt, or may create a prompt in which text data corresponding to the vehicle information S3 or the personal information S4 is added to the text data of the voice.

[0063] Next, microcomputer 4 functions as a first input control unit and transmits the created prompt as input information S1 to conversational AI 10 (Sp12). If microcomputer 4 receives response information S2 from conversational AI 10 (Yes in Sp13), microcomputer 4 converts the received response information S2 into speech and outputs the speech from speaker 5. After reading response information S2 (Sp14), the process proceeds to Sp10.

[0064] In Sp10, if the microcomputer 4 detects that the vehicle dialogue system 1 is turned off, such as when the ignition is turned off (Yes in Sp10), the process ends. If the microcomputer 4 does not detect that the vehicle dialogue system 1 is turned off (No in Sp10), the process returns to Sp1 and the microcomputer 4 enters a standby state until the driver speaks again.

[0065] According to the above embodiment, microcomputer 4 determines whether the text data of the speech indicates an operation instruction for in-vehicle device 11. If microcomputer 4 determines that the text data of the speech indicates an operation instruction for in-vehicle device 11, microcomputer 4 controls in-vehicle device 11 according to the operation instruction. Furthermore, if microcomputer 4 determines that the text data of the speech does not indicate an operation instruction for in-vehicle device 11, microcomputer 4 transmits a prompt including the text data of the speech to conversational AI 10. Thus, vehicular dialogue system 1 can operate in-vehicle device 11 by speech and can respond naturally to utterances not intended to operate in-vehicle device 11, such as small talk.

[0066] According to the above embodiment, the instruction text data is stored in the ROM 7. The microcomputer 4 compares the instruction text data with the text data of the speech and determines whether the text data of the speech represents an operation instruction for the in-vehicle device 11 based on whether there is matching instruction text data. Therefore, the microcomputer 4 can easily determine whether the text data of the speech represents an operation instruction for the in-vehicle device 11.

[0067] According to the above embodiment, the response text data is stored in the ROM 7. If the microcomputer 4 determines that the text data of the speech represents an operation instruction for the in-vehicle device 11, the microcomputer 4 converts the response text data corresponding to the matching instruction text data into speech and reads the speech. Therefore, when the driver issues an operation instruction for the in-vehicle device 11, an appropriate response can be made.

[0068] (Second embodiment)

[0069] Next, a second embodiment will be described.

[0070] Since the dialog system 1 for a vehicle according to the second embodiment has Figure 1 The configuration of the vehicle dialogue system 1 according to the first embodiment is the same as shown, so its detailed description will be omitted here. In the first embodiment, the voice dialogue unit 42 is used as the determination unit. However, in the second embodiment, the conversational AI 10 is used as the determination unit.

[0071] Next, refer to Figure 6 The flowchart shown describes the operation of the dialog system 1 for a vehicle according to the second embodiment. If the microcomputer 4 detects that the dialog system 1 for a vehicle is turned on, such as when the ignition is turned on, the microcomputer 4 starts Figure 6 First, the microcomputer 4 acquires information about in-vehicle devices installed in the vehicle (in-vehicle device information) from the vehicle information S3 acquired via the controller area network (CAN) or the like (Sp21). Next, the microcomputer 4 generates a prompt for notifying the conversational AI 10 of the acquired in-vehicle device information and transmits the prompt (Sp22).

[0072] An example of the prompt sent in step Sp22 will be described below. Microcomputer 4 converts the acquired in-vehicle device information into text data. Thereafter, microcomputer 4 generates a prompt that incorporates the text data of the in-vehicle device information, converting between, for example, the first template sentence pattern, "You are currently in the vehicle," and the second template sentence pattern, "The vehicle is equipped with features. Please answer the following questions based on these features." Thus, the text data, "You are currently in the vehicle. The vehicle is equipped with features such as ACC, wipers, headlights, air conditioning / heater... Please answer the following questions based on these features," is sent to conversational AI 10.

[0073] Thereafter, the microcomputer 4 enters a standby state until the driver starts to speak (Sp23). If the driver speaks (Yes in Sp24), the microcomputer 4 performs a speech recognition process to convert the driver's speech into text data (Sp25).

[0074] If the utterance has not yet ended (No in Sp26), the process returns to Sp25, and the microcomputer 4 continues the voice recognition process. On the other hand, if the utterance has ended (Yes in Sp26), the microcomputer 4 functions as a second input control unit, generates a prompt including text data of the voice converted by the voice recognition process, a determination command as to whether the text data represents an operation instruction for the in-vehicle device 11, and a transmission command for response information corresponding to the text data when the text data does not represent an operation instruction (Sp27), and transmits the generated prompt (Sp28).

[0075] An example of the prompt sent in Sp28 will be described below. Microcomputer 4 generates and sends a prompt in which, following the text data of the voice, text data of a template sentence pattern is added: "Is this content for operation? If so, please answer with 'A'. Otherwise, please answer in a manner that allows the conversation to continue, but do not answer with 'No'."

[0076] In response to receiving response information S2 from conversational AI 10 (Yes in Sp29), microcomputer 4 determines whether the text data of the spoken voice represents an operational instruction based on response information S2 (Sp30). For example, in Sp28, if microcomputer 4 sends a prompt indicating "Is the content 'Turn on the wipers' an operational instruction? If so, please answer 'A'. Otherwise, please answer in a manner that allows the conversation to continue, but do not answer 'No'," and in response, receives response information S2 indicating "A," microcomputer 4 determines that the text data of the spoken voice represents an operational instruction (Yes in Sp30).

[0077] If the microcomputer 4 determines that the text data of the speech is an operation instruction (Yes in Sp30), the microcomputer 4 compares the text data converted by the speech recognition unit 41 with the operation instruction. Figure 2 The microcomputer 4 compares the command text data shown in FIG. 1 with the command text data, and controls the vehicle-mounted device 11 according to the operation command corresponding to the matched command text data (Sp31). Figure 2 The microcomputer 4 performs speech synthesis processing to convert the acquired response text data into speech and output the speech from the speaker 5, and after reading out the response text data (Sp33), the processing proceeds to Sp35.

[0078] For example, in Sp28, when the microcomputer 4 sends a prompt indicating "'Hello.' Is this content for operation? If it is for operation, please answer with 'A'. Otherwise, please answer in a way that allows the conversation to continue, but do not answer with 'No'", and in response, receives a response message S2 indicating "Hello. I am happy to talk to you. Do you have anything to ask or how I can help? ", the microcomputer 4 determines that the text data of the voice is not an operation instruction (No in Sp30).

[0079] If the microcomputer 4 determines that the text data of the voice is not an operation instruction, the microcomputer 4 converts the received response information S2 into voice and outputs the voice from the speaker 5, and after reading out the response information S2 (Sp34), the process proceeds to Sp35.

[0080] In Sp35, if the microcomputer 4 detects that the vehicle dialogue system 1 is turned off, such as when the ignition is turned off (Yes in Sp35), the process ends. If the microcomputer 4 does not detect that the vehicle dialogue system 1 is turned off (No in Sp35), the process returns to Sp23, and the microcomputer 4 enters a standby state until the driver speaks again.

[0081] According to the above embodiment, the conversational AI 10 is used as a determination unit. In the case of the first embodiment, whether the text data is an operation instruction or not, it is necessary to compare the text data uttered by the driver with the text data uttered by the driver. Figure 2 On the other hand, according to the second embodiment, when the text data uttered by the driver is not intended for operation, it is not necessary to compare the text data with the data table shown in FIG. Figure 2 The command text data in the data table shown is compared, thereby improving the processing speed and the response speed of the answer to the driver's words.

[0082] The present disclosure is not limited to the above-described embodiments and can be appropriately modified, improved, etc. In addition, the materials, shapes, sizes, numbers, arrangement positions, etc. of the components in the above-described embodiments are freely selectable and are not limited as long as the present disclosure can be achieved.

[0083] According to the above embodiment, according to Figure 1 In the vehicle dialogue system 1 shown, the conversational AI 10 communicates with the vehicle dialogue system 1 via the Internet communication network. However, the present disclosure is not limited thereto. Figure 7 The vehicle dialogue system 1B shown includes a conversational AI 10B including a microcomputer, which can be mounted on a control board 100. Figure 7 The illustrated vehicle dialogue system 1B enables dialogue even in a poor communication environment.

[0084] according to Figure 1 and Figure 7 In the vehicle dialogue system 1, 1B shown, the microcomputer 4 is used as a voice recognition unit. However, the present disclosure is not limited thereto. Figure 8 and Figure 9 In the vehicle dialogue system 1C and 1D shown in FIG. 1 , the microcomputer 4 can communicate with the server 12 serving as a speech recognition unit and a decision unit. The server 12 can access the stored Figure 2 A database (DB) 13 of data tables is shown.

[0085] In this case, the microcomputer 4 transmits the voice input from the microphone 2 to the server 12. The server 12 converts the received voice into text data, refers to the data table stored in the DB 13, and determines whether the text data represents an operation instruction for the in-vehicle device 11. The server 12 transmits the determination result and the text data of the voice to the microcomputer 4. Therefore, it is not necessary to provide the microcomputer 4 with the determination functions of the voice recognition unit 41 and the voice dialogue unit 42, and the processing load can be reduced.

[0086] according to Figure 8 and Figure 9 In the illustrated vehicular dialogue systems 1C and 1D, the microcomputer 4 does not include the speech recognition unit 41. However, the microcomputer 4 may include the speech recognition unit 41. In this case, the microcomputer 4 converts part of the speech into text data via the speech recognition unit 41, and the server 12 converts the remaining speech into text data. In this case as well, the processing load on the microcomputer 4 can be reduced compared to a case where all speech recognition is performed by the microcomputer 4.

[0087] According to the first and second embodiments described above, the microcomputer 4 compares the text data of the speech with the Figure 2 The plurality of instruction text data shown are compared one by one, and it is determined whether the text data of the voice is intended for operation, and if it is intended for operation, the microcomputer 4 determines the instruction content thereof. However, the present disclosure is not limited thereto.

[0088] like Figure 10 As shown, the ROM 7 stores a data table including operation instructions for the in-vehicle device 11, instruction keywords corresponding to the operation instructions for the in-vehicle device 11, and response text data corresponding to the instruction keywords. The microcomputer 4 extracts keywords from the text data of the speech, compares the extracted keywords with the instruction keywords, and if all keywords match, the microcomputer 4 determines that it is intended for operation and performs the operation instruction corresponding to the matched instruction keyword.

[0089] Here, the features of the embodiment of the dialogue system for a vehicle according to the present disclosure described above are briefly summarized and listed in the following [1] to [5].

[0090] [1] A vehicle dialogue system (1, 1B, 1C, 1D) using a conversational AI (10), wherein the conversational AI (10) outputs response information (S2) consisting of text data in response to receiving input information (S1) consisting of text data, the vehicle dialogue system (1, 1B, 1C, 1D) comprising:

[0091] A voice input unit (2) is configured to input a voice uttered by a driver;

[0092] a speech recognition unit (12) configured to convert speech inputted through the speech input unit (2) into text data;

[0093] a device control unit (42) configured to, in response to determining that the text data converted by the voice recognition unit (12) represents an operation instruction for the in-vehicle device (11), control the in-vehicle device (11) according to the operation instruction;

[0094] a speech synthesis unit (44) configured to convert response information (S2) output from the conversational AI (10) into speech based on input information (S1) consisting of text data converted by the speech recognition unit (12) in response to determining that the text data does not represent an operation instruction for the in-vehicle device (11); and

[0095] The speech output unit (5) is configured to output the speech converted by the speech synthesis unit (44).

[0096] According to the vehicle dialogue system (1, 1B, 1C, 1D) having the configuration described in [1] above, the vehicle-mounted device (11) can be operated by voice, and furthermore, a natural response can be made to words such as small talk that are not intended to operate the vehicle-mounted device (11).

[0097] [2] The vehicle dialogue system (1, 1B, 1C, 1D) according to [1], further comprising:

[0098] a determination unit (42, 12) configured to determine whether the text data converted by the voice recognition unit (12) represents an operation instruction for the in-vehicle device (11); and

[0099] A first input control unit (42) is configured to input the text data converted by the speech recognition unit (12) as input information (S1) to the conversational AI (10) in response to the determination unit (42, 12) determining that the text data does not represent an operation instruction for the vehicle-mounted device (11).

[0100] According to the vehicle dialogue system (1, 1B, 1C, 1D) having the configuration described in [2] above, text data that does not represent an operation instruction can be input to the dialogue AI (10) by the determination unit (42, 12).

[0101] [3] The vehicle dialogue system (1, 1B, 1C, 1D) according to [1], further comprising:

[0102] A second input control unit (42) is configured to input text data converted by the voice recognition unit (12), a determination command as to whether the text data represents an operation instruction for the vehicle-mounted device (11), and a transmission command for response information (S2) corresponding to the text data when the text data does not represent an operation instruction, as input information (S1) to the conversational AI (10), wherein:

[0103] The device control unit (42) controls the vehicle-mounted device (11) according to the operation instruction in response to the response information (S2) input from the conversational AI (10) indicating the operation instruction for the vehicle-mounted device (11), and

[0104] A speech synthesis unit (44) converts response information (S2) corresponding to text data into speech in response to response information (S2) input from a conversational AI (10) that does not represent an operation instruction for an in-vehicle device (11).

[0105] According to the configuration in [3] above, the conversational AI (10) is able to determine whether text data represents an operation instruction, thereby improving processing capabilities.

[0106] [4] The vehicle dialogue system (1, 1B, 1C, 1D) according to [2], further comprising:

[0107] The storage unit (7, 13) is configured to store instruction text data corresponding to an operation instruction for the vehicle-mounted device (11), wherein:

[0108] The determination unit (42, 12) compares the instruction text data stored in the storage unit with the text data converted by the speech recognition unit (12), and makes a determination based on whether there is matching instruction text data.

[0109] According to the vehicle dialogue system (1, 1B, 1C, 1D) having the configuration in [4] above, the determination unit (42, 12) can easily determine whether the text data of the speech represents an operation instruction for the vehicle-mounted device (11).

[0110] [5] The vehicle dialogue system (1, 1B, 1C, 1D) according to [4], wherein

[0111] The storage unit (7, 13) also stores the response text data corresponding to the instruction text data.

[0112] When the determination unit (42, 12) determines that the text data represents an operation instruction for the vehicle-mounted device (11), the determination unit (42, 12) causes the speech synthesis unit (44) to convert the response text data corresponding to the matched instruction text data into speech.

[0113] According to the vehicle dialogue system (1, 1B, 1C, 1D) having the configuration described in [5] above, when the driver issues an operation instruction to the vehicle-mounted device (11), an appropriate response can be made.

[0114] While the present disclosure has been described in detail with reference to specific embodiments, it will be apparent to one skilled in the art that various changes and modifications can be made without departing from the spirit and scope of the present disclosure.

[0115] This application is based on the Japanese patent application filed on October 23, 2023 (Japanese patent application No. 2023-181856) and the Japanese patent application filed on February 19, 2024 (Japanese patent application No. 2024-023172), and the contents of which are incorporated herein by reference.

[0116] Industrial Applicability

[0117] According to the present disclosure, it is possible to provide a vehicle dialogue system that can operate an in-vehicle device by voice and can also respond naturally to utterances not intended to operate the in-vehicle device, such as small talk. The present disclosure having such an effect is useful for a vehicle dialogue system.

Claims

1. A vehicle dialogue system utilizing conversational AI, wherein the conversational AI outputs response information consisting of text data in response to receiving input information consisting of text data, the vehicle dialogue system comprising: a voice input unit configured to input a voice uttered by a driver; a speech recognition unit configured to convert the speech input through the speech input unit into text data; a device control unit configured to, in response to determining that the text data converted by the speech recognition unit represents an operation instruction for an in-vehicle device, control the in-vehicle device according to the operation instruction; a speech synthesis unit configured to, in response to determining that the text data does not represent an operation instruction for the in-vehicle device, convert the response information output from the conversational AI into speech based on the input information consisting of the text data converted by the speech recognition unit; and A speech output unit is configured to output the speech converted by the speech synthesis unit.

2. The vehicle dialogue system according to claim 1, further comprising: a determination unit configured to determine whether the text data converted by the speech recognition unit represents an operation instruction for the in-vehicle device; as well as A first input control unit is configured to input the text data converted by the voice recognition unit as the input information to the conversational AI in response to the determination unit determining that the text data does not represent an operation instruction for the in-vehicle device.

3. The vehicle dialogue system according to claim 1, further comprising: a second input control unit configured to input, as the input information, the text data converted by the voice recognition unit, a determination command as to whether the text data represents an operation instruction for the in-vehicle device, and a transmission command for the response information corresponding to the text data when the text data does not represent the operation instruction, into the conversational AI, wherein: The device control unit controls the in-vehicle device according to the operation instruction in response to the response information indicating the operation instruction for the in-vehicle device input from the conversational AI, and The speech synthesis unit converts the response information corresponding to the text data into speech in response to the response information input from the conversational AI that does not represent an operation instruction for the in-vehicle device.

4. The vehicle dialogue system according to claim 2, further comprising: A storage unit configured to store instruction text data corresponding to an operation instruction for the vehicle-mounted device, wherein The determination unit compares the instruction text data stored in the storage unit with the text data converted by the voice recognition unit, and makes a determination based on whether or not there is matching instruction text data.

5. The vehicle dialogue system according to claim 4, wherein The storage unit also stores response text data corresponding to the instruction text data. When the determination unit determines that the text data represents an operation instruction for the in-vehicle device, the determination unit causes the speech synthesis unit to convert the response text data corresponding to the matched instruction text data into speech.

Citation Information

Patent Citations

  • Interaction support device, interaction system, interaction support method, and program

    JP2014098844A

  • Communication device and control method thereof

    JP2023181856A

  • Integrated fiber for optical shape sensing and spectral tissue sensing

    JP2024023172A