system

The system allows users to discern AI involvement in responses by visually indicating whether AI or a human is responsible, addressing the challenge of user identification in AI-driven interactions.

JP2026070320APending Publication Date: 2026-04-27TOYOTA JIDOSHA KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TOYOTA JIDOSHA KK
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Users find it difficult to determine whether AI is involved in the answer to their voice questions, especially in services using AI chatbots.

Method used

A system comprising a first generation means for generating response content, a second generation means for generating voice, and a display means to indicate visually whether AI or a human is involved in the response, using a vehicle-mounted and center-installed device configuration.

Benefits of technology

Enables users to identify whether AI is involved in the response, enhancing transparency and user control over the interaction mode.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070320000001_ABST
    Figure 2026070320000001_ABST
Patent Text Reader

Abstract

This makes it possible to identify whether or not AI is involved. [Solution] The system (1, 2) includes a first generation means (204) capable of generating response content to a request from a vehicle (10) user, a second generation means (205, 107) capable of generating voice including the response content to the request, and a display means (105) that displays visual information indicating that at least one of the first generation means and the second generation means has been involved in the voice response when at least one of the first generation means and the second generation means has performed at least one of generating the response content to the request and generating voice including the response content to the request.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system, and particularly to the technical field of a system for notifying whether AI (Artificial Intelligence) is involved or not.

Background Art

[0002] For example, an AI chatbot using an avatar has been proposed. As a technique for displaying an avatar, for example, a technique has been proposed for displaying a video including a CG (Computer Graphics) character of a speaker on a display screen along the configuration of dialogue program data based on structure data, content analysis data, and identification data extracted from conversation data (see Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For example, in a service that answers a user's voice question via communication means with voice, it is difficult for the user to determine whether AI is involved in the answer.

[0005] The present invention has been made in view of the above circumstances, and an object thereof is to provide a system that allows a user to identify whether AI is involved in an answer.

Means for Solving the Problems

[0006] A system according to one aspect of the present invention comprises: a first generation means capable of generating response content to a request from a vehicle user; a second generation means capable of generating voice including the response content to the request; and a display means that displays visual information indicating that at least one of the first generation means and the second generation means has been involved in the voice response when at least one of the first generation means and the second generation means has performed at least one of the generation of the response content to the request and the generation of voice including the response content to the request. [Brief explanation of the drawing]

[0007] [Figure 1] This block diagram shows an example of the system configuration according to the embodiment. [Figure 2] A flowchart illustrating an example of the operation of the system according to the embodiment. [Figure 3] This figure shows an example of an image related to the answer settings. [Figure 4] This figure shows an example of an image that will be displayed when you answer a question. [Figure 5] This block diagram shows another example of the system configuration according to the embodiment. [Modes for carrying out the invention]

[0008] Embodiments of the system will be described with reference to Figures 1 to 4.

[0009] (System Configuration) The configuration of System 1 according to this embodiment will be described with reference to Figure 1. In Figure 1, System 1 comprises a notification means identification device 100 mounted on a vehicle 10 and a notification content generation device 200 installed in a center 20. The notification means identification device 100 and the notification content generation device 200 are configured to communicate with each other via a network. For example, the vehicle 10 may be a so-called connected car. The center 20 is a center for supporting user U of the vehicle 10. For example, the center 20 may be called a contact center.

[0010] For example, a user U of vehicle 10 may make a request to the center 20 via the notification means identification device 100 to "find a nearby parking lot." The notification content generation device 200 of the center 20 may send a response to the request to the notification means identification device 100. The notification means identification device 100 may notify user U of the received response. The service provided by such a notification means identification device 100 and notification content generation device 200 will henceforth be referred to as the "agent service."

[0011] In System 1, responses to user U's requests are provided by at least one of operator O (i.e., a human) and AI at center 20. Responses to user U's requests may be provided solely by operator O or solely by AI. Responses to user U's requests may be provided by operator O creating the response content, which is then output as synthesized speech by the AI. Alternatively, responses to user U's requests may be provided by the AI ​​generating the response content, which is then read aloud by operator O.

[0012] The notification means identification device 100 includes an input unit 101, a storage unit 102, a transmission unit 103, a reception unit 104, a first output unit 105, and a second output unit 106.

[0013] The input unit 101 has a voice input function. That is, the input unit 101 can recognize the voice emitted by the user U. The input unit 101 may have a microphone to realize the voice input function. The first output unit 105 outputs visual information. For example, the first output unit 105 may be a display. The second output unit 106 outputs audio information. For example, the second output unit 106 may be a speaker. At least one of the input unit 101, the first output unit 105, and the second output unit 106 may be implemented by the vehicle 10's HMI (Human Machine Interface). That is, part of the notification means identification device 100 may be implemented by the vehicle 10's HMI.

[0014] The storage unit 102 is a storage means. For example, the storage unit 102 may be implemented by at least one of a non-volatile memory and a hard disk drive. For example, the storage unit 102 may store response setting information indicating the user U's preference for the response from the center 20. The response setting information may include at least one of the following: a preference for a response from operator O, a preference for a response from AI, and a preference for a response from both operator O and AI. The transmission unit 103 transmits the request information relating to the user U's request, which was input via the input unit 101, to the notification content generation device 200. The receiving unit 104 receives the response information relating to the response transmitted from the notification content generation device 200.

[0015] The notification content generation device 200 includes a receiving unit 201, an output unit 202, a response input unit 203, a response content generation unit 204, a response voice generation unit 205, and a response transmission unit 206. The output unit 202 and the response input unit 203 are used when operator O is involved in the response. The response content generation unit 204 and the response voice generation unit 205 are used when AI is involved in the response.

[0016] The receiving unit 201 receives request information transmitted from the notification means identification device 100. The output unit 202 may output the request information received by the receiving unit 201. For example, the output unit 202 may have a speaker and a display. The response input unit 203 has a voice input function. The response input unit 203 may have a microphone to realize the voice input function. Operator O may provide a voice response to the request related to the above request information via the response input unit 203. The response transmission unit 206 may transmit the response information related to the voice response by Operator O to the notification means identification device 100.

[0017] The request information received by the receiving unit 201 may be input to the response content generation unit 204. The response content generation unit 204 may use AI to generate a response to the request related to the request information. The AI ​​used by the response content generation unit 204 may include a trained model that generates a response to a request when a request related to the request information is input. The response content generated by the response content generation unit 204 may be output in text data format. The response content generated by the response content generation unit 204 may be input to the response voice generation unit 205. The response voice generation unit 205 may use AI to generate voice information corresponding to the response content generated by the response content generation unit 204. The AI ​​used by the response voice generation unit 205 may include a trained model that generates voice information corresponding to the response when a response content generated by the response content generation unit 204 is input. The response transmission unit 206 may transmit the voice information generated by the response voice generation unit 205 as response information to the notification means identification device 100.

[0018] Information relating to the response content generated by the response content generation unit 204 may be transmitted to the output unit 202. For example, the output unit 202 may display the response content generated by the response content generation unit 204. In this case, operator O may read aloud the response content generated by the response content generation unit 204. Response information may be generated when operator O reads aloud the response content generated by the response content generation unit 204 via the response input unit 203. The response transmission unit 206 may transmit the response information to the notification means identification device 100.

[0019] Operator O may create a response to a request related to the request information output by the output unit 202. For example, Operator O may input the response to the above request via the response input unit 203 using voice input. The information related to the voice-inputted response may be input to the response voice generation unit 205. The response voice generation unit 205 may use AI to generate voice information corresponding to the information related to the voice-inputted response. The response transmission unit 206 may transmit the voice information generated by the response voice generation unit 205 as response information to the notification means identification device 100.

[0020] (Operation of the System) Next, the operation of System 1 will be described with reference to the flowchart of FIG. 2. In FIG. 2, the first output unit 105 of the notification means identification device 100 acquires the current operation mode of the agent service based on the response setting information stored in the storage unit 102 (step S101). Next, the first output unit 105 displays an image indicating the acquired operation mode (step S102).

[0021] In the process of step S102, for example, the image 300 shown in FIG. 3 may be displayed on the display as the first output unit 105. "Content", "voice", and "character" in the image 300 respectively mean "responsible for the response content to the request of user U", "responsible for the voice when answering the request of user U", and "character displayed when answering the request of user U". Whether it is "human (for example, operator O)" or "AI" may be switched by a slider button.

[0022] In the example shown in FIG. 3, AI generates the response content to the request of user U, a human emits the voice corresponding to the response content, and the character generated by AI is displayed when the response is output. Still, user U may switch between "human" and "AI" at any timing by operating the slider button. For example, in the process of step S102, after the current operation mode is displayed and before the process of step S103, user U may switch between "human" and "AI" by operating the slider button. When user U switches between "human" and "AI", the above response setting information may be updated.

[0023] Returning to Figure 2, User U may input their request via voice through the input unit 101 of the notification means identification device 100. The input unit 101 may generate voice information as request information related to User U's request based on the voice spoken by User U. Alternatively, the input unit 101 may generate text information as request information based on the voice spoken by the user. The transmission unit 103 may transmit the request information and mode information indicating the operating mode of the agent service based on the response setting information to the notification content generation device 200. The mode information may be added to the request information. In other words, the mode information may constitute a part of the request information.

[0024] The notification content generation device 200, upon receiving request information and mode information (or request information with mode information added), determines, based on the mode information, whether or not it is "AI mode," in which the AI ​​generates the response content to user U's request (step S103). If it is determined in step S103 that it is AI mode (step S103: Yes), the received request information is input to the response content generation unit 204. In this case, the received request information does not need to be input to the output unit 202.

[0025] The response content generation unit 204 generates response content (for example, a response sentence) corresponding to the request information (step S104). Next, the notification content generation device 200 determines, based on the mode information, whether or not it is "AI mode" in which the AI ​​generates voice corresponding to the response content (step S105). If it is determined in step S105 that it is AI mode (step S105: Yes), the response content generated by the response content generation unit 204 is input to the response voice generation unit 205. The response voice generation unit 205 generates first voice information corresponding to the response content generated by the response content generation unit 204 (step S106). The response transmission unit 206 transmits the first voice information as response information to the notification means identification device 100.

[0026] In step S105, if it is determined that the system is not in AI mode (step S105: No), the response generated by the response content generation unit 204 is input to the output unit 202. The output unit 202 displays the response generated by the response content generation unit 204. Operator O then reads aloud the response generated by the response content generation unit 204 (step S107). The voice emitted by operator O when reading the response is input to the notification content generation device 200 via the response input unit 203. As a result, second voice information including the voice of operator O reading the response is generated. The response transmission unit 206 transmits the second voice information as response information to the notification means identification device 100.

[0027] In the process of step S103, if it is determined that it is not in AI mode (step S103: No), the received request information is input to the output unit 202. In this case, the received request information does not need to be input to the response content generation unit 204. The output unit 202 displays the request related to the request information. Then, operator O speaks the response content to the request (step S108). The voice emitted when operator O speaks the response content is input to the notification content generation device 200 via the response input unit 203. As a result, third voice information including the voice of operator O who spoke the response content is generated.

[0028] The notification content generation device 200 determines, based on the mode information, whether or not it is "AI mode," in which the AI ​​generates voice corresponding to the response content (step S109). If it is determined in step S109 that it is AI mode (step S109: Yes), the third voice information is input to the response voice generation unit 205. The response voice generation unit 205 performs speech recognition processing on the third voice information and outputs text information (step S110). Next, the response voice generation unit 205 generates fourth voice information corresponding to the text information output as a result of the speech recognition processing (step S111). The response transmission unit 206 transmits the fourth voice information as response information to the notification means identification device 100.

[0029] If it is determined in step S109 that the system is not in AI mode (step S109: No), the response transmission unit 206 transmits the third voice information to the notification means identification device 100 as response information.

[0030] Upon receiving the response information, the first output unit 105 of the notification means identification device 100 displays a screen corresponding to the operating mode of the agent service based on the response setting information stored in the storage unit 102 (step S112). In parallel with the processing in step S112, the second output unit 106 outputs audio related to the first audio information, second audio information, third audio information, or fourth audio information as the response information (step S113).

[0031] In the processing of step S112, for example, the image 400 shown in Figure 4 may be displayed on the display, which is the first output unit 105. The image 400 shown in Figure 4 may be displayed when the operating mode of the agent service is set to the state shown in Figure 3. As shown in Figure 4, the image 400 indicates that the response content was generated by AI, that the voice corresponding to the response content is the voice of operator O, and that the character C included in the image 400 was generated by AI. Note that if "Human" is selected for the item "Character" in the image 300 shown in Figure 3, the image corresponding to image 400 may include an image related to operator O instead of character C.

[0032] (Technical effects) In System 1, when the response from Center 20 to User U's request in Vehicle 10 is communicated to User U, an image like Image 400 (i.e., visual information) is displayed. Therefore, according to System 1, User U can identify whether or not AI is involved in the response to their request.

[0033] (modified version) A modified version of the above-described embodiment will be explained with reference to Figure 5. In Figure 5, the modified system 2 comprises a notification means identification device 10a mounted on the vehicle 10 and a notification content generation device 20a installed in the center 20. In the system 1 shown in Figure 1, the notification content generation device 200 in the center 20 comprises a response voice generation unit 205. In contrast, in the modified system 2, the notification means identification device 100a in the vehicle 10 comprises a response voice generation unit 107.

[0034] The operation of System 2 will be explained with reference to the flowchart in Figure 2. However, explanations of System 2 operations that are the same as those of System 1 described above will be omitted as appropriate.

[0035] In the process of step S105, if it is determined that the system is in AI mode (step S105: Yes), the response transmission unit 206 transmits the response content generated by the response content generation unit 204 to the notification means identification device 100a. The notification means identification device 100a inputs the received response content to the response voice generation unit 107. The response voice generation unit 107 generates fifth voice information corresponding to the received response content (step S106). The second output unit 106 outputs the voice related to the fifth voice information (step S113).

[0036] In the process of step S109, if it is determined that the system is in AI mode (step S109: Yes), the response transmission unit 206 transmits the third voice information to the notification means identification device 100a. The notification means identification device 100a inputs the received third voice information to the response voice generation unit 107. The response voice generation unit 107 performs speech recognition processing on the received third voice information and outputs text information (step S110). Next, the response voice generation unit 107 generates sixth voice information corresponding to the text information output as a result of the speech recognition processing (step S111). The second output unit 106 outputs the voice related to the sixth voice information (step S113).

[0037] (others) In systems 1 and 2 described above, a voice response is obtained to the user U's request, but the vehicle 10 may also be remotely controlled in response to the user U's request. Remote control of the vehicle 10 may include remote control performed by an operator (i.e., a human) and remote control performed by an AI. When the vehicle 10 is remotely controlled, an image may be displayed indicating that an operator or an AI is performing the remote control. With this configuration, user U can identify whether or not an AI is involved in the remote control.

[0038] Various aspects of the invention derived from the embodiments and modifications described above are described below.

[0039] A system according to one aspect of the invention comprises: a first generation means capable of generating response content to a request from a vehicle user; a second generation means capable of generating voice including the response content to the request; and a display means that displays visual information indicating that at least one of the first generation means and the second generation means has been involved in the voice response when at least one of the first generation means and the second generation means has performed at least one of the tasks of generating the response content to the request and generating voice including the response content to the request.

[0040] In the above-described embodiment, the "response content generation unit 204" corresponds to an example of the "first generation means," the "response voice generation unit 205" and the "response voice generation unit 107" correspond to an example of the "second generation means," and the "first output unit 105" corresponds to an example of the "display means."

[0041] The system may comprise a vehicle-side device mounted on the vehicle and a center-side device installed at the center, wherein the vehicle-side device may have the display means, and the center-side device may have the first generation means and the second generation means.

[0042] Alternatively, the system may comprise a vehicle-side device mounted on the vehicle and a center-side device installed at the center, wherein the vehicle-side device may have the second generation means and the display means, and the center-side device may have the first generation means.

[0043] In this system, the visual information may include at least one of the following: information indicating whether the first generating means has generated a response to the request, and information indicating whether the second generating means has generated audio containing the response to the request.

[0044] In this system, the request may include at least one of the following: information indicating whether or not the user wishes to use the first generation means, and information indicating whether or not the user wishes to use the second generation means. In the embodiment described above, "mode information" corresponds to an example of "information indicating whether or not the user wishes to use the first generation means" and "information indicating whether or not the user wishes to use the second generation means."

[0045] The present invention is not limited to the embodiments described above, and can be modified as appropriate without contradicting the gist or idea of ​​the invention as can be read from the claims and specification as a whole. Systems involving such modifications are also included within the technical scope of the present invention. [Explanation of symbols]

[0046] 1, 2... System, 10... Vehicle, 20... Center, 100, 100a... Notification means identification device, 105... First output unit, 106... Second output unit, 107, 205... Response voice generation unit, 200, 200a... Notification content generation device, 204... Response content generation unit

Claims

1. A first generation means capable of generating response content to a vehicle user's request, A second generation means capable of generating audio including the content of the response to the aforementioned request, When at least one of the first generation means and the second generation means performs at least one of generating a response to the request and generating an audio response including the response to the request, a display means that displays visual information indicating that at least one of the first generation means and the second generation means was involved in the audio response, A system equipped with these features.

2. The vehicle-side equipment mounted on the aforementioned vehicle, Center-side equipment installed at the center, Equipped with, The vehicle-side device has the display means, The center-side device has the first generation means and the second generation means. The system according to claim 1.

3. The vehicle-side equipment mounted on the aforementioned vehicle, Center-side equipment installed at the center, Equipped with, The vehicle-side device includes the second generating means and the display means, The center-side device has the first generation means The system according to claim 1.

4. The visual information includes at least one of the following: information indicating whether the first generating means has generated a response to the request; and information indicating whether the second generating means has generated audio including the response to the request. The system according to claim 1.

5. The request includes at least one of the following: information indicating whether or not the user wishes to use the first generation means, and information indicating whether or not the user wishes to use the second generation means. The system according to claim 1.

Citation Information

Patent Citations

  • Device and program for video identifying speaker and method of displaying video identifying speaker

    JP2003323628A