Information presentation device, information presentation method, information presentation program, and storage medium
The information presentation system addresses the issue of unsatisfactory user experiences by using a large-scale language model to present answers with accompanying discussions, improving user satisfaction through transparent and engaging responses.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PIONEER IP
- Filing Date
- 2025-06-26
- Publication Date
- 2026-04-23
AI Technical Summary
Existing information presentation systems fail to provide answers to user queries with sufficient persuasiveness, often confusing users with multiple options or making them feel coerced into accepting a single answer, leading to unsatisfactory user experiences.
An information presentation system utilizing a large-scale language model to generate response information that includes both the answer and the considerations behind the answer, presented in a simulated conversation between multiple virtual respondents, enhancing user satisfaction.
The system provides a more convincing and satisfying user experience by offering answers alongside the reasoning process, allowing users to make informed decisions based on simulated discussions, thereby increasing user satisfaction.
Smart Images

Figure JP2025023078_23042026_PF_FP_ABST
Abstract
Description
Information Presentation Device, Information Presentation Method, Information Presentation Program, and Storage Medium
[0001] The present invention relates to an information presentation device, an information presentation method, an information presentation program, and a storage medium.
[0002] An agent device that responds to a query from a user is known. For example, Patent Document 1 discloses an agent device having a plurality of types of agent functions. In Cited Document 1, it is described that a service provider that virtually appears in cooperation with an agent device and an agent server is an agent. Further, Cited Document 1 discloses that each of a plurality of agent function units included in the agent device and each of a plurality of agent servers cooperate to cause a plurality of agents to appear.
[0003] Japanese Unexamined Patent Application Publication No. 2020-148583
[0004] Consider a case where one or more agents are virtually caused to appear using the above-described agent device, and each agent presents an answer to a question from a user. For example, even if each agent simply provides an answer, it is assumed that the answer lacks persuasiveness for the user. For example, even if a plurality of answers are simply presented, it is assumed that the user will be confused about which one to choose. Further, for example, when only one answer is presented, it is assumed that the answer may feel forced to the user or that it may be difficult for the user to be convinced.
[0005] Thus, it is cited as one of the problems that the user's satisfaction may not be sufficiently obtained even if each agent simply provides an answer.
[0006] In view of the above points, the present invention has been made, and one of its objects is to provide an information presentation device, an information presentation method, an information presentation program, and a storage medium that enable provision of information with high user satisfaction in response to a question from a user.
[0007] The invention described in claim 1 is an information presentation device characterized by having a question acquisition unit that acquires questions from a user, and a presentation unit that presents to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating the answer to the question and including the considerations for deriving the answer.
[0008] The invention described in claim 13 is an information presentation method performed by an information presentation device, characterized by comprising: a question acquisition step of acquiring a question from a user; and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in accordance with the question, the response information indicating the answer to the question and including the considerations for deriving the answer.
[0009] The invention described in claim 14 is an information presentation program executed by an information presentation device equipped with a computer, the information presentation program causing the computer to perform a question acquisition step of acquiring a question from a user, and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating the answer to the question and including the considerations for deriving the answer.
[0010] The invention described in claim 15 is a computer-readable storage medium that stores an information presentation program for causing an information presentation device equipped with a computer to perform a question acquisition step of acquiring a question from a user, and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating the answer to the question and including the considerations for deriving the answer.
[0011] This is a schematic diagram showing the outline of an information presentation system according to Embodiment 1 of the present invention. This is a diagram showing the configuration of the front seat area of an automobile according to Embodiment 1. This is a block diagram showing an example of the configuration of an in-vehicle device according to Embodiment 1. This is a block diagram showing an example of the configuration of a server device according to Embodiment 1. This is a diagram showing an example of user information according to Embodiment 1. This is a diagram showing an example of an instruction sentence input to a large-scale language model in Embodiment 1. This is a diagram showing an example of a conversation sentence output from a large-scale language model in Embodiment 1. This is a diagram showing an example of an instruction sentence input to a large-scale language model in Embodiment 1. This is a diagram showing an example of a conversation sentence output from a large-scale language model in Embodiment 1. This is a flowchart showing an example of a routine executed by an in-vehicle device according to Embodiment 1. This is a flowchart showing an example of a routine executed by a server device according to Embodiment 1. This is a diagram showing an example of an instruction sentence input to a large-scale language model in Embodiment 2. This is a diagram showing an example of a summary sentence output from a large-scale language model in Embodiment 2. This is a diagram showing an example of an instruction sentence input to a large-scale language model in Embodiment 2. This is a diagram showing an example of a summary sentence output from a large-scale language model in Embodiment 2. This is a flowchart showing an example of a routine executed by a server device according to Embodiment 2.
[0012] Embodiments of the present invention will be described in detail below. In the following description and accompanying drawings, substantially identical or equivalent parts are denoted by the same reference numerals.
[0013] The configuration of the information presentation system 100, including the information processing device according to Example 1, will be described with reference to the attached drawings.
[0014] Figure 1 shows an overview of the configuration of the information presentation system 100. As shown in Figure 1, the information presentation system 100 consists of an in-vehicle device 10 as an information presentation device, a server 40, and a server 50 on which a Large Language Model (LLM) (hereinafter also simply referred to as LLM) is built. In Figure 1, the in-vehicle device 10 is shown mounted on an automobile M.
[0015] The in-vehicle device 10, server 40, and server 50 can send and receive data to each other via a network NW using communication protocols such as TCP / IP or UDP / IP. The network NW can be constructed using, for example, a mobile communication network, wireless communication such as Wi-Fi®, and internet communication including wired communication.
[0016] The information presentation system 100 is a system that presents information corresponding to a question when a user asks a question of a predetermined nature, and presents to the user not only the answer to the question but also the content of the considerations that may have been made in the process of arriving at that answer. In this embodiment, the information presentation system 100 presents the answer and the content of the considerations in response to a question uttered by a user who is riding in the automobile M.
[0017] The in-vehicle device 10 is a terminal device that performs information processing such as acquiring user speech from inside the vehicle M and transmitting it to the server 40, and presenting information received from the server 40 to the user.
[0018] The in-vehicle device 10 acquires the user's speech via a microphone installed inside the vehicle M. The in-vehicle device 10, for example, performs speech recognition of the user's speech and, if it determines that the question requires a predetermined type of response, transmits information indicating the question to the server 40.
[0019] Furthermore, the in-vehicle device 10 receives information generated in response to user questions from the server 40 and presents it to the user. In this embodiment, the in-vehicle device 10 presents the information received from the server 40 in voice.
[0020] Server 40 is a server device that communicates with the in-vehicle device 10 and Server 50 and performs information processing such as generating information to present to the user in response to user questions. Based on the information indicating the question received from the in-vehicle device 10, Server 40 uses a large-scale language model in Server 50 to generate information including the answer to the question and the content of the consideration.
[0021] The server 40 generates instruction sentences (prompts) in response to user inquiries and inputs them into a large-scale language model, and transmits the information output from the large-scale language model in response to the input of these instruction sentences to the in-vehicle device 10.
[0022] In this specification, information output from a large-scale language model in response to an instruction generated based on a user's question is referred to as response information. This response information indicates the answer to the question and includes the considerations used to derive that answer. For example, the server 40 transmits speech information synthesized based on the response information to the in-vehicle device 10.
[0023] Server 50 is a server device located outside of Server 40, and as described above, a large-scale language model is built within Server 50. When an instruction (prompt) is input, this large-scale language model can output responses to questions and commands contained in the prompt. In this embodiment, the large-scale language model within Server 50 is an interactive natural language processing model that takes the text information of an instruction as input and outputs text information corresponding to the instruction. For example, the large-scale language model within Server 50 is an interactive text generation AI such as GPT-4 (registered trademark) (ChatGPT), PaLM2, or Claude2.1.
[0024] The large-scale language model in server 50 receives an instruction statement generated by server 40 via the network NW, which requests the output of response information. In response to the input of this instruction statement, the response information output from the large-scale language model is transmitted via the network NW and retrieved by server 40.
[0025] In this embodiment, a large-scale language model generates response information that shows an answer to a question from the user, and includes a conversational text that simulates a discussion among multiple respondents that takes place during the process of deriving the answer, and the user is presented with audio corresponding to this response information.
[0026] In other words, the information presentation system of this embodiment is a system that can present to the user a conversation that simulates a discussion taking place between multiple virtual respondents in the process of deriving an answer to a question.
[0027] For example, when a user of the information presentation system in this embodiment is presented with response information containing multiple answers in response to a question requesting destination suggestions, they can refer to the content of a conversation simulating a discussion among multiple respondents to decide whether to choose their destination from among the multiple answers, and if so, which answer to choose. For example, when a single answer is presented, the user can also refer to the content of a conversation simulating a discussion among multiple respondents to decide whether to choose that single answer as their destination.
[0028] In the following description, this embodiment will be explained using the case where the in-vehicle device 10 is a car navigation system as an example. Furthermore, in this embodiment, this will be explained using the case where the in-vehicle device 10 is a terminal device for a so-called cloud-type car navigation system, which receives a destination that the user wishes to be guided to from the user, transmits the destination to the server 40, and the server 40 generates a route to the destination.
[0029] Figure 2 is a perspective view showing the area around the front seats of a car M equipped with the in-vehicle device 10. As an example of installation, Figure 2 shows the case where the in-vehicle device 10 is mounted on the inside of the windshield FG near the rearview mirror RM of the car M.
[0030] The GPS receiver 11 is a device that receives signals (GPS signals) from GPS (Global Positioning System) satellites. The GPS receiver 11 is located, for example, on the dashboard DB. However, the GPS receiver 11 can be located anywhere as long as it can receive GPS signals. The GPS receiver 11 can transmit the received GPS signals to the in-vehicle device 10. The in-vehicle device 10 uses the GPS signals to acquire the current location information of the vehicle M. The GPS receiver 11 may also be built into the in-vehicle device 10.
[0031] Speaker 13 is installed, for example, on the interior side of the left and right A-pillars AP. Speaker 13 is capable of emitting sounds such as music and voices based on the control of the in-vehicle device 10. For example, speaker 13 can emit voice guidance from the car navigation system. Speaker 13 can also output sound, for example, that is reproduced from voice information received by the in-vehicle device 10 from the server 40.
[0032] Figure 2 shows that speaker 13 outputs a speech-synthesized conversation simulating a discussion between multiple respondents, based on the response information. Furthermore, as shown in Figure 2, the voice of one speaker in a conversation between two respondents is emitted from speaker 13 located on the right A-pillar AP, while the voice of the other speaker is emitted from speaker 13 located on the left A-pillar AP. Thus, the speaker locations for outputting speech may differ for each speaker. Note that speaker 13 may be built into the in-vehicle device 10.
[0033] Microphone 15 is a microphone device that receives sounds from inside the vehicle and is, for example, located on the dashboard DB. Microphone 15 may be installed anywhere, such as near the rearview mirror RM, as long as it can receive sounds from inside the vehicle. Microphone 15 acquires the speech of the user inside the vehicle M. In addition, operations to change the conditions for operating the car navigation system or displaying information may be performed by voice via microphone 15.
[0034] The touch panel 17 is a touch panel monitor that combines a display, such as an LCD capable of displaying images, with a touchpad. The touch panel 17 is located, for example, on the center console of the dashboard DB. The touch panel 17 only needs to be located in a place that is visible to the driver and within the driver's reach. For example, the touch panel 17 may be mounted on the dashboard DB.
[0035] The touch panel 17 can display information based on the control of the in-vehicle device 10. For example, the touch panel 17 may display navigation instructions. The touch panel 17 may also display text information corresponding to voice information indicating response information transmitted from the server 40. Alternatively, the touch panel 17 may display an image on a map showing a location (for example, a recommended restaurant) identified by the response information.
[0036] Furthermore, the touch panel 17 can transmit signals to the in-vehicle device 10 that represent input operations received from the user to the touch panel 17. For example, operations related to car navigation functions, such as setting a destination, can be performed via the touch panel 17. In addition, various input operations, such as inputting information about the user's preferences or inputting the number of answers the user wishes to be presented with in response to a question, can be performed via the touch panel 17.
[0037] The external camera 19 is an imaging device that photographs the area in front of the automobile M. In this embodiment, the external camera 19 is built into the in-vehicle device 10. As described above, the in-vehicle device 10 is mounted on the inside of the windshield FG near the rearview mirror RM of the automobile M. The in-vehicle device 10, which incorporates the external camera 19, is mounted so that the shooting direction of the external camera 19 is in front of the automobile M.
[0038] The placement of the external camera 19 is not limited to this; for example, it may be mounted on the dashboard DB. For example, the images captured by the external camera 19 are used in the in-vehicle device 10 to acquire the driving load on the driver of the vehicle M.
[0039] Figure 3 is a block diagram showing the configuration of the in-vehicle device 10. For example, the in-vehicle device 10 is a device in which an input unit 25, a storage unit 27, a control unit 29, a communication unit 31, and an output unit 33 cooperate via a system bus 23.
[0040] In addition, an acceleration sensor 21 is mounted on the vehicle M. The acceleration sensor 21 can measure the acceleration of the vehicle M and output a signal indicating the measured acceleration. The acceleration sensor 21 is a sensor that can detect the acceleration in the traveling direction of the vehicle M, that is, the front-rear direction, when viewed from above the vehicle M. Also, the acceleration sensor can detect, for example, the acceleration in the lateral direction (width direction) perpendicular to the traveling direction of the vehicle M.
[0041] The input unit 25 is an interface unit that communicably connects the in-vehicle device 10 to the GPS receiver 11, the microphone 15, the touch panel 17, the outboard camera 19, and the acceleration sensor 21.
[0042] The in-vehicle device 10 can receive a GPS signal from the GPS receiver 11 via the input unit 25 and acquire information on the current position of the in-vehicle device 10 from the GPS signal.
[0043] The in-vehicle device 10 can acquire voice data of the voice picked up by the microphone 15 via the input unit 25. For example, the in-vehicle device 10 can receive an input operation by voice made by the user via the input unit 25. The in-vehicle device 10 can acquire the voice such as a question spoken by the user via the input unit 25.
[0044] The in-vehicle device 10 can receive a signal indicating an input operation made on the touch pad of the touch panel 17 via the input unit 25. For example, the in-vehicle device 10 can receive a setting input of the destination of car navigation made by the user with respect to the touch panel 17 via the input unit 25.
[0045] For example, the in-vehicle device 10 can receive operations related to the presentation of response information, such as an input of the desired number of answers made by the user with respect to the touch panel 17 via the input unit 25, and an input of selecting one from a plurality of locations (for example, a plurality of restaurants) displayed on the map based on the response information, via the input unit 25.
[0046] In addition, the in-vehicle device 10 can acquire an image captured by the outside vehicle camera 19 via the input unit 25. For example, the in-vehicle device 10 may estimate the magnitude of the driving load on the driver of the automobile M by image recognition using an image of the front of the automobile M acquired from the outside vehicle camera 19. For example, the in-vehicle device 10 estimates the load according to the driving situation, such as estimating that the load related to driving is higher when the visibility is poor than when the visibility is good in the front.
[0047] In addition, the in-vehicle device 10 can receive a signal indicating the acceleration measured by the acceleration sensor 21 via the input unit 25. The in-vehicle device 10 acquires, for example, the acceleration of the automobile M based on the acceleration indicated by the sensor signal of the acceleration sensor 21. The in-vehicle device 10 may acquire the current position information of the automobile M based on, for example, the acceleration indicated by the sensor signal of the acceleration sensor 21 in addition to the GPS signal from the GPS receiver 11.
[0048] The storage unit 27 is, for example, a storage device composed of a hard disk drive, an SSD (Solid State Drive), a flash memory, or the like. The storage unit 27 stores various programs executed in the in-vehicle device 10, such as an operating system and software for the terminal. The various programs include programs for executing processes related to the presentation of response information in the information presentation system in the in-vehicle device 10.
[0049] The various programs may be acquired, for example, from other server devices or the like via a network, or may be recorded on a recording medium and read via various drive devices. That is, the various programs stored in the storage unit 27 can be transmitted via a network and can also be recorded on a computer-readable recording medium and transferred.
[0050] Furthermore, the memory unit 27 stores map information, including road maps. This map information is used, for example, for the guidance display of a car navigation system. This map information is also used when acquiring the driving load on the driver of the vehicle M. For example, map information or images pre-associated with map information may be used to estimate the driving load in place of, or in addition to, the forward image from the external camera 19.
[0051] The control unit 29 is composed of a CPU (Central Processing Unit) 29A, a ROM (Read-Only Memory) 29B, a RAM (Random Access Memory) 29C, etc., and functions as a computer. The CPU 29A reads and executes various programs stored in the ROM 29B and the memory unit 27 to realize various functions.
[0052] In this embodiment, the control unit 29 performs functions such as acquiring questions uttered by the user and transmitting them to the server 40, acquiring response information to the user's questions from the server 40 and presenting it to the user, and providing a car navigation function.
[0053] Furthermore, the control unit 29 has a function to transmit the current location of the in-vehicle device 10 to the server 40. In this embodiment, the control unit 29 transmits the current location of the automobile M on which the in-vehicle device 10 is installed to the server 40 as the current location of the in-vehicle device 10. In addition, when the control unit 29 detects a predetermined question from the user's utterance and transmits it to the server 40, it transmits the current location of the automobile M along with information indicating the question.
[0054] [Question Acquisition] The control unit 29 acquires the user's utterances via the microphone 15, performs speech recognition on the user's utterances, and determines whether or not there are utterances requesting a specific type of answer. Specifically, for example, the control unit 29 is configured to determine whether or not there are specific types of questions in the user's utterances, and if so, to acquire the content of those questions.
[0055] Certain types of questions are those that request suggestions for places (spots) near the user's current location that are recommended as destinations for the in-vehicle device 10, such as, "Are there any recommended spots nearby?" Other types of questions may include those that request suggestions for routes to a destination, such as, "Can you tell me the best route to get to XX?" or those that request suggestions for souvenirs, such as, "What kind of souvenirs would you recommend buying here?" The control unit 29 functions as a question acquisition unit that acquires questions from the user.
[0056] The control unit 29, for example, converts the voice information from the user's utterance into text information, and determines that there was an utterance requesting a specific type of response if the text information contains a specific keyword. For example, utterances such as "Can you recommend a good ramen restaurant around here?" and "Where would be a good ramen restaurant to go around here?" are both determined to be utterances requesting a specific type of response.
[0057] When the control unit 29 determines that the user has made an utterance requesting a specific type of answer, i.e., a specific type of question, it generates question information and sends it to the server 40. The question information includes, for example, the terminal ID of the in-vehicle device 10, the user name, the current location information of the in-vehicle device 10 at the time the utterance was made, and utterance information that has been determined to be an utterance requesting a specific type of answer. The utterance information is, for example, text information. Based on this question information, the server 40 generates an instruction statement to cause the LLM to output response information to the user's question.
[0058] The control unit 29 may also transmit the user's utterance as voice information to the server 40 without performing speech recognition. In that case, the server 40 may perform speech recognition to determine whether or not there was an utterance requesting a specific type of response, and may output response information for the part of the utterance that asks the question.
[0059] [Presentation of Response Information] When the control unit 29 receives presentation information from the server 40 for presenting response information to a question to the user, it presents the presentation information to the user. The presentation information is information generated by the server 40, and is, for example, speech information obtained by speech synthesis of the response information output from the LLM. In this embodiment, the presentation information is an audio recording of a conversation that simulates a discussion between multiple respondents, showing the answer to the question from the user and the considerations for deriving that answer.
[0060] The control unit 29 presents the received presentation information (audio information) to the user by outputting the audio reproduced from the speaker 13. The presentation information is not limited to audio information; for example, it may include map information showing a location identified by the response information, and an image of the map may be displayed on the touch panel 17. The presentation information may also include text information, and characters based on the text information may be displayed on the touch panel 17. The control unit 29 functions as a presentation unit that presents response information to the user by presenting the presentation information received from the server 40 to the user.
[0061] For example, if speech synthesis is not performed on the server 40, the control unit 29 receives text data containing response information from the server 40, performs speech synthesis based on the received text information, and outputs the response information from the speaker 13 to present it to the user.
[0062] In this specification, information transmitted from the server 40 to the in-vehicle device 10 as information for presenting the content of the response information to the user, regardless of whether or not map information or other information is added, is referred to as "presentation information." In other words, presentation information is information that includes at least some or all of the content of the response information. The presentation format of the presentation information may be audio or image.
[0063] Furthermore, the control unit 29 may have a function to acquire the driving load on the driver of the automobile M. The control unit 29 may transmit the acquired driving load to the server 40. The driving load is acquired, for example, by workload estimation. For example, as described above, the control unit 29 performs workload estimation using image recognition of an image taken of the front of the automobile M acquired by the external camera 19, or using information such as map information.
[0064] The communication unit 31 is a communication device that sends and receives data with external devices according to instructions from the control unit 29. The communication unit 31 is, for example, a NIC (Network Interface Card) for connecting the in-vehicle device 10 to a network NW. The communication unit 31 is connected to the aforementioned network NW and sends and receives various data to and from the server 40. For example, the control unit 29 transmits the current location information of the vehicle M to the server 40 via the communication unit 31. The control unit 29 also receives presentation information from the server 40 to present response information to the user.
[0065] The output unit 33 is the output interface unit for the speaker 13 and the touch panel 17. The control unit 29 can transmit an audio signal to the speaker 13 via the output unit 33 to output audio. The control unit 29 can also transmit image data to the touch panel 17 via the output unit 33 to display it.
[0066] Figure 4 is a block diagram showing the configuration of server 40. For example, server 40 is a device in which a large-capacity storage device 43, a control unit 45, and a communication unit 47 cooperate via a system bus 41.
[0067] The server 40 has functions such as acquiring user questions from the in-vehicle device 10, generating instruction sentences to input into a large-scale language model based on the acquired questions, generating presentation information based on response information output from the large-scale language model and transmitting it to the in-vehicle device 10, and route generation.
[0068] The large-capacity storage device 43 is composed of, for example, a hard disk drive and an SSD (solid state drive), and stores various programs such as the operating system and software for the server 40. The large-capacity storage device 43 also includes various databases that store information necessary for processing on the server 40.
[0069] For example, the large-capacity storage device 43 includes a map information database (not shown) in which map information, including road maps, is stored. This map information is used, for example, for route generation. This map information is also used by the server 40 when it generates an image on a map that shows the locations (recommended spots) indicated by the response information, i.e., recommended spots as answers to the user's questions, along with the response information.
[0070] Furthermore, the large-capacity storage device 43 includes a user information database (User Information DB in Figure 4) 43A in which information about the user of the in-vehicle device 10 is stored. The User Information DB 43A stores information indicating the user's preferences and other characteristics, associated with the device ID of the in-vehicle device 10 registered in the server 40. User information is information used by the server 40 to generate instruction statements.
[0071] Figure 5 shows user information UD1, which is an example of user information stored in user information DB 43A. As shown in Figure 5, user information UD1 includes a terminal ID. The terminal ID is information that identifies each in-vehicle device 10 and is the destination information when the server 40 sends response information. The terminal ID is, for example, a MAC address.
[0072] Furthermore, user information UD1 includes the username. The username is information that can identify each user for each terminal ID, and for example, the username entered by the user during registration is stored. For example, the username is the user's real name or nickname. For example, two or more usernames may be registered for one terminal ID, and two or more terminal IDs may be registered for one username.
[0073] Furthermore, in user information UD1, characteristics including user preferences are associated and stored for each terminal ID and username. These characteristics may include user preferences such as liking spicy food and behavioral characteristics such as eating a lot or a little, as shown in Figure 5, as well as information such as user attributes and personality. These "characteristics" may be registered in advance by the user, or they may be estimated from, for example, past driving history.
[0074] Furthermore, user information UD1 includes a setting for the number of answers for each terminal ID and user ID. This number of answers specifies the number of answers to be provided in response to a user's question. For example, the number of answers is registered in advance by the user. Also, for example, the number of answers may be changed depending on the content of each question or the circumstances under which the question is asked. For example, if the question is "Please tell me two recommended tourist spots," even if the number of answers set in advance is "3," the number of answers may be temporarily changed to "2."
[0075] User information may also be stored in the storage unit 27 of the in-vehicle device 10. In that case, for example, user information about the user of the in-vehicle device 10 is included in the question information and sent to the server 40.
[0076] Furthermore, the large-capacity storage device 43 includes an instruction DB 43B in which various standard phrases used by the server 40 when generating instruction statements are stored. The instruction DB 43B stores standard phrases that serve as the basis for instruction statements used to generate response information to various specific types of anticipated questions, such as the question asking for suggestions of recommended spots around the current location or destination, as well as suggestions for recommended routes and items recommended for purchase (souvenirs, etc.).
[0077] For example, the server 40 reads an appropriate predefined phrase from the instruction database 43B according to the content of the question, adds necessary information based on user information to the predefined phrase, and completes the instruction.
[0078] Furthermore, the large-capacity storage device 43 stores a dictionary, which is used for speech recognition and keyword detection from question texts.
[0079] The control unit 45 is composed of a CPU (Central Processing Unit) 45A, ROM (Read Only Memory) 45B, RAM (Random Access Memory) 45C, etc., and functions as a computer. The CPU 45A reads and executes various programs stored in the ROM 45B and the mass storage device 43 to realize various functions.
[0080] The communication unit 47 is a communication device that sends and receives data with external devices according to instructions from the control unit 45. The communication unit 47 is, for example, a NIC (Network Interface Card) for connecting the server 40 to a network NW. The communication unit 47 is connected to the aforementioned network NW and sends and receives various data between the in-vehicle device 10 and the server 50.
[0081] For example, the control unit 45 receives question information from the in-vehicle device 10 via the communication unit 47. The control unit 45 also transmits an instruction to the server 50 via the communication unit 47. The control unit 45 also receives response information from the server 50 via the communication unit 47. The control unit 45 also transmits presentation information generated based on the response information to the in-vehicle device 10 via the communication unit 47.
[0082] [Server Functions] The functions of the control unit 45 of the server 40 will be described in detail below.
[0083] [Question Acquisition] As described above, the control unit 45 receives question information from the in-vehicle device 10. In this embodiment, as described above, the question information includes text information of an utterance that the in-vehicle device 10 has determined to be an utterance requesting a specific type of answer, and further includes the terminal ID of the in-vehicle device 10, the user name, and the current location information of the in-vehicle device 10 at the time of utterance. The control unit 45 functions as a question acquisition unit that acquires questions from the user by receiving question information from the in-vehicle device 10.
[0084] For example, if speech recognition is not performed in the in-vehicle device 10 and the user's utterance is included as voice information in the question information and sent to the server 40, the control unit 45 performs speech recognition on the received voice information and determines whether or not there was an utterance requesting a specific type of answer.
[0085] [Generation of Instructions] When the control unit 45 receives question information, it generates instructions to cause the LLM to output response information. The control unit 45 generates instructions based on, for example, the question information and user information. Specifically, for example, the control unit 45 reads a standard phrase corresponding to the content of the question from the instruction phrase DB 43B of the mass storage device 43, and generates instructions by inputting user information and conditions corresponding to the content of the question into the corresponding parts of the standard phrase.
[0086] When the control unit 45 generates an instruction, it sends the generated instruction to the server 50 and inputs the instruction into the LLM in the server 50. After that, it receives the response information output from the LLM from the server 50.
[0087] Referring to Figures 6A to 7B, we will now explain an example of instruction text generation by the control unit 45 and an example of response information output from the large-scale language model in response to the input of such instruction text. Figure 6A is a diagram showing an example of instruction text generated in response to a user's question, "Can you recommend a ramen restaurant around here?"
[0088] In the instruction text in Figure 6A, the portion F enclosed by the dashed line is a standard phrase. The standard phrase shown in Figure 6A is used to provide an answer to a question requesting a location to be suggested as a destination (place to visit) for the user, and to output a conversational text that simulates a discussion among multiple respondents during the process of deriving the answer.
[0089] The predefined text includes instructions for suggesting multiple locations and for outputting conversational text. It also includes instructions for determining multiple answers to a question and then generating response information containing the considerations used to derive those determined answers. By specifying these procedures to the LLM, for example, it becomes possible to derive more appropriate answers.
[0090] Furthermore, the standard phrases include instructions to consider the user's characteristics when determining the answer to a question. As mentioned above, user information stores characteristics, including the user's preferences, in association with them. Therefore, the instruction sentences generated by the control unit 45 may include instructions to reflect the user's preferences in the response information. For example, as shown in Figure 6A, these instruction sentences may include instructions to identify or distinguish the speaker of each statement by the speaker's name, etc., when outputting conversational text.
[0091] Furthermore, in Figure 6A, the "#User" column is set to the username associated with the received question information. The "#Characteristics" column is set to the user's characteristics stored in the user information database and associated with the username. In other words, "#User" and "#Characteristics" are replaced for each user.
[0092] Furthermore, in Figure 6A, the "#Location" column is set to indicate a location that will serve as the basis for considering recommended spots depending on the content of the question. For example, if the user's question asks for recommended spots near their current location, the current location or the area surrounding the current location will be set as the reference location in the "#Location" column.
[0093] Furthermore, for example, if the user's question is asking for recommended spots around a destination, the destination is set as the reference location in the "#Location" field. For example, the control unit 45 identifies the reference location based on place names and keywords such as "around here" and "destination" included in the question.
[0094] Furthermore, in Figure 6A, the "#Spot" column is set to the type of recommended spot requested by the user, such as "Ramen Restaurant" or "Tourist Spot," depending on the content of the question. For example, "#Spot" is set based on keywords included in the question. The type of recommended spot can be a specific type of facility, such as "Ramen Restaurant," or a broader category, such as "Restaurant."
[0095] Furthermore, the number set in the "#Number" field specifies the number of responses indicated by the response information. For example, the number of responses may be set to the number of responses pre-registered in the user information database. Alternatively, depending on the content of each question or the circumstances under which the question was asked, a different number from the number of responses pre-registered in the user information database may be set in the "#Number" field.
[0096] For example, as mentioned above, if a question includes a specified number of answers, the number of answers specified by the question will be set regardless of the number of answers already registered.
[0097] Furthermore, for example, the control unit 45 may acquire the user's driving load while operating the automobile M, and set the number of responses in the instruction message according to the magnitude of the driving load. For example, if the driving load is large, it is expected that presenting many suggestions would burden the user, so a smaller number of responses may be set than when the driving load is small.
[0098] For example, based on the number of responses recorded in the user information database, a number of responses greater than the standard number may be set when the operating load is less than a predetermined standard value, and a number of responses less than the standard number may be set when the operating load is greater than the predetermined standard value.
[0099] The driving load is acquired by the in-vehicle device 10 as described above and transmitted to the server 40. Alternatively, the control unit 45 in the server 40 may identify the image in front of the vehicle M and the road section on which the vehicle M is currently traveling, and acquire the driving load based on the characteristics of the identified road section.
[0100] For example, the control unit 45 may set the number of responses in the instruction statement according to the traffic congestion on the road the automobile M is traveling on. For example, if the road is congested, it is assumed that the time it takes to move away from the area being traveled is longer than when the road is clear, and therefore more suggestions can be presented. In that case, a larger number of responses may be set when the road is congested than when the road is clear.
[0101] For example, based on the number of responses recorded in the user information database, a smaller number of responses may be set when the road congestion level is less than a predetermined threshold, and a larger number of responses may be set when the road congestion level is greater than a predetermined threshold.
[0102] The control unit 45 may also specify the attributes or characteristics of the respondent in the instruction statement. For example, the instruction statement may include an instruction such as, "Please recreate a conversation between person A, who likes ramen, and person B, who likes spicy food." The control unit 45 may also instruct that the attributes or characteristics of the respondent be reflected in the response information. For example, an explanatory sentence such as, "This is the response from person A, who likes ramen, and person B, who likes spicy food," may be inserted before the conversation, or the instruction statement may instruct that a self-introduction such as, "I'm person A, who likes ramen," be included in the conversation.
[0103] Furthermore, although the control unit 45 specifies two respondents, A and B, in the instruction statement, it is not limited to this, and may specify, for example, three or more respondents in the instruction statement. For example, the number of respondents may be increased or decreased depending on the driving load or road congestion.
[0104] [Response Information] Figure 6B shows an example of response information output from a large-scale language model in response to the instruction sentence in Figure 6A. As shown in Figure 6B, two recommended spots (ramen restaurants) are suggested to the user. In other words, the response information in Figure 6B shows the number of answers specified in the instruction sentence to the user's question.
[0105] Furthermore, in Figure 6B, the response information is output in the form of a conversation that simulates a discussion between two respondents that takes place during the process of deriving multiple spots.
[0106] Furthermore, as shown in the conversation in Figure 6B, the response information may include a comparative analysis of which of several spots is superior.
[0107] Thus, in addition to the answer to the question, the response information is output in the form of a conversation that simulates a discussion among multiple respondents that takes place during the process of deriving multiple answers, as a consideration for how to arrive at multiple answers.
[0108] When such response information is presented, users are more likely to feel satisfied with the answer, for example, because not only the answer but also the considerations are shown. Furthermore, they can feel that multiple agents are working together to answer their question, which can lead to a sense of satisfaction. They can also feel satisfied because, by being presented with multiple answers, they can make their own decisions about their destination.
[0109] For example, if a user who has received response information sees how the answer was derived and feels that it is not the appropriate answer they were looking for, they can change their characteristic utterance, causing a new instruction sentence with the "#characteristic" changed to be generated.
[0110] For example, if a user who has been suggested a spot suitable for someone who "eats a lot" feels that they are not currently hungry and do not need such a suggestion, they can make a statement requesting that the "eating a lot" trait be removed from the user information database. This generates an instruction sentence that does not include "eating a lot" in "#traits," and new response information is obtained in response to the input of this instruction sentence into the LLM.
[0111] Figure 7A shows an example of a command generated in response to a user's question, "Please tell me some recommended tourist spots around here." The command in Figure 7A differs from the command in Figure 6A in that it instructs the system to include in the response information the process of deciding which of several answer candidates to select as the single answer to the user's question. Otherwise, it has the same structure as the command in Figure 6A.
[0112] Furthermore, the instructions in Figure 7A include instructions for a procedure in which, in determining a single answer, the system determines a specified number of answer candidates in response to a question from the user, and then generates response information that includes the process of considering one of those specified answer candidates as the final answer to the question from the user. By specifying a procedure to the LLM in this way, for example, a more appropriate answer can be derived.
[0113] Figure 7B shows an example of response information output from a large-scale language model in response to the instruction text in Figure 7A. As shown in Figure 7B, three candidate spots (tourist spots) are proposed to the user, and the response information includes the considerations leading to the conclusion that one of these spots is the most recommended. In other words, the response information in Figure 7B includes the considerations leading to the decision to select one of several candidate answers to a user's question as the single answer to that question. In short, the response information in Figure 7B ultimately shows a single answer to the user's question.
[0114] As explained in Figures 6A and 7A, the instruction sentences generated by the control unit 45 in this embodiment indicate one or more answers to a question from the user and instruct the output of a conversational text that simulates a discussion among multiple respondents that takes place in the process of deriving the one or more answers.
[0115] Furthermore, as illustrated in Figures 6B and 7B, the response information output from the LLM in response to the input of the above instruction in this embodiment is information that indicates one or more answers to a question from the user and includes conversational text that simulates a discussion among multiple respondents that takes place in the process of deriving the one or more answers.
[0116] The control unit 45 can be said to function as an instruction statement generation unit that generates an instruction statement that indicates an answer to a question and requests the output of response information that includes the considerations for deriving the answer. Alternatively, the control unit 29 of the in-vehicle device 10 may have the function of an instruction statement generation unit.
[0117] [Generation of Presentation Information] When the control unit 45 obtains response information output from the LLM from the server 50, it generates presentation information to present the content of the response information to the user. For example, the control unit 45 performs speech synthesis on the obtained response information to generate audio information of the conversational text included in the response information being read aloud, and uses this as presentation information. For example, the control unit 45 generates audio information of the conversational text included in the response information, with different voice characteristics such as tone of voice for each speaker, and uses this as presentation information.
[0118] Furthermore, the control unit 45 may generate an image showing recommended spots on a map as responses to the conversation and include it in the displayed information. In that case, for example, the displayed information may include audio or images of messages such as "Do you want to set this as your destination?" or "Which spot would you like to set as your destination?"
[0119] When the control unit 45 generates the presentation information, it transmits the generated presentation information to the in-vehicle device 10. As described above, the presentation information is presented to the user via the speaker 13 or touch panel 17 in the in-vehicle device 10.
[0120] [Control Routine] The control routines executed in the information presentation system 100 of this embodiment will be described below.
[0121] Figure 8 is a flowchart showing an information presentation routine RT1, which is an example of a control routine executed by the control unit 29 of the in-vehicle device 10. When power is turned on to the in-vehicle device 10, the control unit 29 repeatedly executes the information presentation routine RT1.
[0122] When the information presentation routine RT1 is started, the control unit 29 determines whether or not there has been an utterance requesting a specific type of response (step S101). In step S101, for example, the control unit 29 performs speech recognition on the user's utterance acquired via the microphone 15 and determines that there has been an utterance requesting a specific type of response if the user's utterance contains a specific keyword. In step S101, the control unit 29 functions as a question acquisition unit that acquires questions from the user.
[0123] In this embodiment, an utterance that seeks a specific type of response is a question asking for a place (spot) that would be recommended as a destination for the user, such as "Can you recommend a good ramen restaurant around here?", and may even be a monologue.
[0124] If the control unit 29 determines that there is no utterance requesting a specific type of response (step S101: NO), it repeatedly executes step S101 to determine again whether or not there was an utterance requesting a specific type of response.
[0125] If the control unit 29 determines that an utterance requesting a specific type of answer has been made (step S101: YES), it generates question information and sends it to the server 40 (step S102). In step S102, the control unit 29 generates question information that associates the text data of the user's utterance obtained in step S101 and the current position of the in-vehicle device 10 at the time of the utterance with the terminal ID of the in-vehicle device 10, and sends it to the server 40.
[0126] After step S102 is executed, the control unit 29 waits for a predetermined time for the server 40 to receive the presentation information and determines whether or not the presentation information has been received (step S103). As described above, for example, the presentation information is information that includes speech information obtained by speech synthesis of the response information.
[0127] If, in step S103, it is determined that the presented information has not been received (step S103: NO), the control unit 29 repeats step S103 to determine again whether or not the presented information has been received.
[0128] In step S103, if it is determined that the presentation information has been received (step S103: YES), the control unit 29 presents the received presentation information to the user in audio (step S104). In step S104, for example, if the control unit 29 receives audio information as presentation information, it outputs the audio of the audio information from the speaker 13.
[0129] In step S104, for example, audio of a conversation between multiple respondents, as illustrated in Figure 6B or Figure 7B, is output from speaker 13. For example, audio of the conversation being read aloud by different speakers with different vocal characteristics (i.e., conversational audio) is output from speaker 13. In step S104, control unit 29 functions as a presentation unit that presents response information to the user.
[0130] In step S104, for example, a conversation voice may be output, and an image showing the recommended spots indicated by the conversation voice on a map may be displayed on the touch panel 17. In this case, for example, a message such as "Do you want to set this as your destination?" or "Which spot do you want to set as your destination?" may be output as voice or displayed on the touch panel 17 along with the image. In this case, for example, when the user's input operation to set a destination is accepted, route generation by the navigation function is started through communication with the server 40.
[0131] After step S104 is executed, the control unit 29 terminates the information presentation routine RT1 and starts a new information presentation routine RT1.
[0132] Figure 9 is a flowchart of control routine RT2, which is an example of a control routine executed by the control unit 45 of the server 40. When the control unit 45 receives question information from the in-vehicle device 10, it starts control routine RT2. When control routine RT2 is started, the control unit 45 generates an instruction statement for input to the LLM, which instructs the output of a conversational sentence (step S201).
[0133] In step S201, the control unit 45 reads a standard phrase from the large-capacity storage device 43, as described above, and generates an instruction sentence that instructs the output of a conversational sentence based on the user's question included in the question information and the user information stored in the user information DB 43A. In step S201, for example, a location that will serve as a reference when considering recommended spots is set and included in the instruction sentence based on the current location of the in-vehicle device 10 or the user's question included in the question information.
[0134] Furthermore, in step S201, for example, based on the content of the user's question, the type of spot requested by the user, such as "ramen restaurant" or "tourist spot," is set.
[0135] In step S201, the control unit 45 functions as an instruction generation unit that generates an instruction sentence that indicates an answer to a question from the user and requests the output of response information that includes the considerations for deriving the answer.
[0136] After step S201 is executed, the control unit 45 inputs the instruction statement generated in step S201 into the Large-Scale Language Model (LLM) (step S202). In step S202, the control unit 45 inputs the instruction statement into the LLM on the external server 50 by sending the instruction statement to the server 50.
[0137] After step S202 is executed, the control unit 45 waits for a predetermined time for the output of response information from the LLM in the server 50 and determines whether or not the output has been obtained (step S203). If it is determined in step S203 that no output has been obtained from the LLM (step S203: NO), the control unit 45 repeats step S203 and determines again whether or not the output has been obtained from the LLM.
[0138] In step S203, if it is determined that output from LLM has been obtained (step S203: YES), the control unit 45 generates presentation information to present the obtained response information to the user (step S204). In step S204, for example, the control unit 45 generates speech information by performing speech synthesis on the response information obtained in step S203. For example, the control unit 45 generates speech information in which the speech characteristics such as tone of voice differ for each speaker for the conversational sentences included in the response information, and uses this as presentation information.
[0139] In step S204, for example, the control unit 45 may generate an image showing the recommended spots indicated by the conversation on a map and include it in the presentation information. In that case, the presentation information may include audio or images of messages such as, for example, "Do you want to set this as your destination?" or "Which spot would you like to set as your destination?".
[0140] After step S204 is executed, the control unit 45 transmits the generated presentation information to the in-vehicle device 10 (step S205). After step S205 is executed, the control unit 45 terminates the control routine RT2.
[0141] As described in detail above, in the information presentation system of this embodiment, the in-vehicle device 10 as an information presentation device includes a question acquisition unit that acquires questions from the user, and a presentation unit that presents to the user response information output from a large-scale language model in response to the input of instruction sentences generated in response to the questions, which indicates the answer to the question and includes the considerations for deriving the answer.
[0142] With this configuration, the information presentation device of the present invention can present not only the user with an answer to a user's question, but also the considerations used to derive that answer. For example, a user who receives response information from the information presentation device can appropriately utilize the answer by referring to the considerations presented along with the answer. For example, the user can accept the answer with greater satisfaction compared to when only the answer is presented.
[0143] Therefore, according to this embodiment, the information presentation device can provide an information presentation device, an information presentation method, an information presentation program, and a storage medium that can provide information that satisfies the user in response to a user's question.
[0144] For example, the information presentation device of this embodiment can present response information to the user in the form of a conversation that simulates a discussion among multiple respondents. As a result, the user can obtain answers derived through the collaboration of multiple virtual respondents, and can also understand the detailed process by which the answers were derived, thereby providing the user with more satisfying information.
[0145] Furthermore, for example, the information presentation device of this embodiment presents multiple answers along with their considerations, allowing the user to decide which answer to adopt, thereby increasing user satisfaction.
[0146] Furthermore, for example, the information presentation device of this embodiment can also present a single answer along with the considerations, thereby preventing the user from becoming confused when making a decision.
[0147] Furthermore, according to this embodiment, it is possible to implement responses from multiple respondents using a single LLM. This makes it possible to implement responses from multiple virtual respondents at a lower cost compared to, for example, using a separate server for each respondent (agent).
[0148] Referring to Figures 10A to 11B, the configuration and functions of the information presentation system 100 according to Embodiment 2 will be described. In Embodiment 2, the information presentation system 100 is configured similarly to the information presentation system 100 of Embodiment 1 in other respects, except that some of the content of the instruction statements generated by the control unit 45 of the server 40 is different.
[0149] Specifically, the instruction generated in Example 1 requests the output of a conversational text that mimics a discussion among multiple respondents, whereas the instruction in Example 2 differs in that it requests the output of a summary of the discussion's content.
[0150] Figure 10A shows an example of instructions generated in response to a user's question, "Can you recommend a ramen restaurant around here?". In the instructions in Figure 10A, the part F enclosed by the dashed line is a standard phrase. As shown in Figure 10A, this standard phrase includes instructions to output a summary that summarizes the discussion between two respondents (Person A and Person B) regarding which of several spots is the best recommendation for the user.
[0151] In Example 2, the number of respondents may be just one, and the instructions may specify that a summary of the content considered by that single respondent be output.
[0152] Figure 10B shows an example of response information output from a large-scale language model in response to the instruction text in Figure 10A. As shown in Figure 10B, the response information is output in the form of a summary sentence that summarizes the discussion between two respondents (Person A and Person B) regarding which of several spots is recommended to the user. Also, as mentioned above, in this embodiment, there may be only one respondent.
[0153] Figure 11A shows an example of instructions generated in response to a user's question, "Can you recommend some tourist spots around here?". In the instructions in Figure 11A, the part F enclosed by the dashed line is a standard phrase.
[0154] The instructions in Figure 11A include instructions to output a summary of the discussion between two respondents (Person A and Person B) regarding which of several spots to recommend to the user, leading them to the conclusion that one spot should be recommended. In this case as well, the number of respondents may be specified as one in the instructions.
[0155] Figure 11B shows an example of response information output from a large-scale language model in response to the instruction text in Figure 11A. As shown in Figure 11B, the response information is output in the form of a summary sentence that summarizes the content of the discussion between two respondents (Person A and Person B) who discussed which of several spots to recommend to the user and reached the conclusion to recommend one spot.
[0156] As explained in Figures 10A and 11A, the instruction statement generated by the control unit 45 in Embodiment 2 indicates one or more answers and instructs the output of a summary of the discussions among the multiple respondents or the considerations by one respondent that take place in the process of deriving the one or more answers.
[0157] Furthermore, as explained in Figures 10B and 11B, the response information output from the LLM in response to the input of the above instruction in Example 2 indicates one or more answers and includes information that summarizes the content of discussions among multiple respondents or considerations by one respondent that take place in the process of deriving one or more answers.
[0158] Figure 12 is a flowchart of control routine RT3, which is an example of a control routine executed by the control unit 45 of the server 40 in Embodiment 2. Control routine RT3 differs from control routine RT2 in that it generates an instruction statement that instructs the output of a summary in step S301, but proceeds in the same way as control routine RT2 for the other steps, so some explanation is omitted.
[0159] When the control unit 45 receives question information from the in-vehicle device 10, it starts the control routine RT3. When the control routine RT3 is started, the control unit 45 generates an instruction statement for input to the LLM, which instructs the output of a summary (step S301). In step S301, for example, an instruction statement like the one illustrated in Figure 10A or Figure 11A is generated.
[0160] Subsequently, similar to the control routine RT2, the control unit 45 inputs the instruction statement to the LLM in the server 50 (step S302), obtains the response information output from the LLM (step S303), and generates presentation information to present the content of the response information to the user (step S304).
[0161] For example, in step S304, audio information is generated of the voice reading the summary sentence included in the response information. For example, only one speaker is needed to read the summary sentence, and there is no need to make the voice characteristics different depending on the speaker, as in a conversation. After that, the control unit 45 transmits the presented information to the in-vehicle device 10 and terminates the control routine RT3.
[0162] As described above, according to this embodiment 2, response information including a summary of the considerations for deriving the answer can be presented to the user. This allows the user to efficiently understand, for example, the process by which the answer was derived in a short amount of time and with minimal burden.
[0163] The routines, configurations, etc., in the above-described embodiments are merely illustrative examples and can be appropriately selected, combined, or modified depending on the application.
[0164] Although this specification primarily describes answers to questions asking for recommended spots near the current location, the present invention is applicable to other questions as well. For example, it is applicable to questions such as asking for recommended spots near a destination, asking for a recommended route to a destination, or asking for recommended souvenirs to purchase near the current location or destination.
[0165] In the above embodiment, the case where the in-vehicle device 10 is a navigation device has been described, but it is not limited to this. The in-vehicle device 10 does not have to have a navigation function, and even in that case, it is possible to present the response information of the present invention to the user. The in-vehicle device 10 may be, for example, a drive recorder. Furthermore, in the above embodiment, the in-vehicle device 10 is not limited to a terminal device mounted in a vehicle, but may be, for example, a terminal device such as a smartphone, tablet, PC, or wearable device.
[0166] In the above embodiment, the driving load is obtained, for example, by performing image recognition using video footage of the front of the vehicle M. For example, workload estimation may be performed to estimate the magnitude of the driving load on the driver of vehicle M by performing image recognition using the image of the front of vehicle M, and a numerical value indicating the magnitude of the load may be calculated as the driving load as a result of said workload estimation.
[0167] In this workload estimation, the workload is estimated according to the driving situation, for example, by estimating that the workload is higher when turning right or left than when driving straight. Also, if a route for vehicle M has been generated, the workload may be estimated to be high when passing through guidance points such as intersections.
[0168] In the above embodiment, an example was described in which server 40 utilizes a large-scale language model built within an external server 50, but the embodiment is not limited to this. For example, the large-scale language model may be built on the large-capacity storage device 43 of server 40, and server 40 may input instructions to the large-scale language model within server 40 and output response information.
[0169] Furthermore, the in-vehicle device 10 may be configured to have some or all of the functions of the server 40 in the above embodiment. For example, the control unit 29 of the in-vehicle device 10 may have a function to generate instruction text or a function to generate presentation information.
[0170] 10 In-vehicle device 11 GPS receiver 13 Speaker 15 Microphone 17 Touch panel 19 External camera 21 Accelerometer 25 Input unit 27 Memory unit 29 Control unit 31 Communication unit 33 Output unit 40, 50 Server 43 Mass storage device 43A User information DB 43B Instruction DB 45 Control unit 47 Communication unit
Claims
1. An information presentation device comprising: a question acquisition unit that acquires questions from a user; and a presentation unit that presents to the user response information output from a large-scale language model in response to the input of instruction sentences generated in accordance with the questions, the response information indicating the answer to the question and including the considerations for deriving the answer.
2. The information presentation device according to claim 1, characterized in that the response information indicates one or more answers to the question and includes a conversational text that mimics a discussion among multiple respondents that takes place in the process of deriving the one or more answers as the content of consideration.
3. The information presentation device according to claim 1, characterized in that the response information indicates one or more answers to the question and includes a summary of the discussion among multiple respondents or the content of the consideration by one respondent that takes place in the process of deriving the one or more answers.
4. The information presentation device according to claim 1, characterized in that the response information indicates the number of responses specified in the instruction statement.
5. The information presentation device according to claim 4, characterized in that the response information indicates a number of responses corresponding to the operating load on the user when the user is operating the mobile body.
6. The information presentation device according to claim 4, characterized in that, when the user is moving by a mobile body, the response information indicates a number of responses corresponding to the congestion status of the road on which the mobile body is traveling.
7. The information presentation device according to claim 1, characterized in that the instruction statement includes instructions for a procedure of determining multiple answers to the question, and then generating response information that includes the considerations for deriving the multiple determined answers.
8. The information presentation device according to claim 1, characterized in that the instruction statement includes instructions for a procedure to generate response information that includes the details of the consideration process up to determining a plurality of answer candidates for the question, and then deciding that one of the determined plurality of answer candidates is the answer to the question.
9. The information presentation device according to claim 1, wherein the instruction statement includes an instruction to reflect the user's preferences in the response information.
10. The information presentation device according to claim 1 or 2, wherein the instruction statement includes an instruction to include information indicating the respective attributes or characteristics of the speakers in the response information when the response information includes a conversational sentence.
11. The information presentation device according to claim 1, characterized in that the presentation unit presents to the user a voice synthesized based on the response information.
12. The information presentation device according to claim 11, characterized in that, when the response information includes a conversational sentence, the presentation unit presents to the user a voice synthesized with different voice characteristics for each speaker.
13. An information presentation method performed by an information presentation device, comprising: a question acquisition step of acquiring a question from a user; and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating an answer to the question and including the considerations for deriving the answer.
14. An information presentation program executed by an information presentation device equipped with a computer, the program causing the computer to perform: a question acquisition step of acquiring a question from a user; and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating the answer to the question and including the considerations for deriving the answer.
15. A computer-readable storage medium that stores an information presentation program for causing a computer-equipped information presentation device to perform the following steps: a question acquisition step of acquiring a question from a user; and a presentation step of presenting to the user response information output from a large-scale language model in response to the input of an instruction sentence generated in response to the question, the response information indicating an answer to the question and including considerations for deriving the answer.
Citation Information
Patent Citations
Information notification device and information notification method
JP2019139514A
Program, method, information processing apparatus, and system
JP2024062499A
Method for providing response to client request and system for chat conversation
JP2024110942A